news

Jul 30, 2026 We release openPangu-2.0-Flash and openPangu-2.0-Pro, together with the accompanying technical report. Flash is a latency-oriented Ascend-native MoE model for long-context reasoning and tool use, while Pro is a 505B-parameter MoE model (18B activated parameters per token) with a 512K context window, designed for stronger general, reasoning, coding, and agentic capabilities. I was a principal contributor to the reinforcement learning (RL) component of this work.
Jun 02, 2026 Our new paper, “What Makes Interaction Trajectories Effective for Training Terminal Agents?”, introduces TERMINAL-LEGO and shows that environment-grounded inspect-act-verify trajectories can provide more effective supervision for training terminal agents than stronger teachers’ outcome-focused traces.
Apr 07, 2026 Our paper “One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment” is accepted in SIGIR 2026. It presents a meta reward modeling approach that adapts alignment to diverse individual preferences.
Jan 08, 2026 We release our survey on Agent-as-a-Judge, which reviews how LLM-based agents can evaluate complex tasks through autonomous interaction, tool use, and environment feedback.
Dec 19, 2025 Our paper “The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs” is accepted in Transactions on Machine Learning Research (TMLR). It examines how long chain-of-thought supervised fine-tuning and reinforcement learning interact when post-training reasoning vision-language models.