Tiezheng Yu
余铁铮 | LLM Post-training, Evaluation, and Agents
Hong Kong SAR
Huawei
I am a researcher at Huawei working on post-training, evaluation, and agentic systems for large language models. My research aims to make LLMs more reliable, capable, and practical for real-world deployment. I earned my Ph.D. from The Hong Kong University of Science and Technology, where I was supervised by Prof. Pascale Fung. I hold a Bachelor’s degree in Electronic Engineering from Zhejiang University.
My research focuses on developing more reliable and practical large language models (LLMs) tailored for real-world deployment scenarios. Currently, my work spans three core directions:
- LLM post-training and alignment
- Comprehensive evaluation of LLM capabilities and reliability
- LLM agents and agentic systems
For detailed information on my publications and research outputs, please refer to my Google Scholar Profile. Should you have any inquiries or wish to collaborate, feel free to contact me at tyuah@connect.ust.hk.
Opportunities
We are looking for motivated interns and full-time researchers interested in LLMs. Our team provides access to GPU computing resources and LLM APIs, with opportunities to work on both research and real-world applications.
If you are interested, please email me with your CV and a brief description of your research interests.
news
| Jul 30, 2026 | We release openPangu-2.0-Flash and openPangu-2.0-Pro, together with the accompanying technical report. Flash is a latency-oriented Ascend-native MoE model for long-context reasoning and tool use, while Pro is a 505B-parameter MoE model (18B activated parameters per token) with a 512K context window, designed for stronger general, reasoning, coding, and agentic capabilities. I was a principal contributor to the reinforcement learning (RL) component of this work. |
|---|---|
| Jun 02, 2026 | Our new paper, “What Makes Interaction Trajectories Effective for Training Terminal Agents?”, introduces TERMINAL-LEGO and shows that environment-grounded inspect-act-verify trajectories can provide more effective supervision for training terminal agents than stronger teachers’ outcome-focused traces. |
| Apr 07, 2026 | Our paper “One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment” is accepted in SIGIR 2026. It presents a meta reward modeling approach that adapts alignment to diverse individual preferences. |
| Jan 08, 2026 | We release our survey on Agent-as-a-Judge, which reviews how LLM-based agents can evaluate complex tasks through autonomous interaction, tool use, and environment feedback. |
| Dec 19, 2025 | Our paper “The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs” is accepted in Transactions on Machine Learning Research (TMLR). It examines how long chain-of-thought supervised fine-tuning and reinforcement learning interact when post-training reasoning vision-language models. |