The Synergy Dilemma of Long-CoT SFT and RL
Our paper “The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs” is accepted in Transactions on Machine Learning Research (TMLR). It examines how long chain-of-thought supervised fine-tuning and reinforcement learning interact when post-training reasoning vision-language models.