Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
TL;DR: We systematically compare Merge, Mix RL, and MOPD for consolidating domain-specific RLVR capabilities into unified models.
Siye Wu (伍思烨) is a third-year M.S. student in Computer Science at Fudan University, advised by Prof. Yanghua Xiao.
His research interests lie in natural language processing (NLP) and large language models (LLMs), with a focus on:
Research Intern on Post-training, Tencent Hunyuan, LLM Department
Research Intern on Post-training, StepFun, Post-Train & Agent GroupSee full list on Google Scholar .
TL;DR: We systematically compare Merge, Mix RL, and MOPD for consolidating domain-specific RLVR capabilities into unified models.
TL;DR: We propose CODA, which scales reasoning by difficulty, reducing overthinking on easy tasks while promoting deeper reasoning on hard ones.
TL;DR: Step 3.5 Flash is our most capable open-source foundation model, built for frontier reasoning and agentic tasks with exceptional efficiency.
TL;DR: We propose ARM, a reasoning model capable of adaptively selecting appropriate reasoning formats based on the task at hand.
TL;DR: We present a comprehensive study of the robustness of LLMs to different types of irrelevant information under various conditions.