Home

Welcome!

Hi, I’m Yancheng He, an LLM Researcher at Alibaba Group.

Research Interests

Currently, I am actively exploring the potential of Large Language Models (LLMs). My research is centered on pursuing the advancement of Agentic Intelligence and Reasoning Capabilities through Reinforcement Learning. I hold a firm belief that LLMs will fundamentally reshape our daily lives and the way we interact with the world, and I am keen to explore the possibilities within this shift.

Key Projects

  • ROME - An open-source agentic model for long-horizon agentic tasks
  • ROLL - High-performance RL training frameworks
  • Agentic Learning Ecosystem (ALE) - A comprehensive framework for agentic AI development

Latest News

Selected Publications

rome

Let It Flow: Agentic Crafting on Rock and Roll

ROLL TEAM

Arxiv

We introduce the Agentic Learning Ecosystem (ALE), a foundational infrastructure that optimizes the production pipeline for agentic model. And, we release ROME, an open-source agent grounded by ALE and trained on over one million trajectories.

[paper] [model]

Liteppo

LitePPO: Tricks or Traps? A Deep Dive into RL for LLM Reasoning

Zihe Liu, Jiashun Liu, Yancheng He, Weixun Wang, Jiaheng Liu, Ling Pan, Xinyu Hu, Shaopan Xiong, Ju Huang, Jian Hu, Shengyi Huang, Siran Yang, Jiamang Wang, Wenbo Su, Bo Zheng

ICLR2026

We present clear guidelines for selecting RL techniques tailored to specific setups and provide a reliable roadmap for practitioners navigating the RL for the LLM domain. Finally, we show that a minimalist combination of two techniques can unlock the learning capability of critic-free policies with a vanilla PPO loss.

[paper] [code] [机器之心]

flash

ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony

Han Lu, Zichen Liu, Shaopan Xiong, Yancheng He, Wei Gao, Yanan Wu, Weixun Wang

Arxiv

We present ROLL Flash, a system that extends ROLL with native support for asynchronous RL post-training. ROLL Flash is built upon two core design principles: fine-grained parallelism and rollout-train decoupling.

[paper] [code] [机器之心]

aysnppo

Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning

Jiashun Liu, Johan Obando-Ceron, Han Lu, Yancheng He, Weixun Wang, Wenbo Su, Bo Zheng, Pablo Samuel Castro, Aaron Courville, Ling Pan

ICLR2026

We revisit this bottleneck from an architectural perspective and introduce Asymmetric Proximal Policy Optimization (AsyPPO), a simple and scalable framework that restores the critics role while remaining efficient in large-model settings.

[paper] [机器之心]

think-j

Think-J: Learning to Think for Generative LLM-as-a-Judge

Hui Huang, Yancheng He, Hongli Zhou, Rui Zhang, Wei Liu, Weixun Wang, Jiaheng Liu, Wenbo Su

AAAI2025

Although generative LLMs have made substantial progress in various tasks, their performance as LLM-Judge still falls short of expectations. In this work, we propose Think-J, which improves generative LLM-as-a-Judge by learning how to think.

[paper] [code]

MuSC

MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training

Hui Huang, Jiaheng Liu, Yancheng He, Shilong Li, Bing Xu, Conghui Zhu, Muyun Yang, Tiejun Zhao

ACL2025

We propose a Multi-granularity Self-Contrastive Training (MuSC) framework, to improve the complex instruction alignment without relying on a stronger model.

[paper] [code]

deltabench

Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?

Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Zhongyuan Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, Bo Zheng

ACL2025

We introduce the DeltaBench including the generated long CoTs from different o1-like models for different reasoning tasks, to measure the ability to detect errors in long COT reasoning. Based on DeltaBench, we first perform fine-grained analysis of the generated long CoTs to discover the effectiveness and efficiency of different o1-like models.

[paper] [github]

Feel free to explore my page for more on our latest research. If you are interested in our research topics, please feel free to reach out via email.