
A Survey of Reinforcement Learning for Large Reasoning Models
We welcome everyone to open an issue for any related work we haven’t discussed, and we’ll try to address it in the next release!
🎉 News
- [2026-07-31] 🎉 First OpenRSI release: Frontis-MA1 (35B / 30B, with GGUF derivatives), the OpenMLE stack (Gym / RL / Evo), and the OpenMLE Tasks and OpenMLE SFT Traces datasets. Check it out: GitHub.
- [2026-06-25] 🎉 Our survey Self-Improving Agents in the Era of Experience: A Survey of Self- to Meta-Evolution is now available on OpenReview. Check it out: GitHub and OpenReview.
- [2025-11-05] 🔥 Excited to release our paper list about Memory for Agents, covering breakthroughs in Context Management and Learning from Experience powering self-improving AI agents. Check it out: GitHub
- [2025-10] 🎉 Honored to give talks at BAAI, Qingke Talk and Tencent Wiztalk! Here are the slides.
- [2025-09-18] 🎉 We update the full list of papers in the category structure of the survey!
- [2025-09-12] 🎉 Our survey was ranked #1 Paper of the Day on 🤗 Hugging Face Daily Papers!
- [2025-09-11] 🔥 Excited to release our RL for LRMs Survey! We’ll be updating the full list of papers in with a new category structure soon. Check it out: Paper.
- [2025-08-15] 🔥 Introducing SSRL: an investigation for Agentic Search RL without reliance on external search engine. Check it out: GitHub and Paper.
- [2025-05-27] 🔥 Introducing MARTI: A Framework for LLM-based Multi-Agent Reinforced Training and Inference. Check it out: Github.
- [2025-04-23] 🔥 Introducing TTRL: an open-source solution for online RL on data without ground-truth labels, especially test data. Check it out: Github and Paper.
- [2025-03-20] 🔥 We are excited to introduce collection of papers and projects on RL for reasoning models!
🎈 Citation
If you find this survey helpful, please cite our work:
@article{zhang2025survey,
title={A survey of reinforcement learning for large reasoning models},
author={Zhang, Kaiyan and Zuo, Yuxin and He, Bingxiang and Sun, Youbang and Liu, Runze and Jiang, Che and Fan, Yuchen and Tian, Kai and Jia, Guoli and Li, Pengfei and others},
journal={arXiv preprint arXiv:2509.08827},
year={2025}
}
📖 Contents
- A Survey of Reinforcement Learning for Large Reasoning Models
- 🎉 News
- 🎈 Citation
- 📖 Contents
- 🗺️ Overview
- 📄 Paper List
- Frontier Models
- Reward Design
- Policy Optimization
- Sampling Strategy
- Training Resource
- Static Corpus (Code)
- Static Corpus (STEM)
- Static Corpus (Math)
- Static Corpus (Agent)
- Static Corpus (Mix)
- Dynamic Environment (Rule-based)
- Dynamic Environment (Code-based)
- Dynamic Environment (Game-based)
- Dynamic Environment (Model-based)
- Dynamic Environment (Ensemble-based)
- RL Infrastructure (Primary)
- RL Infrastructure (Secondary)
- Applications
- 🌟 Acknowledgment
- ✨ Star History
🗺️ Overview
Our survey provides a comprehensive examination of Reinforcement Learning for Large Reasoning Models.

We organize the survey into five main sections:
- Foundational Components: Reward design, policy optimization, and sampling strategies
- Foundational Problems: Key debates and challenges in RL for LRMs
- Training Resources: Static corpora, dynamic environments, and infrastructure
- Applications: Real-world implementations across diverse domains
- Future Directions: Emerging research opportunities and challenges
📄 Paper List
Frontier Models
Reward Design
Generative Rewards
Dense Rewards
Unsupervised Rewards
Rewards Shaping
Policy Optimization
Policy Gradient Objective
Critic-based Algorithms
Critic-Free Algorithms
Off-policy Optimization
Off-policy Optimization (Exp replay)
Regularization Objectives
Sampling Strategy
Dynamic and Structured Sampling
Sampling Hyper-Parameters
Training Resource
Static Corpus (Code)
Static Corpus (STEM)
Static Corpus (Math)
Static Corpus (Agent)
Static Corpus (Mix)
Dynamic Environment (Rule-based)
Dynamic Environment (Code-based)
| Date | Name | Title | Paper | Github |
|---|---|---|---|---|
| 2025-06 | AgentCPM-GUI |
AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning | ||
| 2025-06 | MedAgentGym |
MedAgentGym: Training LLM Agents for Code-Based Medical Reasoning at Scale | ||
| 2025-05 | MLE-Dojo |
MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering | ||
| 2025-05 | SWE-rebench |
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents | - | |
| 2025-05 | ZeroGUI |
ZeroGUI: Automating Online GUI Learning at Zero Human Cost | ||
| 2025-04 | R2E-Gym |
R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents | ||
| 2025-03 | ReSearch |
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning | ||
| 2025-02 | MLGym |
MLGym: A New Framework and Benchmark for Advancing AI Research Agents | ||
| 2024-07 | AppWorld |
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents | [![GitHub Stars](https://img.shields.io/githu |