← Open Source
TsinghuaC3I

Awesome-RL-for-LRMs

A Survey of Reinforcement Learning for Large Reasoning Models

ListsPaper collectionsTeX
Open on GitHub
Momentum
+0stars in 24 hours0.0%
2.49k
Stars
135
Forks
+0
This week
28
Contributors
Created 2025-03-20 · Updated 2026-09-29 · #8452 today
Top developers
README

A Survey of Reinforcement Learning for Large Reasoning Models

Awesome Survey Github HF Papers Twitter

We welcome everyone to open an issue for any related work we haven’t discussed, and we’ll try to address it in the next release!

🎉 News

  • [2026-07-31] 🎉 First OpenRSI release: Frontis-MA1 (35B / 30B, with GGUF derivatives), the OpenMLE stack (Gym / RL / Evo), and the OpenMLE Tasks and OpenMLE SFT Traces datasets. Check it out: GitHub.
  • [2026-06-25] 🎉 Our survey Self-Improving Agents in the Era of Experience: A Survey of Self- to Meta-Evolution is now available on OpenReview. Check it out: GitHub and OpenReview.
  • [2025-11-05] 🔥 Excited to release our paper list about Memory for Agents, covering breakthroughs in Context Management and Learning from Experience powering self-improving AI agents. Check it out: GitHub
  • [2025-10] 🎉 Honored to give talks at BAAI, Qingke Talk and Tencent Wiztalk! Here are the slides.
  • [2025-09-18] 🎉 We update the full list of papers in the category structure of the survey!
  • [2025-09-12] 🎉 Our survey was ranked #1 Paper of the Day on 🤗 Hugging Face Daily Papers!
  • [2025-09-11] 🔥 Excited to release our RL for LRMs Survey! We’ll be updating the full list of papers in with a new category structure soon. Check it out: Paper.
  • [2025-08-15] 🔥 Introducing SSRL: an investigation for Agentic Search RL without reliance on external search engine. Check it out: GitHub and Paper.
  • [2025-05-27] 🔥 Introducing MARTI: A Framework for LLM-based Multi-Agent Reinforced Training and Inference. Check it out: Github.
  • [2025-04-23] 🔥 Introducing TTRL: an open-source solution for online RL on data without ground-truth labels, especially test data. Check it out: Github and Paper.
  • [2025-03-20] 🔥 We are excited to introduce collection of papers and projects on RL for reasoning models!

🎈 Citation

If you find this survey helpful, please cite our work:

@article{zhang2025survey,
  title={A survey of reinforcement learning for large reasoning models},
  author={Zhang, Kaiyan and Zuo, Yuxin and He, Bingxiang and Sun, Youbang and Liu, Runze and Jiang, Che and Fan, Yuchen and Tian, Kai and Jia, Guoli and Li, Pengfei and others},
  journal={arXiv preprint arXiv:2509.08827},
  year={2025}
}

📖 Contents

🗺️ Overview

Our survey provides a comprehensive examination of Reinforcement Learning for Large Reasoning Models.

![Overview of RL for LRMs Survey](figs/teaser.png) 

We organize the survey into five main sections:

  1. Foundational Components: Reward design, policy optimization, and sampling strategies
  2. Foundational Problems: Key debates and challenges in RL for LRMs
  3. Training Resources: Static corpora, dynamic environments, and infrastructure
  4. Applications: Real-world implementations across diverse domains
  5. Future Directions: Emerging research opportunities and challenges

📄 Paper List

Frontier Models

Date Name Title Paper Github
2025-08 Intern-S1 Intern-S1: A Scientific Multimodal Foundation Model Paper GitHub Stars
2025-08 GLM-4.5 GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Paper GitHub Stars
2025-08 gpt-oss gpt-oss-120b & gpt-oss-20b Model Card Paper GitHub Stars
2025-08 InternVL3.5 InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Paper GitHub Stars
2025-07 Kimi K2 Kimi K2: Open Agentic Intelligence Paper GitHub Stars
2025-07 Step 3 Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding Paper GitHub Stars
2025-07 GLM-4.1V-Thinking GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Paper GitHub Stars
2025-07 Skywork-R1V3 Skywork-R1V3 Technical Report Paper GitHub Stars
2025-07 GLM-4.5V GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Paper GitHub Stars
2025-06 Magistral Magistral Paper -
2025-06 Minimax-M1 MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Paper GitHub Stars
2025-05 MiMo MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining Paper GitHub Stars
2025-05 Qwen3 Qwen3 Technical Report Paper GitHub Stars
2025-05 Llama-Nemotron-Ultra Llama-Nemotron: Efficient Reasoning Models Paper GitHub Stars
2025-05 INTELLECT-2 INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning Paper -
2025-05 Hunyuan-TurboS Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought Paper GitHub Stars
2025-05 Skywork OR-1 Skywork Open Reasoner 1 Technical Report Paper GitHub Stars
2025-04 Phi-4 Reasoning Phi-4-reasoning Technical Report Paper -
2025-04 Skywork-R1V2 Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning Paper GitHub Stars
2025-04 InternVL3 InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Paper GitHub Stars
2025-03 ORZ Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model Paper GitHub Stars
2025-01 DeepSeek-R1 DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning Paper GitHub Stars
- QwQ QwQ-32B: Embracing the Power of Reinforcement Learning Blog GitHub Stars
- Seed-OSS Seed-OSS Open-Source Models Paper GitHub Stars
- ERNIE-4.5-Thinking ERNIE 4.5 Technical Report Blog -

Reward Design

Generative Rewards

Date Name Title Paper Github
2026-06 - Steer, Don't Solve: Training Small Critic Models for Large Code Agents Paper GitHub Stars
2025-08 CAPO CAPO: Towards Enhancing LLM Reasoning through Verifiable Generative Credit Assignment Paper GitHub Stars
2025-08 CompassVerifier CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward Paper GitHub Stars
2025-08 Cooper Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models Paper GitHub Stars
2025-08 ReviewRL ReviewRL: Towards Automated Scientific Review with RL Paper GitHub Stars
2025-08 Rubicon Reinforcement Learning with Rubric Anchors Paper -
2025-08 RuscaRL Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning Paper -
2025-07 OMNI-THINKER OMNI-THINKER: Scaling Cross-Domain Generalization in LLMs via Multi-Task RL with Hybrid Rewards Paper -
2025-07 URPO URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Paper -
2025-07 RaR Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Paper -
2025-07 RLCF Checklists Are Better Than Reward Models For Aligning Language Models Paper -
2025-07 PCL Post-Completion Learning for Language Models Paper -
2025-07 K2 KIMI K2: OPEN AGENTIC INTELLIGENCE Paper -
2025-07 LIBRA LIBRA: ASSESSING AND IMPROVING REWARD MODEL BY LEARNING TO THINK Paper -
2025-07 TP-GRPO Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner Paper GitHub Stars
2025-06 RewardAnything RewardAnything: Generalizable Principle-Following Reward Models Paper Blog
2025-06 Writing-Zero Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards Paper -
2025-06 Critique-GRPO Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback Paper GitHub Stars
2025-06 PAG PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Paper -
2025-06 GRAM GRAM: A Generative Foundation Reward Model for Reward Generalization Paper GitHub Stars
2025-06 ProxyReward From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation Paper -
2025-06 QA-LIGN QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA Paper -
2025-05 RM-R1 RM-R1: Reward Modeling as Reasoning Paper GitHub Stars
2025-05 J1 J1: Incentivizing Thinking in LLM-as-a-Judge via RL Paper -
2025-05 TinyV TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Paper GitHub Stars
2025-05 General-Reasoner General-reasoner: Advancing llm reasoning across all domains Paper -
2025-05 RRM Reward Reasoning Model Paper -
2025-05 RL Tango RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning Paper GitHub Stars
2025-05 Think-RM Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Paper GitHub Stars
2025-04 JudgeLRM JudgeLRM: Large Reasoning Models as a Judge Paper GitHub Stars
2025-04 GenPRM GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning Paper GitHub Stars
2025-04 DeepSeek-GRM Inference-Time Scaling for Generalist Reward Modeling Paper -
2025-04 AIR AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset Paper -
2025-04 Pairwise-RL A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization Paper -
2025-04 xVerify xVerify: Efficient Answer Verifier for Reasoning Model Evaluations Paper GitHub Stars
2025-04 Seed-Thinking-v1.5 Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning Paper -
2025-04 ThinkPRM Process Reward Models That Think Paper GitHub Stars
2025-03 - Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains Paper -
2025-02 - Self-rewarding correction for mathematical reasoning Paper GitHub Stars
2024-10 GenRM Generative Reward Models Paper -
2024-08 CLoud Critique-out-Loud Reward Models Paper GitHub Stars
2024-08 Generative Verifier Generative Verifiers: Reward Modeling as Next-Token Prediction Paper -
2024-01 Self-Rewarding LM Self-Rewarding Language Models Paper -
2023-10 Auto-J Generative Judge for Evaluating Alignment Paper GitHub Stars
2023-06 Judge LLM-as-a-Judge Judging llm-as-a-judge with mt-bench and chatbot arena Paper GitHub Stars

Dense Rewards

Date Name Title Paper Github
2026-09 DRACO DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training Paper GitHub Stars
2025-09 Tree-GRPO Tree Search for LLM Agent Reinforcement Learning Paper GitHub Stars
2025-09 AttnRL Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models Paper GitHub Stars
2025-09 TARL Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents Paper -
2025-09 PROF Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Paper GitHub Stars
2025-09 HICRA Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning Paper -
2025-08 KlearReasoner Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization Paper GitHub Stars
2025-08 CAPO CAPO: Towards Enhancing LLM Reasoning through Verifiable Generative Credit Assignment Paper GitHub Stars
2025-08 GTPO & GRPO-S GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy Paper -
2025-08 VSRM Promoting Efficient Reasoning with Verifiable Stepwise Reward Paper -
2025-08 G-RA Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards Paper -
2025-08 SSPO SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression Paper -
2025-08 AIRL-S Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS Paper -
2025-08 TreePO TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling Paper GitHub Stars
2025-08 MUA-RL MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Paper -
2025-07 SPRO Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Paper -
2025-07 FR3E First Return, Entropy-Eliciting Explore Paper -
2025-07 ARPO Agentic Reinforced Policy Optimization Paper GitHub Stars
2025-07 TP-GRPO Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner Paper GitHub Stars
2025-06 TreeRPO TreeRPO: Tree Relative Policy Optimization Paper GitHub Stars
2025-06 TreeRL TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Paper GitHub Stars
2025-06 Entropy Advantage Reasoning with Exploration: An Entropy Perspective on Reinforcement Learning for LLMs Paper -
2025-06 ReasonFlux-PRM ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs Paper GitHub Stars
2025-05 S-GRPO S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models Paper -
2025-05 GiGPO Group-in-Group Policy Optimization for LLM Agent Training Paper GitHub Stars
2025-05 - Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Credit Assignment Paper -
2025-05 Tango RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning Paper GitHub Stars
2025-05 StepSearch StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization Paper GitHub Stars
2025-05 - Aligning Dialogue Agents with Global Feedback via Large Language Model Reward Decomposition Paper -
2025-05 Tool-Star Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Paper GitHub Stars
2025-05 SPA-RL SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution Paper GitHub Stars
2025-05 SPO Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Mode Paper GitHub Stars
2025-04 GenPRM GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning Paper GitHub Stars
2025-04 PURE Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning Paper GitHub Stars
2025-03 MRT Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning Paper GitHub Stars
2025-03 SWEET-RL SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks Paper GitHub Stars
2025-02 PRIME Process Reinforcement through Implicit Rewards Paper GitHub Stars
2024-12 Implicit PRM Free Process Rewards without Process Labels Paper GitHub Stars
2024-10 VinePPO VinePPO: Refining Credit Assignment in RL Training of LLMs Paper GitHub Stars
2024-10 PAV Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning Paper -
2024-04 - From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function Paper -
2024-03 GELI Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback Paper -
2023-12 Math-Shepherd Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations Paper -
2023-05 PRM800K Let's Verify Step by Step Paper GitHub Stars
2022-11 - Solving math word problems with process- and outcome-based feedback Paper -

Unsupervised Rewards

Date Name Title Paper Github
2025-09 Vision-Zero Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play Paper GitHub Stars
2025-08 Co-Reward Co-Reward: Self-supervised Reinforcement Learning for Large Language Model Reasoning via Contrastive Agreement Paper GitHub Stars
2025-08 SQLM Self-Questioning Language Models Paper GitHub Stars
2025-08 R-zero R-Zero: Self-Evolving Reasoning LLM from Zero Data Paper GitHub Stars
2025-08 ETTRL ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism Paper -
2025-07 RLSF Post-Training Large Language Models via Reinforcement Learning from Self-Feedback Paper -
2025-06 RLSC Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models Paper -
2025-06 RPT Reinforcement Pre-Training Paper -
2025-06 CoVo Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Paper GitHub Stars
2025-06 SEAL Self-Adapting Language Models Paper -
2025-06 Spurious Rewards Spurious Rewards: Rethinking Training Signals in RLVR Paper GitHub Stars
2025-06 No Free Lunch No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Paper -
2025-05 Absolute Zero Absolute Zero: Reinforced Self-play Reasoning with Zero Data Paper GitHub Stars
2025-05 EM-RL The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Paper GitHub Stars
2025-05 SSR-Zero SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation Paper GitHub Stars
2025-05 - Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers Paper GitHub Stars
2025-05 RLIF Learning to Reason without External Rewards Paper GitHub Stars
2025-05 SeRL SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data Paper GitHub Stars
2025-05 SRT Can Large Reasoning Models Self-Train? Paper GitHub Stars
2025-05 RENT-RL Maximizing Confidence Alone Improves Reasoning Paper GitHub Stars
2025-04 EMPO Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization Paper GitHub Stars
2025-04 TRANS-ZERO TRANS-ZERO: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data Paper GitHub Stars
2025-04 TTRL TTRL: Test-Time Reinforcement Learning Paper GitHub Stars
2025-04 One-Shot-RLVR Reinforcement Learning for Reasoning in Large Language Models with One Training Example Paper GitHub Stars
2025-02 CAGSR A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals Paper -
2024-07 MINIMO Learning Formal Mathematics From Intrinsic Motivation Paper GitHub Stars

Rewards Shaping

Date Name Title Paper Github
2025-09 CDE CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Paper -
2025-09 DARLING Jointly Reinforcing Diversity and Quality in Language Model Generations Paper GitHub Stars
2025-09 DRER Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL Paper -
2025-09 OBE Outcome-based Exploration for LLM Reasoning Paper -
2025-08 Pass@kTraining Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Paper GitHub Stars
2025-05 PKPO Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems Paper -
2025-05 rl-without-gt Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers Paper GitHub Stars
2025-03 CrossDomain-RLVR Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains Paper -
2025-01 DeepSeek-R1 DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning Paper GitHub Stars
2024-09 Qwen2.5-Math Qwen2. 5-math technical report: Toward mathematical expert model via self-improvement Paper GitHub Stars

Policy Optimization

Policy Gradient Objective

Date Name Title Paper Github
2017-07 PPO Proximal policy optimization algorithms Paper -
- PG Policy gradient methods for reinforcement learning with function approximation. Paper -
- REINFORCE Simple statistical gradient-following algorithms for connectionist reinforcement learning Paper -
- TRPO Trust region policy optimization Paper -

Critic-based Algorithms

Date Name Title Paper Github
2025-08 VL-DAC Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success Paper GitHub Stars
2025-08 VRPO VRPO:Rethinking Value Modeling for Robust RL Training under Noisy Supervision Paper -
2025-05 VerIPO VerIPO: Long Reasoning Video-R1 Model with Iterative Policy Optimization Paper GitHub Stars
2025-04 VAPO Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks Paper -
2025-03 VCPPO What’s Behind PPO’s Collapse in Long-CoT? Value Optimization Holds the Secret Paper -
2025-03 Open reasoner-zero open reasoner-zero: An open source approach to scaling up reinforcement learning on the base model Paper GitHub Stars
2025-02 PRIME PROCESS REINFORCEMENT THROUGH IMPLICIT REWARDS Paper GitHub Stars
2024-12 Implicit PRM FREE PROCESS REWARDS WITHOUT PROCESS LABELS Paper GitHub Stars
2023-12 Math-shepherd Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations Paper -
2015-06 GAE High-dimensional continuous control using generalized advantage estimation Paper -
- Autopsv Autopsv: Automated process-supervised verifier. Paper GitHub Stars

Critic-Free Algorithms

Date Name Title Paper Github
2025-09 UPGE Towards a Unified View o fLarge Language Model Post-Training Paper GitHub Stars
2025-09 SPO Single-stream Policy Optimization Paper -
2025-08 LitePPO Part I: Tricks or Traps? A Deep Dive into RLfor LLM Reasoning Paper -
2025-07 R1-RE R1-RE: Cross-Domain Relation Extraction with RLVR Paper -
2025-07 GSPO Group Sequence Policy Optimization Paper -
2025-06 CISPO MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Paper GitHub Stars
2025-05 KRPO Kalman Filter Enhanced Group Relative Policy Optimization for Language Model Reasoning Paper GitHub Stars
2025-05 CPGD CPGD:Toward Stable Rule-based Reinforcement Learning for Language Models Paper GitHub Stars
2025-05 NFT Bridging Supervised Learning and Reinforcement Learning in Math Reasoning Paper -
2025-05 Clip-Cov/KL-Cov The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models Paper GitHub Stars
2025-03 OpenVLThinker OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Paper GitHub Stars
2025-03 DAPO DAPO: an Open-Source LLM Reinforcement Learning System at Scale Paper GitHub Stars
2025-03 Dr. GRPO Understanding R1-Zero-Like Training: A Critical Perspectiv Paper GitHub Stars
2025-01 Kimi k1.5 Kimi k1.5: Scaling Reinforcement Learning with LLMs Paper -
2024-02 RLOO Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms Paper -
2024-02 GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Paper GitHub Stars
2023-10 ReMax ReMax: A Simple, Effective, and Efficient Method for Aligning Large Language Models Paper GitHub Stars
- REINFORCE Simple statistical gradient-following algorithms for connectionist reinforcement learning Paper -
- REINFORCE++ REINFORCE++: An Efficient RLHF Algorithm with Robustnessto Both Prompt and Reward Models Paper GitHub Stars
- VinePPO VINEPPO: UNLOCKING RL POTENTIAL FOR LLM REASONING THROUGH REFINED CREDIT ASSIGNMENT Paper GitHub Stars
- FlashRL Fast RL training with Quantized Rollouts Paper GitHub Stars

Off-policy Optimization

Date Name Title Paper Github
2025-09 BRIDGE Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning Paper GitHub Stars
2025-09 HPT Towards a Unified View of Large Language Model Post-Training Paper GitHub Stars
2025-08 DFT On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification Paper GitHub Stars
2025-08 RED Recall-Extend Dynamics: Enhancing Small Language Models through Controlled Exploration and Refined Offline Integration Paper GitHub Stars
2025-07 Prefix‑RFT Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling Paper -
2025-07 ReMix Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Paper GitHub Stars
2025-06 ReLIFT Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions Paper GitHub Stars
2025-06 BREAD BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Paper -
2025-06 SRFT SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Paper -
2025-05 AMPO Adaptive Thinking via Mode Policy Optimization for Social Language Agents Paper GitHub Stars
2025-05 UFT UFT: Unifying Supervised and Reinforcement Fine-Tuning Paper GitHub Stars
2025-04 LUFFY Learning to Reason under Off-Policy Guidance Paper GitHub Stars
2025-03 SPO Soft Policy Optimization: Online Off-Policy RL for Sequence Models Paper GitHub Stars
2025-03 TOPR TAPERED OFF-POLICY REINFORCE Stable and efficient reinforcement learning for LLMs Paper -
2024-05 IFT Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process Paper GitHub Stars
2023-05 DPO Direct Preference Optimization: Your Language Model is Secretly a Reward Model Paper -
2015-11 - Fixed point quantization of deep convolutional networks Paper -
- - Your Efficient RL Framework Secretly Brings You Off-Policy RL Training Paper GitHub Stars

Off-policy Optimization (Exp replay)

Date Name Title Paper Github
2025-09 SAPO Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing Paper -
2025-09 SEELE Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding Paper GitHub Stars
2025-08 Memory-R1 Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning Paper -
2025-07 RLEP RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Paper GitHub Stars
2025-06 EFRame EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework Paper GitHub Stars
2025-05 ARPO ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Paper GitHub Stars
2025-04 - Improving RL Exploration for LLM Reasoning through Retrospective Replay Paper -

Regularization Objectives

Date Name Title Paper Github
2025-10 ASPO ASPO: Asymmetric Importance Sampling Policy Optimization Paper GitHub Stars
2025-09 CE-GPPO CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning Paper GitHub Stars
2025-09 CDE CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models Paper -
2025-09 DPH RL The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward Paper GitHub Stars
2025-09 empgseed-seed Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Paper -
2025-07 Archer Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Paper GitHub Stars
2025-06 Bingo Bingo: Boosting Efficient Reasoning of LLMs via Dynamic and Significance-based Reinforcement Learning Paper GitHub Stars
2025-06 HighEntropy RL Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Paper -
2025-06 Entropy RL Reasoning with Exploration: An Entropy Perspective on Reinforcement Learning for LLMs Paper -
2025-06 ALP RL Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Paper -
2025-05 DisCO DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization Paper GitHub Stars
2025-05 Skywork OR1 Skywork Open Reasoner 1 Technical Report Paper GitHub Stars
2025-05 Entropy Mechanism The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models Paper GitHub Stars
2025-05 ProRL ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Paper -
2025-05 Short RL Efficient RL Training for Reasoning Models via Length-Aware Optimization Paper GitHub Stars
2025-03 DAPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale Paper -
2025-03 L1 L1: Controlling how long a reasoning model thinks with reinforcement learning Paper GitHub Stars

Sampling Strategy

Dynamic and Structured Sampling

Date Name Title Paper Github
2025-10 EEPO EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget Paper GitHub Stars
2025-09 AttnRL Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models Paper -
2025-09 DACE Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Paper -
2025-09 Parallel-R1 Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Paper GitHub Stars
2025-08 G^2RPO-A G^2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidanc Paper GitHub Stars
2025-08 RuscaRL Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning Paper -
2025-08 TreePO TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling Paper GitHub Stars
2025-07 ARPO Agentic Reinforced Policy Optimization Paper GitHub Stars
2025-06 TreeRPO TreeRPO: Tree Relative Policy Optimization Paper GitHub Stars
2025-06 E2H Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning Paper -
2025-06 TreeRL TreeRL: LLM Reinforcement Learning with On-Policy Tree Search Paper GitHub Stars
2025-05 ToTRL ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving Paper -
2025-03 DARS DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal Paper GitHub Stars
2025-03 DAPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale Paper GitHub Stars
2025-02 PRIME Process Reinforcement through Implicit Rewards Paper GitHub Stars
- POLARIS POLARIS: A POst-training recipe for scaling reinforcement Learning on Advanced ReasonIng modelS Blog GitHub Stars

Sampling Hyper-Parameters

Date Name Title Paper Github
2025-08 GFPO Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning Paper -
2025-06 AceReason-Nemotron 1.1 AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy Paper -
2025-06 T-PPO Truncated Proximal Policy Optimization Paper -
2025-06 Confucius3-Math Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning Paper GitHub Stars
2025-05 E3-RL4LLMs Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs Paper GitHub Stars
2025-05 AceReason-Nemotron AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Paper -
2025-05 Pro-RL ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Paper -
2025-03 - Output Length Effect on DeepSeek-R1's Safety in Forced Thinking Paper -
2025-03 DAPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale Paper GitHub Stars
2025-02 PRIME Process Reinforcement through Implicit Rewards Paper GitHub Stars
2025-02 - Training Language Models to Reason Efficiently Paper -
- DeepScaleR DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL Paper GitHub Stars
- POLARIS POLARIS: A POst-training recipe for scaling reinforcement Learning on Advanced ReasonIng modelS Paper GitHub Stars

Training Resource

Static Corpus (Code)

Date Name Title Paper Github
2025-05 rStar-Coder rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset Paper GitHub Stars
2025-04 Z1 Z1: Efficient Test-time Scaling with Code Paper GitHub Stars
2025-04 OpenCodeReasoning OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Paper -
2025-04 LeetCodeDataset LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs Paper GitHub Stars
2025-03 KodCode KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding Paper -
2025-01 SWE-Fixer SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution Paper GitHub Stars
2024-12 SWE-Gym Training Software Engineering Agents and Verifiers with SWE-Gym Paper GitHub Stars
- Code-R1 Code-R1: Reproducing R1 for Code with Reliable Rewards Paper GitHub Stars
- codeforces-cots CodeForces CoTs Paper -
- DeepCoder DeepCoder: A Fully Open-Source 14B Coder at O3-mini Level Blog GitHub Stars

Static Corpus (STEM)

Date Name Title Paper Github
2025-09 SSMR-Bench Synthesizing Sheet Music Problems for Evaluation and Reinforcement Learning Paper GitHub Stars
2025-09 Loong Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers Paper GitHub Stars
2025-07 MegaScience MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning Paper -
2025-06 ReasonMed ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning Paper GitHub Stars
2025-05 ChemCoTDataset Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations Paper -
2025-02 NaturalReasoning NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions Paper -
2025-01 SCP-116K SCP-116K: A High-Quality Problem-Solution Dataset and a Generalized Pipeline for Automated Extraction in the Higher Education Science Domain Paper -

Static Corpus (Math)

Date Name Title Paper Github
2025-07 MiroMind-M1-RL-62K MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Paper GitHub Stars
2025-04 DeepMath DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning Paper GitHub Stars
2025-04 OpenMathReasoning AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset Paper GitHub Stars
2025-03 STILL-3-RL An Empirical Study on Eliciting and Improving R1-like Reasoning Models Paper GitHub Stars
2025-03 Light-R1 Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond Paper -
2025-03 DAPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale Paper GitHub Stars
2025-03 OpenReasoningZero Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model Paper GitHub Stars
2025-02 PRIME Process Reinforcement through Implicit Rewards Paper GitHub Stars
2025-02 LIMO Limo: Less is more for reasoning Paper GitHub Stars
2025-02 LIMR Limr: Less is more for rl scaling Paper GitHub Stars
2025-02 Big-MATH Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Paper -
- NuminaMath 1.5 Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions Paper GitHub Stars
- OpenR1-Math Open R1: A fully open reproduction of DeepSeek-R1 Blog GitHub Stars
- DeepScaleR DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL Paper -

Static Corpus (Agent)

Date Name Title Paper Github
2025-08 ASearcher Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL Paper -
2025-07 WebShaper WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization Paper -
2025-05 ZeroSearch ZeroSearch: Incentivize the Search Capability of LLMs without Searching Paper GitHub Stars
2025-04 ToolRL ToolRL: Reward is All Tool Learning Needs Paper GitHub Stars
2025-03 Search-R1 Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning Paper GitHub Stars
2025-03 ToRL ToRL: Scaling Tool-Integrated RL Paper GitHub Stars
- MicroThinker MiroVerse V0.1: A Reproducible, Full-Trajectory, Ever-Growing Deep Research Dataset Paper -
2025-03 DeepRetrieval DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning Paper GitHub Stars

Static Corpus (Mix)

Date Name Title Paper Github
2025-08 Graph-R1 Graph-R1: Unleashing LLM Reasoning with NP-Hard Graph Problem Paper -
2025-06 RewardAnything RewardAnything: Generalizable Principle-Following Reward Models Paper Blog
2025-06 guru-RL-92k Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective Paper -
2025-05 Llama-Nemotron-PT Llama-Nemotron: Efficient Reasoning Models Paper -
2025-05 SkyWork OR1 Skywork Open Reasoner 1 Technical Report Paper GitHub Stars
2025-03 OpenVLThinker OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Paper GitHub Stars
- AM-DS-R1-0528-Distilled AM-DeepSeek-R1-0528-Distilled Paper GitHub Stars
- dolphin-r1 Dolphin R1 Dataset Paper -
- SYNTHETIC-1/2 SYNTHETIC-1 Release: Two Million Collaboratively Generated Reasoning Traces from Deepseek-R1 Blog -

Dynamic Environment (Rule-based)

Date Name Title Paper Github
2025-06 ProtoReasoning ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs Paper -
2025-05 SynLogic SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond Paper GitHub Stars
2025-05 Reasoning Gym REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards Paper GitHub Stars
2025-05 Enigmata Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles Paper GitHub Stars
2025-02 AutoLogi AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models Paper GitHub Stars
2025-02 Logic-RL Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning Paper GitHub Stars

Dynamic Environment (Code-based)

Date Name Title Paper Github
2025-06 AgentCPM-GUI AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning Paper GitHub Stars
2025-06 MedAgentGym MedAgentGym: Training LLM Agents for Code-Based Medical Reasoning at Scale Paper GitHub Stars
2025-05 MLE-Dojo MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering Paper GitHub Stars
2025-05 SWE-rebench SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents Paper -
2025-05 ZeroGUI ZeroGUI: Automating Online GUI Learning at Zero Human Cost Paper GitHub Stars
2025-04 R2E-Gym R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents Paper GitHub Stars
2025-03 ReSearch ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning Paper GitHub Stars
2025-02 MLGym MLGym: A New Framework and Benchmark for Advancing AI Research Agents Paper GitHub Stars
2024-07 AppWorld AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents Paper [![GitHub Stars](https://img.shields.io/githu