← Open Source
thinkwee

AgentsMeetRL

Awesome List for Agentic RL

ListsOtherHTML
Open on GitHub
Momentum
+0stars in 24 hours0.0%
1.85k
Stars
75
Forks
+4
This week
9
Contributors
Created 2025-06-09 · Updated 2026-10-06 · #8651 today
Top developers
README

Logo

Base Framework General Search & RAG Web & GUI

Tool Code & SWE Reasoning Multi-Agent

Memory Embodied Domain-Specific Reward & Training

Safety VLM Agent Self-Evolution Environment

Interactive Dashboard

When LLM Agents Meet Reinforcement Learning

AgentsMeetRL is an awesome list that summarizes open-source repositories for training LLM Agents using reinforcement learning:

  • 🤖 The criteria for identifying an agent project are that it must have at least one of the following: multi-turn interactions or tool use (so TIR projects, Tool-Integrated Reasoning, are considered in this repo).
  • ⚠️ This project is based on code analysis from open-source repositories using LLM coding agents, which may contain unfaithful cases. Although manually reviewed, there may still be omissions. If you find any errors, please don't hesitate to let us know immediately through issues or PRs - we warmly welcome them!
  • 🚀 We particularly focus on the reinforcement learning frameworks, RL algorithms, rewards, and environments that projects depend on, for everyone's reference on how these excellent open-source projects make their technical choices. See [Click to view technical details] under each table.
  • 📅 Last updated: 2026-08-26
  • 🤗 Feel free to submit your own projects anytime - we welcome contributions!
  • 📚 If you find this repository helpful for your research, please cite it via the "Cite this repository" button on the right sidebar.

Taxonomy:

  • Base Framework: General-purpose RL training frameworks for LLM agents (e.g., veRL, OpenRLHF, trl)
  • General/MultiTask: Agent systems trained/evaluated across multiple tasks or environments
  • Search & RAG: Search-augmented reasoning agents that use retrieval tools to enhance LLM reasoning
  • Web & GUI: Agents that interact with web browsers, mobile/desktop GUIs, or operating systems
  • Tool-Use: Agents trained to invoke external tools (APIs, code executors, MCP, etc.)
  • Code & SWE: Software engineering and code generation agents
  • Reasoning: Reasoning agents with tool-integrated or multi-turn reasoning (math, QA, visual)
  • Multi-Agent RL: Multi-agent collaboration, negotiation, or credit assignment via RL
  • Memory: Agents that learn to manage, retrieve, or evolve memory
  • Embodied: Agents operating in embodied/physical simulation environments
  • Domain-Specific: RL agents for specialized domains (medical, OS tuning, etc.)
  • Reward & Training: Process/outcome reward models and training methodologies for agents
  • Safety: RL for agent safety alignment, adversarial red-teaming, and jailbreak defense/attack
  • VLM Agent: Vision-language model agents trained with RL for multimodal interaction
  • Self-Evolution: Agents that self-evolve via RL feedback loops (⚠️ definition still evolving in the community)
  • Environment: Benchmarks, gyms, and sandbox environments for agent training/evaluation

Some Enumeration:

  • Enumeration for Reward Type:
    • External Verifier: e.g., a compiler or math solver
    • Rule-Based: e.g., a LaTeX parser with exact match scoring
    • Model-Based: e.g., a trained verifier LLM or reward LLM
    • Custom

Updates

  • 📢 2026-08 Update: Added 23 new repositories across 9 categories (Environment +8 [Echoverse/PAST-Bench/PatientAgentBench/LegalWorld/Evo-Bench/DocOps/ScrambleToolBench/DigiWorld], Search & RAG +5 [EviSD/GTA-RAG/LAPO, plus catch-up of GrepSeek/PyRAG], Base Framework +3 [Molt, AReno, and catch-up of Microsoft Orchard], Self-Evolution +2 [AgentOPSD/BaT], Code & SWE +1 [Lego-RL, harness-native RL inside Claude Code/OpenHands/OpenCode], Reward & Training +1 [Agent-G²], VLM Agent +1 [InSight-doc], Memory +1 [MemPrism], Tool-Use +1 [MUA-RL, promoted from Under Review — its code had in fact been public since 2025.8]). Every repo was opened and confirmed to contain real RL-training (or executable-environment) code. Papers whose code is still unreleased were left out (Qwen-UI-Agent, Qwen-CUA, UI-Mate, SearchMaster, RoMeRL, Agon, SINKFLEX-RL, GRASP, MAVEN, EviBack, ChemWorld) — see Under Review. Notably no qualifying new Safety, Embodied, or Multi-Agent RL repos appeared this window: that crop was uniformly SFT-only, inference-only, or code-withheld.
  • 📢 2026-07 Update: Added 13 new repositories from late-Jun–Jul 2026 across 8 categories (Self-Evolution +3 [SEED/OPID/UCOB, the on-policy-distillation-for-agentic-RL line], VLM Agent +2 [VTS/VSeek, long-video search agents], Tool-Use +2 [Tool-RL-Box; plus catch-up of AWorld-RL], Environment +2 [SETA terminal envs, OpenAgent tool-generalization sandbox], Web & GUI +1 [SCALE-CUA], Embodied +1 [REAL], Memory +1 [Supersede], Domain-Specific +1 [FaithMed]). Every entry was verified by opening the repo and confirming real RL-training (or environment) code — papers whose code is not yet released (EvoCUA-1.5, DeepSearch-World, CompactionRL, GUICrafter, VideoSearcher, Xiaomi-GUI-0) were deliberately left out.
  • 📢 2026-06 Update: Added 43 new repositories across 11 categories (VLM Agent +8, Search & RAG +7, Environment +6, Reward & Training +4, Base Framework/Tool-Use/Self-Evolution/Embodied +3 each, Web & GUI/Code & SWE/Domain-Specific +2 each). New since the last update: Harness-1, FastContext, OpenWebRL, Polar, AgentJet, HarnessX, APPO, SPADER, DeepRubric, Embodied-R1.5, SIRI; plus catch-up of earlier-2026 misses (Vision-DeepResearch, ARM-Thinker, PyVision-RL, Gen-Searcher, DataMind, Tool-R0, Agent World Model, VisGym, Gym-Anything, ChemCraft, OpAgent, etc.).
  • 📢 2026-05 Update: Added 17 new repositories from Apr–May 2026 across 11 categories (notably General/MultiTask +4 [SkillZero/T²PO/SDAR/StraTA, mostly ZJU-REAL & related agentic RL methods], VLM Agent +3 [MTA-Agent/ParaVT/OpenSearch-VL, multimodal deep search & video tool use], Web & GUI +2 [ClawGUI/ToolCUA]). Moved CoEvolve to "Under Review" (code not yet released).
  • 📢 2026-04 Update: Added 67 new repositories covering Apr 2025 – Apr 2026 across nearly every category (notably VLM Agent +9, Search & RAG +10, Web & GUI +7, Tool-Use +7). Also reclassified SkyRL (→ General) and SPIRAL (→ Multi-Agent), and updated the VAGEN entry to its NeurIPS'25 upstream repo.
  • 📢 2026-03 Update: Restructured taxonomy from 12 to 16 categories (added Multi-Agent RL, Reward & Training, Safety, VLM Agent, Self-Evolution, Domain-Specific; merged GUI into Web & GUI; retired TextGame/Biomedical). Added ~70 new repositories covering Sep 2025 – Mar 2026, growing the total from ~134 to 205.

🤖 Use as a Claude Code Skill

Logo

This list is also packaged as a Claude Code Skill — agents-meet-rl — that turns the corpus into an on-demand assistant for agentic-RL training, evaluation, and experiment design: reward not moving, KL / entropy / length blow-ups, GRPO / PPO / DAPO knobs, retokenization drift, tool-call parse failures, long-horizon credit assignment, LLM-judge inconsistency, benchmark contamination, and framework / benchmark / algorithm selection — each answer anchored to specific papers and repos from this list. Backed by a machine-readable corpus of 405 projects (snapshot 2026-08-26). Once installed, Claude Code auto-invokes it whenever your question matches.

Install as a plugin (recommended):

/plugin marketplace add thinkwee/claude-plugins
/plugin install agents-meet-rl@thinkwee

Or install manually:

git clone https://github.com/thinkwee/AgentsMeetRL
cp -r AgentsMeetRL/skills/agents-meet-rl ~/.claude/skills/

Then just ask, e.g. "my GRPO search agent's reward is flat but eval keeps dropping" or "which RL framework should I pick for a multi-turn tool-use agent?" — the skill routes your symptom to fixes grounded in this corpus.

🔧 Base Framework

Github Repo 🌟 Stars Date Org Paper Link
Libra Stars 2026.8 NetX Lab Paper
Molt Stars 2026.7 NVIDIA (NeMo Labs) Paper
Orchard Stars 2026.7 Microsoft Paper
AgentJet Stars 2026.6 ModelScope (Alibaba) Paper
HarnessX Stars 2026.6 Darwin-Agent Paper
Dressage Stars 2026.6 Accio-Lab --
AReno Stars 2026.6 Ant Group (inclusionAI) --
Polar Stars 2026.5 NVIDIA (NeMo) Paper
uni-agent Stars 2026.4 verl-project --
VeRL-Omni Stars 2026.4 verl-project --
OpenClaw-RL Stars 2026.3 Gen-Verse Paper
Claw-R1 Stars 2026.3 USTC --
Open-AgentRL Stars 2026.2 Gen-Verse Paper
NeMo-RL Stars 2026.1 NVIDIA --
RLinf Stars 2025.8 Tsinghua/Infinigence AI/PKU Paper
siiRL Stars 2025.7 Shanghai Innovation Institute Paper
slime 2025.6 Tsinghua University (THUDM) blog
agent-lightning Stars 2025.6 Microsoft Research Paper
AReaL Stars 2025.6 AntGroup/Tsinghua Paper
ROLL Stars 2025.6 Alibaba Paper
MARTI Stars 2025.5 Tsinghua --
Tunix Stars 2025.4 Google --
RL2 Stars 2025.4 Accio –
verifiers Stars 2025.3 Individual --
prime-rl Stars 2025.2 Prime Intellect --
oat Stars 2024.11 NUS/Sea AI Paper
veRL Stars 2024.10 ByteDance Paper
OpenRLHF Stars 2023.7 OpenRLHF Paper
trl Stars 2019.11 HuggingFace --

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
Libra Async GRPO Single Outcome Multi Agentic RL post-training with resource-aware training and rollout All (Custom/External/Rule) Yes (search, code execution, validation tools)
Orchard Online RL (vendored slime) Single Both Multi Harness-native RL (browser/computer-use/SWE) Model/Rule Yes (live browser, containers, agent harnesses)
AReno GSPO/GRPO (+SFT/DPO) Single Outcome Multi Single-node RL post-training (Math/SWE-style/Games) Custom Yes (inspect_tree/read_file/rg/apply_patch/run_command)
Molt REINFORCE/RLOO/GRPO/Dr.GRPO/GAE + On-policy Distillation Single Outcome Multi Domain-agnostic agentic RL (Math/Geometry/Chat) Custom Yes (Python exec, HTTP, VLM tools)
AgentJet GRPO/PPO (swarm, multi-dim reward) Both Both Multi Swarm agentic RL (heterogeneous multi-agent, multi-task) All (Custom/External/Rule) Yes (tool calls, agent frameworks)
HarnessX GRPO/PPO (slime/verl recipes) Single Outcome Multi Composable agent-harness foundry (ALFWorld/GAIA/WebShop/SWE-bench) External + Custom Yes (harness orchestrates tools/memory)
Dressage GRPO Both Outcome Multi Agentic RL for any agent and sandbox (SWE-Gym/ALFWorld/HotpotQA) External/Rule Yes (whitebox: code/shell/file/retrieval; blackbox: opencode/openclaw/claude_code/codex)
Polar GRPO Both Outcome Multi Agentic RL on any harness (SWE-Bench/SWE-Gym) External Verifier Yes (real agent harnesses: shell/Codex/Claude Code)
uni-agent GRPO/GSPO (partial rollout, fully-async) Single Outcome Multi SWE-Bench/Search/General Agent (1000+ concurrent) All Yes (unified model/tool/env abstractions)
VeRL-Omni FlowGRPO/DanceGRPO/Diffusion DPO Single Outcome Single Multimodal generation RL (image/video/omni) Model/External No
OpenClaw-RL GRPO/OPD Both Both Multi Terminal/GUI/SWE/Tool-call Model/External Yes
Claw-R1 Generic RL Framework Multi Both Multi General Agent All Yes (Framework-agnostic)
Open-AgentRL GRPO-TCR Single Both Multi Reasoning/GUI/Coding Model (PRM) Yes (SandboxFusion)
NeMo-RL GRPO/DAPO/GDPO/DPO Single Outcome Multi Math/Reasoning/Code Rule/External No
RLinf PPO/GRPO/DAPO/SAC/REINFORCE++/CrossQ/RLPD Both Both Multi Robotics/Math/Code/QA/VQA All (Rule/Model/External) Yes
siiRL PPO/GRPO/CPGD/MARFT Multi Both Multi LLM/VLM/LLM-MAS PostTraining Model/Rule Planned
slime GRPO/GSPO/REINFORCE++ Single Both Both Math/Code External Verifier Yes
agent-lightning PPO/Custom/Automatic Prompt Optimization Multi Outcome Multi Calculator/SQL Model/External/Rule Yes
AReaL PPO Both Outcome Both Math/Code External Yes
ROLL PPO/GRPO/Reinforce++/TOPR/RAFT++ Multi Both Multi Math/QA/Code/Alignment All Yes
MARTI PPO/GRPO/REINFORCE++/TTRL Multi Both Multi Math All Yes
Tunix PPO/GRPO/GSPO-Token/DAPO/Dr.GRPO Single Outcome Multi Math/Code/Game Rule/External Yes
RL2 Dr. GRPO/PPO/DPO Single Both Both QA/Dialogue Rule/Model/External Yes
verifiers GRPO Multi Outcome Both Reasoning/Math/Code All Code
prime-rl GRPO/PPO Multi Outcome Multi Math/Code/Search Model/External Yes
oat PPO/GRPO Single Outcome Multi Math/Alignment External No
veRL PPO/GRPO Single Outcome Both Math/QA/Reasoning/Search All Yes
OpenRLHF PPO/REINFORCE++/GRPO/DPO/IPO/KTO/RLOO Multi Both Both Dialogue/Chat/Completion Rule/Model/External Yes
trl PPO/GRPO/DPO Single Both Single QA Custom No

💪 General/MultiTask

Github Repo 🌟 Stars Date Org Paper Link RL Framework
T2PO Stars 2026.5 Academic (ICML 2026 Spotlight) Paper veRL
StraTA Stars 2026.5 Shanghai AI Lab / Oxford / Multi-institution Paper rLLM
SDAR Stars 2026.5 Zhejiang University (ZJU-REAL) Paper veRL (GiGPO-based)
SkillZero Stars 2026.4 Zhejiang University (ZJU-REAL) Paper veRL (GiGPO-based)
MetaClaw Stars 2026.3 UNC-Chapel Hill (AIMING Lab) Paper Custom
SkillRL Stars 2026.2 UNC-Chapel Hill (AIMING Lab) Paper Custom
LLM-in-Sandbox Stars 2026.1 RUC/MSRA/THU Paper rllm (w/ veRL)
youtu-agent Stars 2025.12 Tencent Youtu Lab Paper Custom
DEPO Stars 2025.11 HKUST/SJTU Paper LLaMA-Factory
SPEAR Stars 2025.10 Tencent Youtu Lab Paper veRL/verl-agent
DeepAgent Stars 2025.10 RUC/Xiaohongshu Paper Custom
AgentRL Stars 2025.9 Tsinghua Paper veRL
AgentGym-RL Stars 2025.9 Fudan University Paper veRL
Agent_Foundation_Models Stars 2025.8 OPPO Personal AI Lab Paper veRL
Trinity-RFT Stars 2025.5 Alibaba Paper veRL
SPA-RL-Agent Stars 2025.5 PolyU Paper TRL
verl-agent Stars 2025.5 NTU/Skywork Paper veRL
SkyRL Stars 2025.4 UC Berkeley / NovaSky-AI Paper Self (skyrl-train)
VAGEN Stars 2025.3 Northwestern University (mll-lab-nu) Paper veRL
ART Stars 2025.3 OpenPipe Paper TRL
OpenManus-RL Stars 2025.3 UIUC/MetaGPT -- Custom

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
T2PO T²PO (token+turn uncertainty-guided) Single Both Multi WebShop/ALFWorld/SearchQA/Embody/Game Rule Yes (search, web, embodied)
StraTA Hierarchical GRPO + Strategic Trajectory Abstraction Single Outcome Multi ALFWorld (93.1%)/WebShop (84.2%)/SciWorld (63.5%) Rule + Model (self-judge) Yes (interactive long-horizon envs)
SDAR Self-Distilled Agentic RL (GRPO + gated OPSD) Single Outcome Multi ALFWorld/WebShop/Search-QA Rule Yes (interactive envs)
SkillZero In-Context Agentic RL (GRPO + skill-context curriculum withdrawal) Single Outcome Multi ALFWorld/WebShop/Search-QA Rule Yes (interactive envs + skill library)
MetaClaw GRPO (LoRA) Single Process Multi General Agentic Model (PRM) Yes (Skill-augmented)
SkillRL GRPO Single Outcome Multi ALFWorld/WebShop/Search Rule Yes (Web search, actions)
LLM-in-Sandbox GRPO++ Single Outcome Multi Math/Physics/Chemistry/Biomedicine/Long-context/IF/SWE Rule Yes (Code Sandbox w/ Terminal, File, Internet)
youtu-agent Training-Free GRPO Single Outcome Multi Deep Research/Data Analysis/Tool-use Model/External Yes (Web search, code, file)
DEPO KTO + Efficiency Loss Single Both Multi Agent (BabyAI/WebShop) Rule Yes
SPEAR GRPO/GiGPO + SIL Single Both Multi Math/Agent Rule/External Yes (Search, Sandbox, Browser)
DeepAgent ToolPO Single Outcome Multi ToolBench/ALFWorld/WebShop/GAIA/HLE Model Yes (16,000+ RapidAPIs)
AgentRL GRPO/REINFORCE++/RLOO/ReMax/GAE Single Outcome Multi Agent Tasks External Yes
AgentGym-RL PPO/GRPO/RLOO/REINFORCE++ Single Outcome Multi Web/Search/Game/Embodied/Science Rule/Model/External Yes (Web, Search, Env APIs)
Agent_Foundation_Models DAPO/PPO Single Outcome Single QA/Code/Math Rule/External Yes
Trinity-RFT PPO/GRPO Single Outcome Both Math/TextGame/Web All Yes
SPA-RL-Agent PPO Single Process Multi Navigation/Web/TextGame Model No
verl-agent PPO/GRPO/GiGPO/DAPO/RLOO/REINFORCE++ Multi Both Multi Phone Use/Math/Code/Web/TextGame All Yes
SkyRL GRPO/PPO Single Both Multi Long-horizon Agents (SWE-Bench/Search/Math/SQL) Rule/External/Custom Yes
VAGEN PPO/GRPO (World Modeling RL) Single Both Multi Navigation/TextGame/Multimodal All Yes
ART GRPO Multi Both Multi TextGame All Yes
OpenManus-RL PPO/DPO/GRPO Multi Outcome Multi TextGame All Yes

🔍 Search & RAG Agent

Github Repo 🌟 Stars Date Org Paper Link RL Framework
EviSD Stars 2026.8 Academic Paper veRL
GTA-RAG Stars 2026.8 Academic (EMNLP'26 Findings) Paper veRL
LAPO Stars 2026.7 Academic Paper veRL
Harness-1 Stars 2026.6 UIUC Paper Custom
SlimSearcher Stars 2026.6 Ant Group / ZJU Paper Custom (agentic RL)
DeepRubric Stars 2026.6 Shandong University Paper verl-tool
SAAS Stars 2026.5 Xiamen University Paper slime
CuSearch Stars 2026.5 Academic Paper Custom
GrepSeek Stars 2026.5 UMass Amherst (CIIR) Paper veRL
PyRAG Stars 2026.5 Academic Paper veRL
ORBIT Stars 2026.4 University of Waterloo Paper Custom
LiteResearcher Stars 2026.4 Simplex AI / ZJU / PolyU Paper Custom
DR-Venus Stars 2026.4 Ant Group (inclusionAI) Paper veRL (IGPO-based)
MR-Search Stars 2026.3 Academic Paper Custom
ProRAG Stars 2026.1 RUC Paper Custom
O-Researcher Stars 2026.1 OPPO PersonalAI Lab Paper Custom
Agentic-RAG-R1 Stars 2025.12 PKU -- Custom
MemSearcher Stars 2025.11 CAS Paper Custom
DR Tulu Stars 2025.11 AI2 / UW / CMU / MIT Paper Open-Instruct
IGPO Stars 2025.10 Ant Group Paper (ICLR 2026) veRL
ReSeek Stars 2025.10 Tencent PCG BAC/Tsinghua University Paper veRL
AutoGraph-R1 Stars 2025.10 HKUST KnowComp Paper Custom
WebSeer Stars 2025.10 Individual Paper veRL
HiPRAG Stars 2025.10 Individual Paper veRL
Tree-GRPO Stars 2025.9 AMAP Paper veRL
DeepResearch Stars 2025.9 Alibaba/Tongyi Lab Paper Custom
DeepDive Stars 2025.9 Tsinghua/THUDM Paper Custom
ASearcher Stars 2025.8 Ant Research RL Lab
Tsinghua University & UW Paper RealHF/AReaL
SSRL Stars 2025.8 Tsinghua Paper Custom
Research-Venus Stars 2025.8 Ant Group Paper Custom
Graph-R1 Stars 2025.7 BUPT/NTU/NUS Paper veRL
Kimi-Researcher Stars 2025.6 Moonshot AI blog Custom
R-Search Stars 2025.6 Individual -- veRL
R1-Searcher-plus Stars 2025.5 RUC Paper Custom
StepSearch Stars 2025.5 SenseTime Paper veRL
AutoRefine Stars 2025.5 USTC Paper veRL
ZeroSearch Stars 2025.5 Alibaba Paper veRL
ReasonRAG Stars 2025.5 CityU HK / Huawei Paper Custom
VRAG Stars 2025.5 USTC / Tongyi Lab, Alibaba Paper veRL
MaskSearch Stars 2025.5 Tongyi Lab, Alibaba Paper DAPO / veRL
R3-RAG Stars 2025.5 Fudan NLP Paper OpenRLHF
O2-Searcher Stars 2025.5 KnowledgeXLab Paper veRL
s3 Stars 2025.5 UIUC Paper veRL
knowledge-r1 Stars 2025.5 CAS / UCAS Paper veRL
WebThinker Stars 2025.4 RUC Paper Custom
DeepResearcher Stars 2025.4 SJTU Paper veRL
Search-R1 Stars 2025.3 UIUC/Google paper1, paper2 veRL
R1-Searcher Stars 2025.3 RUC Paper OpenRLHF
C-3PO Stars 2025.2 Alibaba Paper OpenRLHF
DeepRetrieval Stars 2025.2 UIUC Paper veRL

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
PyRAG Curriculum shared-parameter GRPO (LoRA) Multi (Decompose/Plan/Answer) Both Multi Multi-hop RAG via executable Python synthesis Rule-Based + Execution Yes (Python exec + E5 retriever)
GrepSeek SFT cold-start + GRPO Single Outcome Multi Direct corpus interaction (shell/grep, no index) Rule-Based (token-F1 x format gate) Yes (shell pipelines over raw corpus)
LAPO GRPO + Leave-One-Turn Attribution Single Both Multi Multi-turn search QA (NQ/TriviaQA/HotpotQA/2Wiki) Rule-Based (self-generated) Yes (retrieval)
GTA-RAG 3-stage GRPO Single Both Multi Multi-hop QA over entity-document graph Rule-Based (support-doc + EM) Yes (graph + dense retrieval)
EviSD GRPO + Evidence-Conditioned Self-Distillation Single Outcome Multi Search-augmented multi-hop QA Rule-Based Yes (search/retrieval)
Harness-1 GRPO Single Outcome Multi Long-horizon search (web/finance/patents) w/ state-externalizing harness External + Rule Yes (search/retrieval/rerank)
SlimSearcher GRPO + Adaptive Reward Gating Single Outcome Multi Efficiency-aware deep research (GAIA/BrowseComp/xBench) Custom + Rule Yes (web search, browse)
DeepRubric GRPO + rubric rewards Single Process Multi Deep research report synthesis (evidence-tree rubric) Model + Rule (rubric) Yes (search/browse/scholar)
SAAS RL w/ boundary-aware reward (2-stage curriculum) Single Outcome Multi Self-aware agentic search (over-search mitigation, 7 QA sets) Rule-Based Yes (search)
CuSearch GRPO + Search-Depth curriculum rollout Single Outcome Multi Agentic RAG multi-hop QA Rule-Based (EM) Yes (retrieval/search)
ORBIT GRPO Single Outcome Multi Verifiable data-gen + RL for web search (Qwen3-4B) External + Rule Yes (web search)
LiteResearcher Scalable Agentic RL (curriculum w/ lite virtual world) Single Outcome Multi Deep Research (GAIA 71.3% / Xbench-DS 78.0%, 4B SOTA) Rule/External Yes (local search/browse env, Milvus+PostgreSQL)
DR-Venus GRPO + IGPO (info-gain turn-level) w/ agentic SFT Single Both Multi Edge-scale Deep Research (4B) Intrinsic (info-gain) + Rule (format) Yes (Search/Browse)
MR-Search In-context Meta-RL (multi-episode credit) Single Outcome Multi Agentic search w/ self-reflection Rule-Based Yes (search)
ProRAG GRPO + DGA (dual-granularity advantage) Single Both Multi Multi-hop RAG Model (PRM via MCTS) Yes (Retrieval)
O-Researcher GRPO + RLAIF Multi Process Multi Deep Research (Zhihu-KOL/WideSearch/ELI5) Model (LLM-as-Judge) Yes (Search/Crawl)
Agentic-RAG-R1 GRPO Single Outcome Multi Knowledge-intensive QA Rule/Model Yes (Wiki/Doc search)
MemSearcher Multi-context GRPO Single Outcome Multi Search/QA + Memory Rule/Model Yes (Web search + Memory)
DR Tulu GRPO + evolving rubrics Single Outcome Multi Long-form Deep Research Model (rubrics) Yes (Search/MCP)
IGPO GRPO + IGPO (Information Gain turn-level reward) Single Both Multi Multi-turn Search Agent (BrowseComp/-ZH) Intrinsic (belief Δ) + Outcome Yes (Search)
ReSeek GRPO/PPO Single Both Multi QA/Search Rule Search/JUDGE
AutoGraph-R1 GRPO (via VeRL) Single Outcome Multi KG Construction for QA Rule Yes (Graph retrieval)
WebSeer GRPO-style Single Outcome Multi Web Search QA (w/ self-reflection) Rule/Model Yes (Search)
HiPRAG PPO Single Process Multi Efficient Agentic RAG Model/Rule Yes (Retrieval)
Tree-GRPO GRPO/Tree-GRPO Single Outcome Multi Search Rule Search
DeepResearch RL-based Single Outcome Multi Deep Research Model Yes (Search, Browse)
DeepDive GRPO Single Outcome Multi KG-augmented Search Rule Yes (KG + Search)
ASearcher PPO/GRPO + Decoupled PPO Single Outcome Multi Math/Code/SearchQA External/Rule Yes
SSRL GRPO Single Outcome Multi Self-Search Rule Yes (Self-search)
Research-Venus GRPO Single Both Multi Deep Research Model (atomic thought) Yes (Search)
Graph-R1 GRPO/REINFORCE++/PPO Single Outcome Multi KGQA Rule (EM/F1) Yes (Graph retrieval)
Kimi-Researcher REINFORCE Single Outcome Multi Research Outcome Search, Browse, Coding
R-Search PPO/GRPO Single Both Multi QA/Search All Yes
R1-Searcher-plus Custom Single Outcome Multi Search Model Search
StepSearch PPO Single Process Multi QA Model Search
AutoRefine PPO/GRPO Multi Both Multi RAG QA Rule Search
ZeroSearch PPO/GRPO/REINFORCE Single Outcome Multi QA/Search Rule Yes
ReasonRAG DPO + MCTS-based PRM Single Process Multi Multi-hop QA Model (PRM) Yes (Wikipedia search)
VRAG GRPO Single Both Multi Visually-rich RAG Rule/Model Yes (Visual retrieval)
MaskSearch DAPO Single Outcome Multi RAMP Pretraining + QA Rule/Model Yes (Search)
R3-RAG PPO Single Both Multi Multi-hop QA Rule Yes (Retrieval)
O2-Searcher GRPO Single Outcome Multi Open-ended QA Rule/Model Yes (Search)
s3 GRPO Single Outcome Multi RAG / Medical QA Model (Gain-Beyond-RAG) Yes (Retrieval)
knowledge-r1 GRPO Single Outcome Multi Knowledge-intensive QA (KB-aware) Rule Yes (Retrieval)
WebThinker DPO Single Outcome Multi Reasoning/QA/Research Model/External Web Browsing
DeepResearcher PPO/GRPO Multi Outcome Multi Research All Yes
Search-R1 PPO/GRPO Single Outcome Multi Search All Search
R1-Searcher PPO/DPO Single Both Multi Search All Yes
C-3PO PPO Multi Outcome Multi Search Model Yes
DeepRetrieval GRPO Single Outcome Multi Query Generation/IR Rule Yes (Search)

🌐 Web & GUI Agent

Github Repo 🌟 Stars Date Org Paper Link RL Framework
SCALE-CUA Stars 2026.7 Tsinghua (THUDM) Paper Custom (Ray + vLLM + Megatron-LM)
OpenWebRL Stars 2026.6 UIUC / Microsoft Research Paper slime
ToolCUA Stars 2026.5 Alibaba Tongyi Lab (X-PLUG) Paper Custom
ClawGUI Stars 2026.4 Zhejiang University (ZJU-REAL) Paper Custom (veRL-based)
OpAgent Stars 2026.2 Codefuse AI (Ant Group) Paper Agent-R1 (veRL)
GUI-Libra Stars 2026.2 GUI-Libra (MS-affiliated) Paper Custom
MobileAgent Stars 2025.9 X-PLUG (TongyiQwen) paper veRL
UI-TARS Stars 2025.9 ByteDance Seed Paper Custom
MobileRL Stars 2025.9 Tsinghua / Zhipu AI (THUDM) Paper Custom
DART-GUI Stars 2025.9 Computer-use-agents Paper veRL
Mano-P Stars 2025.9 Mininglamp AI Paper Mano-SDK
InfiGUI-G1 Stars 2025.8 InfiX AI Paper veRL
gui-rcpo Stars 2025.8 Zhejiang University Paper Custom
UI-AGILE Stars 2025.7 Xiamen University Paper Custom
GUI-G2 Stars 2025.7 Zhejiang University (ZJU-REAL) Paper Custom (VLM-R1)
MagicGUI Stars 2025.7 Honor (MagicAgent-GUI) Paper Custom
Grounding-R1 Stars 2025.6 Salesforce blog trl
AgentCPM-GUI Stars 2025.6 OpenBMB/Tsinghua/RUC Paper Huggingface
TTI Stars 2025.6 CMU Paper Custom
GTA1 Stars 2025.6 Salesforce / ANU Paper Custom (DeepSpeed)
SE-GUI Stars 2025.5 Nankai University/vivo Paper trl
ARPO Stars 2025.5 CUHK/HKUST Paper veRL
GUI-G1 Stars 2025.5 RUC Paper TRL
WebAgent-R1 Stars 2025.5 Amazon/UVA Paper Custom
ZeroGUI Stars 2025.5 Shanghai AI Lab Paper Custom
GUI-R1 Stars 2025.4 CAS/NUS Paper veRL
InfiGUI-R1 Stars 2025.4 Zhejiang University Paper Custom
UI-R1 Stars 2025.3 vivo/CUHK Paper TRL
CollabUIAgents Stars 2025.2 Tsinghua/Alibaba/HKUST Paper Custom
DigiQ Stars 2025.2 UC Berkeley/CMU/Amazon Paper Custom
GUI-Agent-RL Stars 2025.2 Microsoft Paper Custom
WebAgent Stars 2025.1 Alibaba paper1, paper2 LLaMA-Factory

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
SCALE-CUA GRPO (fully async, Frontier Sampling + visual context segmentation) Single Outcome Multi Computer Use (OSWorld, ScienceBoard; 24K+ synthesized verifiable tasks) External Verifier (executable judge functions) Yes (GUI actions in Docker desktop)
OpenWebRL GRPO (online multi-turn) Single Both Multi Visual web browsing on live sites (WebVoyager/Online-Mind2Web) Rule + Model (format + LLM-judge) Yes (Playwright browser)
ToolCUA Tool-Bootstrapped GUI RFT + Online Agentic RL (Tool-Efficient Path Reward) Single Both Multi Computer Use (OSWorld-MCP, hybrid GUI+tool) Rule (path-efficiency) Yes (GUI actions + structured tool calls)
ClawGUI GiGPO + Process Reward Model Single Both Multi Mobile GUI (Android/HarmonyOS/iOS, MobileWorld) Rule + Model (PRM) Yes (GUI + hybrid CLI-GUI + persistent memory)
OpAgent Online agentic RL (GRPO/PPO) Multi Both Multi Web navigation (WebArena 71.6% pass@5) Rule + Model (RDTree + WebJudge) Yes (Playwright browser)
GUI-Libra KL-regularized GRPO (Partially Verifiable RL) Single Outcome Multi GUI (AndroidWorld/WebArena/Online-Mind2Web) Rule Yes
MobileAgent semi-online RL Single Both Multi MobileGUI/Automation Rule Yes
UI-TARS Multi-turn RL Single Both Multi GUI (Cross-platform) Model Yes (GUI actions)
MobileRL AdaGRPO (Difficulty-Adaptive) Single Outcome Multi Mobile GUI (AndroidWorld/AndroidLab) Rule Yes (Android)
DART-GUI Decoupled GRPO Single Outcome Multi GUI (OSWorld) Rule Yes
Mano-P Three-stage SFT→Offline RL→Online RL Single Both Multi GUI (OSWorld) Rule Yes
InfiGUI-G1 AEPO Single Outcome Single GUI/Grounding Rule No
gui-rcpo RCPO Single Outcome Single GUI Grounding Rule (self-supervised) No
UI-AGILE GRPO Single Outcome Single GUI Grounding Rule (continuous) No
GUI-G2 GRPO (Gaussian Reward) Single Outcome Single GUI Grounding Rule (continuous) No
MagicGUI Reinforcement Fine-Tuning (RFT) Single Outcome Multi Mobile GUI Model/Rule Yes
Grounding-R1 GRPO Single Outcome Multi GUI Grounding Model Yes
AgentCPM-GUI GRPO Single Outcome Multi Mobile GUI Model Yes
TTI REINFORCE/BC Single Outcome Multi Web External Web Browsing
GTA1 GRPO-style (click-success reward) Single Outcome Multi GUI Grounding (OSWorld/ScreenSpot-Pro) Rule Yes
SE-GUI GRPO Single Both Single GUI Grounding Rule Yes
ARPO GRPO Single Outcome Multi GUI External Computer Use
GUI-G1 GRPO Single Outcome Single GUI Rule/External No
WebAgent-R1 M-GRPO Single Outcome Multi Web Navigation (WebArena-Lite) Rule (task success) Yes (Web browsing)
ZeroGUI Online RL Single Outcome Multi GUI Agent Rule Yes (GUI actions)
GUI-R1 GRPO Single Outcome Multi GUI Rule No
InfiGUI-R1 RL + sub-goal guidance Single Both Multi GUI Reasoning Rule Yes
UI-R1 GRPO Single Process Both GUI Rule Computer/Phone Use
CollabUIAgents DPO (credit re-assignment) Multi Process Multi GUI (Mobile + Web) Model (LLM) Yes (GUI interaction)
DigiQ Value-based offline RL Single Outcome Multi Android Device Control Model (Q-function) Yes
GUI-Agent-RL Value-based RL (VEM) Single Outcome Multi GUI (Web Shopping) Model Yes
WebAgent DAPO Multi Process Multi Web Model Yes

🔨 Tool-Use Agent

Github Repo 🌟 Stars Date Org Paper Link RL Framework
Tool-RL-Box Stars 2026.6 Harbin Institute of Technology Paper veRL (w/ verl-tool)
SPADER Stars 2026.6 Zhejiang University Paper veRL
APPO Stars 2026.6 Alibaba AMAP (AMAP-ML) Paper veRL
AgenticQwen Stars 2026.4 Alibaba PAI Paper veRL (w/ EasyDistill)
Agent-STAR Stars 2026.3 CUHK Paper veRL
ToolOrchestra Stars 2025.11 NVIDIA / HKU Paper Custom (veRL-based)
ToolMaster Stars 2025.11 Northeastern University (NEUIR) Paper Custom
MATPO Stars 2025.10 MiroMind AI Paper Custom
AWorld-RL Stars 2025.10 Ant Group (inclusionAI) -- AWorld + veRL
CodeGym Stars 2025.9 Academic Paper Custom
UserRL Stars 2025.9 Salesforce AI Research Paper veRL
ToolBrain Stars 2025.9 ToolBrain (AAMAS 2026) Paper Custom
Tool-R1 Stars 2025.9 Individual (YBYBZhang) Paper Custom
MiroRL Stars 2025.8 MiroMindAI HF Repo veRL
MUA-RL Stars 2025.8 Alibaba (Tongyi) Paper veRL
verl-tool Stars 2025.6 TIGER-Lab X veRL
Multi-Turn-RL-Agent Stars 2025.5 University of Minnesota Paper Custom
Tool-N1 Stars 2025.5 NVIDIA Paper veRL
Tool-Star Stars 2025.5 RUC Paper LLaMA-Factory
RL-Factory Stars 2025.5 Simple-Efficient model veRL
calculator_agent_rl Stars 2025.5 Individual (Danau5tin) -- Verifiers
ReTool Stars 2025.4 ByteDance Paper veRL
ToolRL Stars 2025.4 UIUC Paper veRL
AWorld Stars 2025.3 Ant Group (inclusionAI) Paper veRL
Agent-R1 Stars 2025.3 USTC Paper veRL
ReCall Stars 2025.3 BaiChuan Paper veRL

📋 Click to view technical details

| Github Repo | RL Algorithm | Single/Multi Agent | Outcome/Process Reward | Single/Multi Turn | Task | Reward Type | Tool usage |

MUA-RL GRPO Single Outcome Multi Multi-turn user-interacting agentic tool use (tau-bench/tau2-bench/ACEBench) Rule-Based (task completion) Yes (simulated user + tool APIs)
Tool-RL-Box GRPO + supervisory signals (anti format-collapse) Single Process Multi Multi-step function calling (FCL / ToolACE, pluggable tool servers) Model (LLM-judge error taxonomy) + Rule Yes (function-calling tools)
SPADER GRPO + Step-wise Peer Advantage (SPA) Single Both Multi Long-horizon tool-augmented multi-answer QA (QAMPARI) Rule-Based (entity-match + diversity) Yes (search)
APPO APPO (procedure-aware branching; extends ARPO/GRPO) Single Process Multi Multi-turn TIR (reasoning+search+code, 13 benchmarks) Rule-Based Yes (search + code)
AgenticQwen Multi-round RL (Reasoning RL + Agentic RL w/ dual data flywheels) Single Outcome Multi Industrial Tool Use (search, data analysis, tau-bench airline/retail/telecom) Rule + Model (rubric) Yes (Python interpreter, web search, mock tools)
Agent-STAR GRPO + dense/curriculum reward (STAR recipe) Single Both Multi Long-horizon tool-using agents (TravelPlanner, ReAct up to 60 turns) Rule + External Yes (planning APIs)
ToolOrchestra End-to-end RL (outcome+efficiency+preference) Single Both Multi Tool orchestration / agentic workflows All Yes (Search/Code/LLMs)
ToolMaster SFT + GRPO (trial-then-execute) Single Outcome Multi Tool trialing + execution (ToolHop/TMDB/StableToolBench) Rule/External Yes (Simulated tools)
MATPO GRPO (multi-agent) Multi Outcome Multi Tool-use/Search Rule Yes (MCP: Serper, Web scraping)
AWorld-RL Collection: RODS / HardGen / FunReason-MT / Environment Tuning / V2P / RAG-R1 Both Both Multi Multi-turn function calling + GUI grounding + deep search (BFCL etc.) Rule + Model (progress reward) Yes (function calls, GUI, search)
CodeGym GRPO-family Single Outcome Multi Synthetic Multi-turn Tool-Use Rule (verifiable) Yes (Synthesized tools)
UserRL GRPO (multi-turn credit) Single Both Multi User-centric (Function/Persuade/Search/Tau Gyms) Model/External Yes
ToolBrain GRPO/DPO Single Outcome Multi Agentic tool training Rule/Model Yes (User-defined tools)
Tool-R1 Policy optimization (PPO-style) Single Outcome Multi Agentic Tool Use (GAIA) Model + External Yes (Python exec)
MiroRL GRPO Single Both Multi Reasoning/Planning/ToolUse Rule-based MCP
verl-tool PPO/GRPO Single Both Both Math/Code Rule/External Yes
Multi-Turn-RL-Agent GRPO Single Both Multi Tool-use/Math Rule/External Yes
Tool-N1 PPO Single Outcome Multi Math/Dialogue All Yes
Tool-Star PPO/DPO/ORPO/SimPO/KTO Single Outcome Multi Multi-modal/Tool Use/Dialogue Model/External Yes
RL-Factory GRPO Multi Both Multi Tool-use/NL2SQL All MCP
calculator_agent_rl GRPO Single Outcome Multi Calculator Tool Use Model (Claude-judge) Yes
ReTool PPO Single Outcome Multi Math External Code
ToolRL GRPO/PPO Single Outcome Multi Tool Learning Rule/External Yes
AWorld GRPO Both Outcome Multi Search/Web/Code External/Rule Yes
Agent-R1 PPO/GRPO Single Both Multi Tool-use/QA Model Yes
ReCall PPO/GRPO/RLOO/REINFORCE++/ReMax Single Outcome Multi Tool-use/Math/QA All Yes

💻 Code & SWE Agent

Github Repo 🌟 Stars Date Org Paper Link RL Framework
Lego-RL Stars 2026.8 LegoX Paper veRL
FastContext Stars 2026.6 Microsoft Paper Custom
SWE-Edit Stars 2026.4 Microsoft Research Paper Custom
CodeScout Stars 2026.3 OpenHands Paper SkyRL
CUDA-Agent Stars 2026.2 ByteDance/Tsinghua Paper Custom
SWE-World Stars 2026.2 RUC (RUCAIBox) Paper OpenRLHF + veRL
LLM-in-Sandbox Stars 2026.1 RUC/MSRA/THU Paper rllm (w/ veRL)
CUDA-L2 Stars 2026.1 DeepReinforce AI Paper Custom
PPP-Agent Stars 2025.11 CMU/OpenHands Paper veRL
DeepAnalyze Stars 2025.10 RUC/Tsinghua Paper Custom
RepoDeepSearch Stars 2025.8 PKU, Bytedance, BIT Paper veRL
CUDA-L1 Stars 2025.7 DeepReinforce AI Paper Custom
SWE-Swiss Stars 2025.7 Tsinghua / ByteDance -- veRL
MedAgentGym Stars 2025.6 Emory/Georgia Tech Paper Hugginface
CURE Stars 2025.6 University of Chicago
Princeton/ByteDance Paper Huggingface
Time-R1 Stars 2025.5 UIUC Paper veRL
ML-Agent Stars 2025.5 MASWorks Paper Custom
R1-Code-Interpreter Stars 2025.5 MIT Paper Custom
digitalhuman Stars 2025.4 Tencent Paper veRL
Skywork-OR1 Stars 2025.4 Skywork AI Paper Custom (veRL fork)
sweet_rl Stars 2025.3 Meta/UCB Paper OpenRLHF
swe-rl Stars 2025.2 Meta/UIUC/CMU Paper Custom
CTRL Stars 2025.2 HKU/ByteDance Paper Custom
AceCoder Stars 2025.2 Waterloo (TIGER-Lab) Paper Custom
rllm Stars 2025.1 Berkeley Sky Computing Lab
BAIR / Together AI Notion Blog veRL
open-r1 Stars 2025.1 HuggingFace -- TRL

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
Lego-RL PPO/GRPO/GSPO (sync & async) Single Outcome Multi SWE-bench Verified inside real coding harnesses External Verifier (Harbor sandbox) Yes (native harness bash/file-edit/test)
FastContext Task-grounded RL Single Outcome Multi Repo-explorer subagent (context gathering + citations) Rule-Based Yes (Read/Glob/Grep, parallel)
SWE-Edit GRPO (adaptive mode selection) Multi (Viewer + Editor subagents) Outcome Multi SWE-bench Verified (find-replace vs whole-file rewrite) Rule/External (test-based) Yes (bash, file ops, viewer subagent)
CodeScout GSPO Single Outcome Multi Repo-level code search/localization (terminal) Rule-Based (F1) Yes (terminal: rg/sed/cat)
CUDA-Agent Agentic RL (staged) Single Outcome Multi CUDA Kernel Generation Rule (correctness + performance) Yes (compile/verify/profile)
SWE-World RL with learned world model (SWT + SWR) Single Both Multi Docker-free SWE (SWE-Bench Verified) Model (surrogate) + Rule Yes
LLM-in-Sandbox GRPO++ Single Outcome Multi Code/SWE + General (Math/Sci/Bio) Rule Yes (Code Sandbox w/ Terminal, File, Internet)
CUDA-L2 Contrastive RL Single Outcome Single HGEMM / CUDA Matmul Rule (TFLOPs) Yes (compile/benchmark)
PPP-Agent PPP-RL Single Both Multi SWE/Research Rule+Model Search, Ask, Browse
DeepAnalyze Curriculum RL Single Outcome Multi Data Science Rule/External Yes (Code exec)
RepoDeepSearch GRPO Single Both Multi Search/Repair Rule/External Yes
CUDA-L1 Contrastive RL Single Outcome Single CUDA Optimization Rule (performance) No
SWE-Swiss Two-stage RL curriculum Single Outcome Multi SWE (Localization/Repair/Unit-Test) Rule (test-based) Yes
MedAgentGym SFT/DPO/PPO/GRPO Single Outcome Multi Medical/Code External Yes
CURE PPO Single Outcome Single Code External No
Time-R1 PPO/GRPO/DPO Multi Outcome Multi Temporal All Code
ML-Agent Custom Single Process Multi Code All Yes
R1-Code-Interpreter GRPO Single Outcome Multi Code Interpretation Rule/External Yes (Code exec)
digitalhuman PPO/GRPO/ReMax/RLOO Multi Outcome Multi Empathy/Math/Code/MultimodalQA Rule/Model/External Yes
Skywork-OR1 Large-scale rule-based RL (GRPO variant) Single Outcome Single Math + Code (AIME/LiveCodeBench) Rule (verifiable) No
sweet_rl DPO Multi Process Multi Design/Code Model Web Browsing
swe-rl RL-based Single Outcome Single SWE (SWE-bench) Rule (similarity) No
CTRL RL (critique-revision) Single Process Multi Code Refinement Model Yes (Code exec)
AceCoder GRPO Single Outcome Single Code Generation External (test cases) Yes
rllm PPO/GRPO Single Outcome Multi Code Edit External Yes
open-r1 GRPO Single Outcome Single Math/Code All Yes

🤔 Reasoning Agent

Github Repo 🌟 Stars Date Org Paper Link RL Framework
Agent0 Stars 2025.10 UNC‑Chapel Hill / Salesforce Research / Stanford University Paper veRL
KG-R1 Stars 2025.9 UIUC/Google Paper1, Paper2 veRL
AgentFlow Stars 2025.09 Stanford University arXiv veRL
THOR Stars 2025.9 USTC / iFLYTEK Paper veRL
Tool-Light Stars 2025.9 RUC (RUC-NLPIR) Paper LLaMA-Factory
ARPO Stars 2025.7 RUC, Kuaishou Paper veRL
terminal-bench-rl Stars 2025.7 Individual (Danau5tin) N/A rLLM
AutoTIR Stars 2025.7 Beihang University / BAAI Paper veRL
MOTIF Stars 2025.6 University of Maryland Paper trl
cmriat/l0 Stars 2025.6 CMRIAT Paper veRL
agent-distillation Stars 2025.5 KAIST Paper Custom
EasyR1 Stars 2025.4 Individual repo1/paper2 veRL
AutoCoA Stars 2025.3 BJTU Paper veRL
ToRL Stars 2025.3 SJTU Paper veRL
ReMA Stars 2025.3 SJTU, UCL Paper veRL
Agentic-Reasoning Stars 2025.2 Oxford Paper Custom
SimpleTIR Stars 2025.2 NTU, Bytedance Notion Blog veRL
openrlhf_async_pipline Stars 2024.5 OpenRLHF Paper OpenRLHF

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
Agent0 ADPO Multi Process Multi Math/Visual Model/Verifier Yes
KG-R1 GRPO/PPO Single Both Multi KGQA Rule/Model KG Retrieval
AgentFlow Flow-GRPO Single Outcome Multi Search/Math/QA Model/External Yes
THOR Hierarchical GRPO (trajectory+step) Single Both Multi Math (MATH500/AIME/Olympiad) External (SandboxFusion) Yes (Python)
Tool-Light Self-Evolved DPO Single Outcome Multi Tool-Integrated Reasoning Model (preference) Yes (FlashRAG/Python)
ARPO GRPO Single Outcome Multi Math/Coding Model/Rule Yes
terminal-bench-rl GRPO Single Outcome Multi Coding/Terminal Model+External Verifier Yes
AutoTIR PPO Single Outcome Multi Autonomous Tool Selection (QA/Math/IF) Rule Yes (Search/Python)
MOTIF GRPO Single Outcome Multi QA Rule No
cmriat/l0 PPO Multi Process Multi QA All Yes
agent-distillation PPO Single Process Multi QA/Math External Yes
EasyR1 GRPO Single Process Multi Vision-Language Model Yes
AutoCoA GRPO Multi Outcome Multi Reasoning/Math/QA All Yes
ToRL GRPO Single Outcome Single Math Rule/External Yes
ReMA PPO Multi Outcome Multi Math Rule No
Agentic-Reasoning Custom Single Process Multi QA/Math External Web Browsing
SimpleTIR PPO/GRPO (with extensions) Single Outcome Multi Math, Coding All Yes
openrlhf_async_pipline PPO/REINFORCE++/DPO/RLOO Single Outcome Multi Dialogue/Reasoning/QA All No

👥 Multi-Agent RL

Github Repo 🌟 Stars Date Org Paper Link RL Framework
Maestro Stars 2026.5 Tsinghua / Multi-institution Paper veRL + verl-tool
DrMAS Stars 2026.2 NTU Paper Custom
MarsRL Stars 2025.11 Academic Paper veRL
PettingLLMs Stars 2025.10 Intel / UCSD Paper Custom
MASPRM Stars 2025.10 UBC / Huawei Paper Custom
MrlX Stars 2025.10 Ant Group (AQ-MedAI) Paper Custom (SGLang + Megatron)
CoMAS Stars 2025.10 Shanghai AI Lab / CUHK / Oxford / NUS Paper Custom
MAPoRL Stars 2025.8 Academic -- Custom
CoMLRL Stars 2025.8 OpenMLRL Paper TRL
ARIA Stars 2025.6 Fudan University Paper Custom
SPIRAL Stars 2025.6 NUS / A*STAR / Sea AI Lab Paper Oat
AMPO Stars 2025.5 Tongyi Lab, Alibaba Paper veRL
FlowReasoner Stars 2025.4 Sea AI Lab / NUS Paper Custom
MARFT Stars 2025.4 SII / SJTU Paper Custom

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
Maestro Outcome RL (lightweight orchestrator over frozen expert ensembles) Multi (orchestrator + frozen experts) Outcome Multi 10 multimodal benchmarks (math/chart/HR/domain — 70.1% avg, beats GPT-5 & Gemini-2.5-Pro) External Yes (expert models + 2-tier skill library: OCR/detection/visual)
DrMAS GRPO (agent-wise) Multi Outcome Multi Multi-agent LLM Systems Rule No
MarsRL RLVR (agent-specific rewards) Multi Both Multi Math Reasoning (AIME/BeyondAIME) Rule (verifiable) No
PettingLLMs AT-GRPO Multi Both Multi Game/Code/Math/Planning Rule (verifiable) No
MASPRM PRM (trained from MCTS rollouts) Multi Process Multi Reasoning (GSM8K/MATH/MMLU) Learned PRM No
MrlX M-GRPO (hierarchical) Multi Outcome Multi Deep Research (GAIA/XBench) Rule + Model Yes (Search)
CoMAS RL w/ LLM-Judge intrinsic reward Multi Process Multi Co-evolving Reasoning Model No
MAPoRL PPO Multi Outcome Multi Collaborative LLM Tasks Rule No
CoMLRL MAGRPO / MAREINFORCE / MARLOO Multi Outcome Multi Writing / Code / Minecraft Custom Minimal
ARIA REINFORCE Both Process Multi Negotiation/Bargaining Other No
SPIRAL Role-conditioned Advantage Estimation (RAE) Multi Outcome Multi Zero-sum Games (TicTacToe/Kuhn/Negotiation) Rule No
AMPO BC/AMPO(GRPO improvement) Multi Outcome Multi Social Interaction Model-based No
FlowReasoner GRPO Multi Outcome Multi Multi-agent Workflow Design Rule Yes
MARFT MARFT paradigm (action+token level) Multi Both Multi Research / Math Rule Yes

🧠 Memory

Github Repo 🌟 Stars Date Org Paper Link RL Framework
MemPrism Stars 2026.8 Academic Paper veRL
Supersede Stars 2026.6 Vrin Paper verifiers + prime-rl
AgeMem Stars 2026.4 Multi-institution (incl. Alibaba DAMO) Paper Trinity-RFT
Mem-alpha Stars 2025.9 UCSD / USTC Paper veRL
MEM1 Stars 2025.7 MIT Paper veRL (based on Search-R1)
M3-Agent Stars 2025.7 ByteDance Seed / Zhejiang University Paper Custom
Memento Stars 2025.6 UCL, Huawei Paper Custom
MemAgent Stars 2025.6 Bytedance, Tsinghua-SIA Paper veRL

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
MemPrism GRPO/GiGPO Single Both Multi Memory-view selection for ALFWorld/ALFRED + Mind2Web Rule-Based Yes (memory-view action + env actions)
Supersede GRPO (+ LoRA) Single Outcome Multi Memory-update gap: keeping notes current across sessions (LongMemEval knowledge-update) Rule-Based (answered_current / stale_penalty) Yes (capped notes memory as action space)
AgeMem Step-wise GRPO (3-stage progressive RL) Single Process Multi Unified LTM/STM management (memory ops as tools) Rule (task accuracy + memory quality) Yes (store/retrieve/update/summarize/discard memory tools)
Mem-alpha GRPO Single Outcome Multi Long-context QA + Memory Construction Rule (downstream QA) Yes (memory tools)
MEM1 PPO/GRPO Single Outcome Multi WebShop/GSM8K/QA Rule/Model Yes
M3-Agent RL-based Single Outcome Multi Long-video QA (M3-Bench) Rule/Model Yes (multimodal memory graph)
Memento soft Q-Learning Single Outcome Multi Research/QA/Code/Web External/Rule Yes
MemAgent PPO, GRPO, DPO Multi Outcome Multi Long-context QA Rule/Model/External Yes

🦾 Embodied

Github Repo 🌟 Stars Date Org Paper Link RL Framework
REAL Stars 2026.7 InternRobotics Paper Custom (GSPO/GRPO over MCP)
Embodied-R1.5 Stars 2026.6 Tianjin University Paper EasyR1 / veRL
AVA-VLA Stars 2026.6 UCAS Paper Custom (PPO)
WorldVLN Stars 2026.5 Tsinghua (EmbodiedCity) Paper Custom
Embodied-R1 Stars 2025.6 Tianjing University Paper veRL
VIKI-R Stars 2025.6 MARS-EAI (NeurIPS 2025 D&B) Paper veRL + LLaMA-Factory
STeCa Stars 2025.2 The Hong Kong Polytechnic University Paper FastChat/TRL

📋 Click to view technical details

Github Repo RL Algorithm Single/Multi Agent Outcome/Process Reward Single/Multi Turn Task Reward Type Tool usage
REAL GRPO/GSPO (online RL over an MCP tool interface) Single Outcome Multi Open-world mobile manipulation in Isaac Sim (REAL-Bench, 241 tasks) External Verifier (target world-state check) Yes (8 MCP tools: navigate_to/pick/place/ask/...)
Embodied-R1.5 RFT (GRPO-family multimodal) Single Outcome Multi Embodied foundation model w/ Planner-Grounder-Corrector closed-loop Rule-Based No (closed-loop PGC)
AVA-VLA PPO (latent reasoning as sequential decision) Single Both Multi VLA manipulation (LIBERO/ALOHA), latent CoT w/ early-exit External (task success) + Custom No (closed-loop manipulation)
WorldVLN Action-aware GRPO Single Both Multi Aerial (UAV) vision-language navigation (closed-loop) Rule + Model No (closed-loop UAV control)
Embodied-R1 GRPO Single Outcome Single Grounding/Waypoint Rule No
VIKI-R GRPO (RFT after SFT) Multi Outcome Multi Embodied Multi-Robot Cooperation (VIKI-Bench) Rule + Model No
STeCa DPO (RFT) Single Both Multi Embodied/Household Rule/MC Environment Actions

🏷️ Domain-Specific

| Github Repo | 🌟 Stars | Date | Org | Paper Link | RL Framework | Domain | | :----: | :----: | :----: | :----: | :----: | :----: |