← Open Source
verl-project

verl-omni

Multimodal RL training framework for diffusion & omni models

Model DevelopmentDomain-specific trainingReinforcement LearningPython
Open on GitHub
Momentum
+6stars in 24 hours+0.5%
1.17k
Stars
228
Forks
—
This week
39
Contributors
Created 2026-04-21 · Updated 2026-10-05 · #1383 today
Top developers
README

VeRL-Omni

Easy, fast, and stable RL training for diffusion and omni-modality models

Ask DeepWiki.com Docs License

VeRL-Omni is a general RL training framework focused on multimodal generative models, built on top of verl.

It originated from the multi-modal generation RL effort in verl, and now has a dedicated home so it can evolve in a more focused way.

News 🔥

  • [2026-09] The v1 trainer (TransferQueue + ReplayBuffer) is now the default for diffusion models; the legacy v0 trainer is deprecated, while offline DPO remains on v0 by design.
  • [2026-08] DiffusionOPD (on-policy distillation, including multi-teacher MOPD) is now supported.
  • [2026-08] 🔥 MiniMax-H3 now supports T2VA, FL2VA, and Ref2VA with both FlowGRPO and DiffusionNFT.
  • [2026-08] 🎉 We have released v0.2.0 for faster diffusion rl and more stable Qwen3-Omni multimodal training. Blog: VeRL-Omni v0.2.0
  • [2026-08] LTX2.3 text-to-video+audio model is now supported with FlowGRPO.
  • [2026-07] Team-proposed algorithm FlowGRPO with DiNa-LRM is available. Training skips VAE decoding by scoring clean diffusion latents directly for faster and more resource-efficient model alignment.
  • [2026-07] VeRL-Omni is presented in QingKe AI, vLLM community, and verl x Ascend Beijing meetup. Slides are shared.
  • [2026-06] Qwen3-Omni GSPO Trainer is available! Flow-DPPO is integrated. vLLM-Omni rollout backend is upgraded to v0.22 for higher throughput, with default actor attn backend switched to FA3.
  • [2026-06] DiffusionNFT and Diffusion DPO are integrated with verified recipes on Qwen-Image/SD3.5. Wan2.2 is now supported for video generation tasks.

Why VeRL-Omni

Multimodal generative RL training differs from text-only LLM RL not only in model structure, but also in I/O patterns, compute characteristics, and runtime bottlenecks. As this space grows, it deserves a dedicated training repository that can evolve quickly around its own constraints.

Scope

VeRL-Omni targets RL post-training for three families of generative models:

  1. Diffusion generative models for image, video, and audio — e.g., Qwen-Image, Wan2.2.
  2. Unified multimodal understanding + generation models — e.g., BAGEL, HunyuanImage-3.0.
  3. Omni-modality models that jointly handle text, image, audio, and video — e.g., Qwen3-Omni.

What we focus on

  • Fast multi-modal rollout: Adopt vLLM-Omni backend and accelerate generation via rollout routing, rollout batching, embed caching optimizations, and more.

  • Flexible & async multi-reward serving: Support multi-reward serving (HPSv3, GenRM-OCR, UnifiedReward, etc.), HTTP scorer, and asynchronous reward computation to overlap the rollout phase.

  • Modular training backends: Selectable VeOmni and FSDP2 backends with combinable parallelism (USP/TP/DP) for distributed training.

  • Stability: Boost stability and speed in diffusion RL pipelines via rollout correction to skip logP recomputation, and achieve reproducible E2E training with deterministic RL. Reward, rollout and actor update are composable and extensible, via Hydra configs.

  • Efficient and convergent training recipes: On our reference Qwen-Image FlowGRPO setup, VeRL-Omni achieves ~25% higher end-to-end throughput than the diffusers-based flow_grpo implementation, driven by vLLM-Omni rollout, FSDP2 trainer, overlapped reward computation (asynchronous), etc.

    verl-omni architecture diagram

Getting Started 🚀

Visit our documentation to learn more.

  • Installation

  • Quickstart

     ![Training step 0](https://github.com/user-attachments/assets/cb242db0-4233-4a1b-8274-5b3d87d69a36) 
    
    
    
    
     ![Training step 120](https://github.com/user-attachments/assets/af82391b-081d-46a3-9878-9113b89452cd) 
    

Training step 0

Training step 120

Example: Optimizing Qwen-Image Text Rendering Accuracy with FlowGRPO (recipe  |  wandb)

Model and Algorithm Support 🎨

Model

Category

Modality

Algorithm

Status

Qwen-Image & Qwen-Image-Edit

Diffusion generator

Text/Image → Image

FlowGRPO (+ CPS/SDE)

✅

Flow-DPPO

✅

MixGRPO

✅

GRPO-Guard

✅

DiffusionNFT

✅

DPO

✅

Wan2.2

Diffusion generator

Text → Video

DanceGRPO

✅

LTX2.3

Diffusion generator

Text → Video + Audio

FlowGRPO

✅

MiniMax-H3

Diffusion generator

Any → Video + Audio

DiffusionNFT

✅

FlowGRPO

✅

Boogu-Image

Diffusion generator

Text/Image → Image

FlowGRPO (+ CPS/SDE)

✅

DiffusionNFT

✅

BAGEL

Unified understand + gen

Text + Image

FlowGRPO

✅

SD3.5

Diffusion generator

Text → Image

DPO

✅

FlowGRPO

✅

FlowGRPO w/ DiNa-LRM

✅

DiffusionOPD (incl. MOPD)

✅

HunyuanImage-3.0

Unified understand + gen

Text + Image

MixGRPO

Planned

SRPO

Planned

Qwen3-Omni-Thinker

Omni-modality

Text / Image / Video / Audio

DPO

✅

GSPO

✅

DAPO (Phase 1, LoRA)

✅

Qwen3-TTS

Audio-modality

Text → Audio

DPO

WIP

GSPO

WIP

GRPO

✅

Ascend NPU Support 💠

VeRL-Omni now supports Ascend NPU. For instructions on how to install and get started with FlowGRPO training on Ascend NPU, please refer to our Ascend NPU Quickstart Guide.

Roadmap 🗺

Future work is tracked in VeRL-Omni Q3 Roadmap

Contributing 🤝

Contributions are welcome.

See the contribution guide.

Join the Community: Feel free to ask questions, provide feedback, and discuss with fellow users of VeRL-Omni in our WeChat group.

Acknowledgement 🌟

verl-omni builds on the engineering foundations developed in verl and is closely aligned with multimodal inference systems such as vLLM-Omni.

Citation 📚

If you find the project helpful, please cite and star ⭐

@misc{verlomni_github,
  title        = {{VeRL-Omni: Easy, Fast, and Stable RL Training for Diffusion and Omni-Modality Models}},
  author       = {Yongxiang Huang and Cheung Kawai and Jingan Zhou and Yingshu Chen and {openYuanrong Team} and Xibin Wu},
  year         = {2026},
  howpublished = {\url{https://github.com/verl-project/verl-omni}},
  urldate      = {2026-04-28}
}