🌍 Awesome World Models
📜 A Curated List of Amazing Works in World Modeling, spanning applications in Embodied AI, Autonomous Driving, Natural Language Processing and Agents. Based on Awesome-World-Model-for-Autonomous-Driving and Awesome-World-Model-for-Robotics.

Photo Credit: Gemini-Nano-Banana🍌.
🚩 News & Updates
Major updates and announcements are shown below. Scroll for full timeline.
🚀 [2025-11] 1k+ Stars ⭐️ Under 30 Days — 🌍 Awesome World Models reached 1k github stars within 30 days of initial release, let's go!!!
🗺️ [2025-10] Enhanced Visual Navigation — Introduced badge system for papers! All entries now display
for quick access to resources.
🔥 [2025-10] Repository Launch — Awesome World Models is now live! We're building a comprehensive collection spanning Embodied AI, Autonomous Driving, NLP, and more. See CONTRIBUTING.md for how to contribute.
💡 [Ongoing] Community Contributions Welcome — Help us maintain the most up-to-date world models resource! Submit papers via PR or contact us at email.
⭐ [Ongoing] Support This Project — If you find this useful, please cite our work and give us a star. Share with your research community!
Overview
- 🎯 Aim of the project
- 📚 Definition of World Models
- 🧭 Tutorials & Starter Resources
- 📖 Surveys of World Models
- 🎮 World Models for Game Simulation
- 🚗 World Models for Autonomous Driving
- 🤖 World Models for Embodied AI
- 🔬 World Models for Science
- 💭 Positions on World Models
- 📐 Theory & World Models Explainability
- 🛠️ General Approaches to World Models
- 📊 Evaluating World Models
- 🙏 Acknowledgements
- 📝 Citation
Aim of the Project
World Models have become a hot topic in both research and industry, attracting unprecedented attention from the AI community and beyond. However, due to the interdisciplinary nature of the field (and because the term "world model" simply sounds amazing), the concept has been used with varying definitions across different domains.

This repository aims to:
- 🔍 Organize the rapidly growing body of world model research across multiple application domains
- 🗺️ Provide a minimalist map of how world models are utilized in different fields (Embodied AI, Autonomous Driving, NLP, etc.)
- 🤝 Bridge the gap between different communities working on world models with varying perspectives
- 📚 Serve as a one-stop resource for researchers, practitioners, and enthusiasts interested in world modeling
- 🚀 Track the latest developments and breakthroughs in this exciting field
Whether you're a researcher looking for related work, a practitioner seeking implementation references, or simply curious about world models, we hope this curated list helps you navigate the landscape!
Definition of World Models
While world models' outreach has been expanded again and again, it is widely adopted that the original sources of world models come from these two papers:
- [⭐️] World Models, World Models.
- [⭐️] Yann Lecun's Speech, "A Path Towards Autonomous Machine Intelligence".
Some other great blogposts on world models include:
- [⭐️] Towards Video World Models, "Towards Video World Models".
- Status of World Models in 2025, "Beyond the Hype: How I See World Models Evolving in 2025".
- [⭐️] Jim Fan's tweet.
- [⭐️] World Model Workshop at Montreal, "Keynote: Yoshua Bengio, Yann Lecun, Jurgen Schmidhuber, Sherry Yang, Shirley Ho, etc. (Youtube Streamning included)".
Tutorials & Starter Resources
- [⭐️] Nano World Models, "Nano World Models: A Minimalist Implementation of Future Video Prediction".
- [⭐️] minWM, "minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models".
- StableWM, "stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation".
- WorldFoundry, "WorldFoundry: Unified World Model Inference & Evaluation Infrastructure".
- world-models.io, "A structured knowledge platform for discovering and comparing AI world models across robotics, model-based reinforcement learning, simulation engines, embodied AI, and autonomous systems, with model profiles, comparisons, research, benchmarks, and a practical taxonomy."
Surveys of World Models
1. World Models and Video Generation:
- [⭐️] Is Sora a World Simulator, "Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond".
- Physics Cognition in Video Generation, "Exploring the Evolution of Physics Cognition in Video Generation: A Survey".
- From Generation to Simulation, "From Generation to Simulation: How Far Are World Models from Being True Simulators?".
2. World Models and 3D Generation:
- [⭐️] 3D and 4D World Modeling: A Survey, "3D and 4D World Modeling: A Survey".
- [⭐️] Understanding World or Predicting Future?, "Understanding World or Predicting Future? A Comprehensive Survey of World Models".
- From 2D to 3D Cognition, "From 2D to 3D Cognition: A Brief Survey of General World Models".
3. World Models and Embodied Artificial Intelligence:
- [⭐️] World Models for Embodied AI, "A Comprehensive Survey on World Models for Embodied AI".
- World Models and Physical Simulation, "A Survey: Learning Embodied Intelligence from Physical Simulators and World Models".
- Embodied AI Agents: Modeling the World, "Embodied AI Agents: Modeling the World".
- Aligning Cyber Space with Physical World, "Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI".
- Physical Grounding in World Models, "From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models".
- [⭐️] World Action Models v.s. VLA, "Do World Action Models Generalize Better than VLAs? A Robustness Study".
- WorldScape, "WorldScape: A Unified Real-time World Model Integrating Locomotion And Manipulation".
- [⭐️] World Model for Robot Learning: "World Model for Robot Learning: A Comprehensive Survey".
4. World Models for Autonomous Driving:
- [⭐️] A Survey of World Models for Autonomous Driving, "A Survey of World Models for Autonomous Driving".
- World Models for Autonomous Driving: An Initial Survey, "World Models for Autonomous Driving: An Initial Survey".
- Interplay Between Video Generation and World Models in Autonomous Driving, "Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey".
- Progressive Robustness-Aware World Models in Autonomous Driving: A Survey, "Progressive Robustness-Aware World Models in Autonomous Driving: A Survey".
5. Other Good Surveys:
- From Masks to Worlds, "From Masks to Worlds: A Hitchhiker's Guide to World Models".
- The Safety Challenge of World Models, "The Safety Challenge of World Models for Embodied AI Agents: A Review".
- World Models in AI: Like a Child, "World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a Child".
- World Model Safety, "World Models: The Safety Perspective".
- Model-based reinforcement learning: "A survey on model-based reinforcement learning".
- On Memory: A comparison of memory mechanisms in world models: "On Memory: A comparison of memory mechanisms in world models".
World Models for Game Simulation
Pixel Space:
- [⭐️] GameNGen, "Diffusion Models Are Real-Time Game Engines".
- [⭐️] DIAMOND, "Diffusion for World Modeling: Visual Details Matter in Atari".
- MineWorld, "MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft".
- Oasis, "Oasis: A Universe in a Transformer".
- AnimeGamer, "AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction".
- [⭐️] Matrix-Game, "Matrix-Game: Interactive World Foundation Model."
- [⭐️] Matrix-Game 2.0, Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World Model.
- [⭐️] Matrix-Game 3.0, "Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory".
- Matrix-Game 3.5, "Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory".
- RealPlay, "From Virtual Games to Real-World Play".
- GameFactory, "GameFactory: Creating New Games with Generative Interactive Videos".
- WORLDMEM, "Worldmem: Long-term Consistent World Simulation with Memory".
- Waypoint-1, "The Path to Real-Time Worlds and Why It Matters".
- SCOPE, "SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models".
- MIRA, "Multiplayer Interactive World Models with Representation Autoencoders".
- Solaris, "Solaris: Building a Multiplayer Video World Model in Minecraft".
- ABot-World-0: "ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU".
- ForgeWM: "ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models".
- Programmable World Model: "Programmable World Model".
3D Mesh Space:
- [⭐️] HunyuanWorld 1.0, HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels.
- [⭐️] Matrix-3D, Matrix-3D: Omnidirectional Explorable 3D World Generation.
World Models for Autonomous Driving
Refer to https://github.com/LMD0311/Awesome-World-Model for full list.
[!NOTE] 📢 [Call for Maintenance] The repo creator is no expert of autonomous driving, so this is a more-than-concise list of works without classification. We anticipate community effort on turning this section cleaner and more well-sorted.
- [⭐️] Cosmos-Drive-Dreams, "Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models".
- [⭐️] OmniDreams, "NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation".
- [⭐️] GAIA-2, "GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving".
- WorldLens, "WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World".
- Copilot4D, "Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion".
- OmniNWM: "OmniNWM: Omniscient Driving Navigation World Models".
- GAIA-1, "Introducing GAIA-1: A Cutting-Edge Generative AI Model for Autonomy".
-
PWM, "From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction".
-
Dream4Drive, "Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks".
-
SparseWorld, "SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries".
-
DriveVLA-W0: "DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving".
-
"Enhancing Physical Consistency in Lightweight World Models".
-
IRL-VLA: "IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model".
-
LiDARCrafter: "LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences".
-
FASTopoWM: "FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models".
-
Orbis: "Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models".
-
"World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving".
-
NRSeg: "NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models"
-
World4Drive: "World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model".
-
Epona: "Epona: Autoregressive Diffusion World Model for Autonomous Driving".
-
"Towards foundational LiDAR world models with efficient latent flow matching".
-
SceneDiffuser++: "SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model".
-
COME: "COME: Adding Scene-Centric Forecasting Control to Occupancy World Model"
-
STAGE: "STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation".
-
ReSim: "ReSim: Reliable World Simulation for Autonomous Driving".
-
"Ego-centric Learning of Communicative World Models for Autonomous Driving".
-
V2XCrafter: "V2XCrafter: Learning to Generate Driving Scene Across Agents".
-
Dreamland: "Dreamland: Controllable World Creation with Simulator and Generative Models".
-
LongDWM: "LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model".
-
GeoDrive: "GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control".
-
FutureSightDrive: "FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving".
-
Raw2Drive: "Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)".
-
VL-SAFE: "VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving".
-
PosePilot: "PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth".
-
"World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks".
-
DriVerse: "DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment".
-
"End-to-End Driving with Online Trajectory Evaluation via BEV World Model".
-
"Knowledge Graphs as World Models for Semantic Material-Aware Obstacle Handling in Autonomous Vehicles".
-
MiLA: "MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving".
-
SimWorld: "SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model".
-
UniFuture: "Seeing the Future, Perceiving the Future: A Unified Driving World Model for Future Generation and Perception".
-
EOT-WM: "Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space".
-
InDRiVE: "InDRiVE: Intrinsic Disagreement based Reinforcement for Vehicle Exploration through Curiosity Driven Generalized World Model".
-
MaskGWM: "MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction".
-
Dream to Drive: "Dream to Drive: Model-Based Vehicle Control Using Analytic World Models".
-
"Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving".
-
HERMES: "HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation".
-
AdaWM: "AdaWM: Adaptive World Model based Planning for Autonomous Driving".
-
AD-L-JEPA: "AD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR Data".
-
DrivingWorld: "DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT".
-
DrivingGPT: "DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers".
-
"An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training".
-
GEM: "GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control".
-
GaussianWorld: "GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction".
-
Doe-1: "Doe-1: Closed-Loop Autonomous Driving with Large World Model".
-
InfiniCube: "InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models".
-
InfinityDrive: "InfinityDrive: Breaking Time Limits in Driving World Models".
-
ReconDreamer: "ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration".
-
Imagine-2-Drive: "Imagine-2-Drive: High-Fidelity World Modeling in CARLA for Autonomous Vehicles".
-
DynamicCity: "DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes".
-
DriveDreamer4D: "World Models Are Effective Data Machines for 4D Driving Scene Representation".
-
DOME: "Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model".
-
SSR: "Does End-to-End Autonomous Driving Really Need Perception Tasks?".
-
"Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models".
-
LatentDriver: "Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving".
-
OccLLaMA: "An Occupancy-Language-Action Generative World Model for Autonomous Driving".
-
DriveGenVLM: "Real-world Video Generation for Vision Language Model based Autonomous Driving".
-
Drive-OccWorld: "Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving".
-
CarFormer: "Self-Driving with Learned Object-Centric Representations".
-
BEVWorld: "A Multimodal World Model for Autonomous Driving via Unified BEV Latent Space".
-
TOKEN: "Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving".
-
UMAD: "Unsupervised Mask-Level Anomaly Detection for Autonomous Driving".
-
AdaptiveDriver: "Planning with Adaptive World Models for Autonomous Driving".
-
UnO: "Unsupervised Occupancy Fields for Perception and Forecasting".
-
LAW: "Enhancing End-to-End Autonomous Driving with Latent World Model".
-
Delphi: "Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation".
-
OccSora: "4D Occupancy Generation Models as World Simulators for Autonomous Driving".
-
MagicDrive3D: "Controllable 3D Generation for Any-View Rendering in Street Scenes".
-
Vista: "A Generalizable Driving World Model with High Fidelity and Versatile Controllability".
-
CarDreamer: "Open-Source Learning Platform for World Model based Autonomous Driving".
-
DriveSim: "Probing Multimodal LLMs as World Models for Driving".
-
DriveWorld: "4D Pre-trained Scene Understanding via World Models for Autonomous Driving".
-
LidarDM: "Generative LiDAR Simulation in a Generated World".
-
SubjectDrive: "Scaling Generative Data in Autonomous Driving via Subject Control".
-
DriveDreamer-2: "LLM-Enhanced World Models for Diverse Driving Video Generation".
-
Think2Drive: "Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving".
-
MARL-CCE: "Modelling Competitive Behaviors in Autonomous Driving Under Generative World Model".
-
GenAD: "Generalized Predictive Model for Autonomous Driving".
-
NeMo: "Neural Volumetric World Models for Autonomous Driving".
-
MARL-CCE: "Modelling-Competitive-Behaviors-in-Autonomous-Driving-Under-Generative-World-Model".
-
ViDAR: "Visual Point Cloud Forecasting enables Scalable Autonomous Driving".
-
Drive-WM: "Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving".
-
Cam4DOCC: "Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications".
-
Panacea: "Panoramic and Controllable Video Generation for Autonomous Driving".
-
OccWorld: "Learning a 3D Occupancy World Model for Autonomous Driving".
-
DrivingDiffusion: "Layout-Guided multi-view driving scene video generation with latent diffusion model".
-
SafeDreamer: "Safe Reinforcement Learning with World Models".
-
MagicDrive: "Street View Generation with Diverse 3D Geometry Control".
-
DriveDreamer: "Towards Real-world-driven World Models for Autonomous Driving".
-
SEM2: "Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model".
-
COMPARATIVE STUDY OF WORLD MODELS: "COMPARATIVE STUDY OF WORLD MODELS, NVAE- BASED HIERARCHICAL MODELS, AND NOISYNET- AUGMENTED MODELS IN CARRACING-V2".
-
Knowledge Graphs as World Models: "Knowledge Graphs as World Models for Material-Aware Obstacle Handling in Autonomous Vehicles".
-
Uncertainty Modeling: "Uncertainty Modeling in Autonomous Vehicle Trajectory Prediction: A Comprehensive Survey".
-
Divide and Merge: "Divide and Merge: Motion and Semantic Learning in End-to-End Autonomous Driving".
-
RDAR: "RDAR: Reward-Driven Agent Relevance Estimation for Autonomous Driving".
World Models for Embodied AI
1. Foundation Embodied World Models
- [⭐️] Genie Envisioner: "Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation".
- [⭐️] WoW, "WoW: Towards a World omniscient World model Through Embodied Interaction".
- UnifoLM-WMA-0, "UnifoLM-WMA-0: A World-Model-Action (WMA) Framework under UnifoLM Family".
- [⭐️] iVideoGPT, "iVideoGPT: Interactive VideoGPTs are Scalable World Models".
- Direct Robot Configuration Space Construction: "Direct Robot Configuration Space Construction using Convolutional Encoder-Decoders".
- ViPRA: "ViPRA: Video Prediction for Robot Actions".
- ROPES: "ROPES: Robotic Pose Estimation via Score-based Causal Representation Learning".
- [⭐️] Qwen-RobotWorld, "Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation".
- [⭐️] Xiaomi-Robotics-U0: "Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model".
- Masked Visual Actions: "Masked Visual Actions for Unified World Modeling".
- Hydra-0: "Hydra-0: Action Flow for Generalist World Modeling and Control".
- Zero-WAM: "Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization".
- WALL-SS: "WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression".
- Riemann-1.0: "Riemann-1.0: An Embodied World Action Model for Physical AI".
- ZimaBlue: "ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training".
- [⭐️] OpenWAM: "OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining".
- SyncWorld: "SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators".
- LeWAM: "Latent evolving World Action Model".
- InternW0-Δ: "InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data".
2. World Models for Manipulation
- [⭐️] PointWorld,
- [⭐️] Dex-WM, "".
- [⭐️] FLARE, "FLARE: Robot Learning with Implicit World Modeling".
- [⭐️] Enerverse, "EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation".
- [⭐️] AgiBot-World, "AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems".
- [⭐️] DyWA: "DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation"
- [⭐️] TesserAct, "TesserAct: Learning 4D Embodied World Models".
- [⭐️] DreamGen: "DreamGen: Unlocking Generalization in Robot Learning through Video World Models".
- [⭐️] HiP, "Compositional Foundation Models for Hierarchical Planning".
- [⭐️] VLA-JEPA, "VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model".
- PAR: "Physical Autoregressive Model for Robotic Manipulation without Action Pretraining".
- iMoWM: "iMoWM: Taming Interactive Multi-Modal World Model for Robotic Manipulation".
- WristWorld: "WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation".
- "A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models".
- EMMA: "EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer".
- PhysTwin, "PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos".
- [⭐️] KeyWorld: "KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models".
- World4RL: "World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation".
- [⭐️] SAMPO: "SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models".
- PhysicalAgent: "PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models".
- "Empowering Multi-Robot Cooperation via Sequential World Models".
- [⭐️] "Learning Primitive Embodied World Models: Towards Scalable Robotic Learning".
- [⭐️] GWM: "GWM: Towards Scalable Gaussian World Models for Robotic Manipulation".
- [⭐️] Flow-as-Action, "Latent Policy Steering with Embodiment-Agnostic Pretrained World Models".
- EmbodieDreamer: "EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling".
- WEAVER: "WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation".
- PhysisForcing: "PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation".
- RynnWorld-4D: "RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation".
- RynnWorld-Teleop: "RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation".
- RoboScape: "RoboScape: Physics-informed Embodied World Model".
- FWM, "Factored World Models for Zero-Shot Generalization in Robotic Manipulation".
- [⭐️] ParticleFormer: "ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation".
- ManiGaussian++: "ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model".
- ReOI: "Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control".
- GAF: "GAF: Gaussian Action Field as a Dynamic World Model for Robotic Manipulation".
- "Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins".
- "Time-Aware World Model for Adaptive Prediction and Control".
- [⭐️] 3DFlowAction: "3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model".
- [⭐️] ORV: "ORV: 4D Occupancy-centric Robot Video Generation".
- [⭐️] WoMAP: "WoMAP: World Models For Embodied Open-Vocabulary Object Localization".
- "Sparse Imagination for Efficient Visual World Model Planning".
- [⭐️] OSVI-WM: "OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation".
- [⭐️] LaDi-WM: "LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation".
- FlowDreamer: "FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation".
- PIN-WM: "PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation".
- RoboMaster, "Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control".
- ManipDreamer: "ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance".
- [⭐️] AdaWorld: "AdaWorld: Learning Adaptable World Models with Latent Actions"
- "Towards Suturing World Models: Learning Predictive Models for Robotic Surgical Tasks"
- [⭐️] EVA: "EVA: An Embodied World Model for Future Video Anticipation".
- "Representing Positional Information in Generative World Models for Object Manipulation".
- DexSim2Real$^2$: "DexSim2Real$^2: Building Explicit World Model for Precise Articulated Object Dexterous Manipulation".
- "Physically Embodied Gaussian Splatting: A Realtime Correctable World Model for Robotics".
- [⭐️] LUMOS: "LUMOS: Language-Conditioned Imitation Learning with World Models".
- [⭐️] "Object-Centric World Model for Language-Guided Manipulation"
- [⭐️] DEMO^3: "Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning"
- "Strengthening Generative Robot Policies through Predictive World Modeling".
- RoboHorizon: "RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation.
- Dream to Manipulate: "Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination".
- [⭐️] RoboDreamer: "RoboDreamer: Learning Compositional World Models for Robot Imagination".
- [⭐️] Vidar: "Vidar: Embodied Video Diffusion Model for Generalist Manipulation".
- ManiGaussian: "ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation".
- [⭐️] WHALE: "WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making".
- [⭐️] VisualPredicator: "VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning".
- [⭐️] "Multi-Task Interactive Robot Fleet Learning with Visual World Models".
- PIVOT-R: "PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation".
- Video2Action, "Grounding Video Models to Actions through Goal Conditioned Exploration".
- Diffuser, "Planning with Diffusion for Flexible Behavior Synthesis".
- Decision Diffuser, "Is Conditional Generative Modeling all you need for Decision-Making?".
- Potential Based Diffusion Motion Planning, "Potential Based Diffusion Motion Planning".
-
GRIM: "GRIM: Task-Oriented Grasping with Conditioning on Generative Examples".
-
World4Omni: "World4Omni: A Zero-Shot Framework from Image Generation World Model to Robotic Manipulation".
-
In-Context Policy Iteration: "In-Context Policy Iteration for Dynamic Manipulation".
-
HDFlow: "HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon Robotic Assembly".
-
Mobile Manipulation with Active Inference: "Mobile Manipulation with Active Inference for Long-Horizon Rearrangement Tasks".
- [⭐️] Hierarchical World Model, "H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model".
- Cosmos Policy: "Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning".
- DiT4DiT: "DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control".
- GEM-4D, "GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation".
- GaussianDream, "GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation".
- Demo-JEPA, "Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation".
- Action-State Consistency, "Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models".
- DriftWorld: "DriftWorld: Fast World Modeling through Drifting".
- Robot-Factored World Models: "Robot-Factored World Models via Robot Rendering".
- DreamX-Phi 1.0: "DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation".
- EgoGenesis: "EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE".
- N0-TWAM, "N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation".
3. World Models for Navigation
- [⭐️] NWM, "Navigation World Models".
- EfficientNWM, "An Efficient and Multi-Modal Navigation System with One-Step World Model".
- [⭐️] MindJourney: "MindJourney: Test-Time Scaling with World Models for Spatial Reasoning".
-
Test-Time Scaling: "Test-Time Scaling with World Models for Spatial Reasoning".
-
Scaling Inference-Time Search: "Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension".
-
FalconWing: "FalconWing: An Ultra-Light Fixed-Wing Platform for Indoor Aerial Applications".
-
Foundation Models as World Models: "Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds".
-
Geosteering Through the Lens of Decision Transformers: "Geosteering Through the Lens of Decision Transformers: Toward Embodied Sequence Decision-Making".
-
Latent Weight Diffusion: "Latent Weight Diffusion: Generating reactive policies instead of trajectories".
-
Abstract Sim2Real: "Abstract Sim2Real through Approximate Information States".
-
FLAM: "FLAM: Scaling Latent Action Models with Factorization".
- NavMorph: "NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments".
- Unified World Models: "Unified World Models: Memory-Augmented Planning and Foresight for Visual Navigation".
[code]
- RECON, "Rapid Exploration for Open-World Navigation with Latent Goal Models".
- WMNav: "WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation".
- NavCoT, "NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning".
- NaVi-WM, "Deductive Chain-of-Thought Augmented Socially-aware Robot Navigation World Model".
- AIF, "Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation".
- "Kinodynamic Motion Planning for Mobile Robot Navigation across Inconsistent World Models".
- "World Model Implanting for Test-time Adaptation of Embodied Agents".
- "Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation".
- [⭐️] Persistent Embodied World Models, "Learning 3D Persistent Embodied World Models".
- "Perspective-Shifted Neuro-Symbolic World Models: A Framework for Socially-Aware Robot Navigation"
- X-MOBILITY: "X-MOBILITY: End-To-End Generalizable Navigation via World Modeling".
- MWM, "Masked World Models for Visual Control".
4. World Models for Locomotion
Locomotion:
- [⭐️] Ego-VCP, "Ego-Vision World Model for Humanoid Contact Planning".
- [⭐️] RWM-O, "Offline Robotic World Model: Learning Robotic Policies without a Physics Simulator".
- [⭐️] DWL: "Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning".
- HRSSM: "Learning Latent Dynamic Robust Representations for World Models".
- WMP: "World Model-based Perception for Visual Legged Locomotion".
- TrajWorld, "Trajectory World Models for Heterogeneous Environments".
- Puppeteer: "Hierarchical World Models as Visual Whole-Body Humanoid Controllers".
- ProTerrain: "ProTerrain: Probabilistic Physics-Informed Rough Terrain World Modeling".
- Occupancy World Model, "Occupancy World Model for Robots".
- [⭐️] "Accelerating Model-Based Reinforcement Learning with State-Space World Models".
- [⭐️] "Learning Humanoid Locomotion with World Model Reconstruction".
- [⭐️] Robotic World Model: "Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics.
Loco-Manipulation:
- [⭐️] 1X World Model, 1X World Model.
- [⭐️] GROOT-Dreams, "Dream Come True — NVIDIA Isaac GR00T-Dreams Advances Robot Training With Synthetic Data and Neural Simulation".
- Humanoid World Models: "Humanoid World Models: Open World Foundation Models for Humanoid Robotics".
- Ego-Agent, "EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds".
- D^2PO, "World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning"
- COMBO: "COMBO: Compositional World Models for Embodied Multi-Agent Cooperation.
-
Scalable Humanoid Whole-Body Control: "Scalable Humanoid Whole-Body Control via Differentiable Neural Network Dynamics".
-
HuWo: "HuWo: Building Physical Interaction World Models for Humanoid Robot Locomotion".
-
Bridging the Sim-to-Real Gap: "Bridging the Sim-to-Real Gap in Humanoid Dynamics via Learned Nonlinear Operators".
- GigaBrain-WBC-0.5: "GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction".
5. World Models x VLAs
Unifying World Models and VLAs in one model:
- [⭐️] CoT-VLA: "CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models".
- [⭐️] UP-VLA, "UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent".
- [⭐️] VPP, "Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations".
- [⭐️] FLARE: "FLARE: Robot Learning with Implicit World Modeling".
- [⭐️] MinD: "MinD: Unified Visual Imagination and Control via Hierarchical World Models".
- [⭐️] DreamVLA, "DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge".
- [⭐️] WorldVLA: "WorldVLA: Towards Autoregressive Action World Model".
- 3D-VLA: "3D-VLA: A 3D Vision-Language-Action Generative World Model".
- LAWM: "Latent Action Pretraining Through World Modeling".
- [⭐️] UniVLA: "UniVLA: Unified Vision-Language-Action Model".
- [⭐️] dVLA, "dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought".
- [⭐️] Vidar, "Vidar: Embodied Video Diffusion Model for Generalist Manipulation".
- [⭐️] UD-VLA, "Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process".
- Goal-VLA: "Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation".
- Vidarc: "Vidarc: Embodied Video Diffusion Model for Closed-loop Control".
- [⭐️] VideoVLA: "VideoVLA: Video Generators Can Be Generalizable Robot Manipulators".
- [⭐️] Motus: "Motus: A Unified Latent Action World Model".
- [⭐️] mimic-video: "mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs".
- InternVLA-A1.5, "InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization".
Combining World Models and VLAs:
- [⭐️] Ctrl-World: "Ctrl-World: A Controllable Generative World Model for Robot Manipulation".
- VLA-RFT: "VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators".
- World-Env: "World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training".
- [⭐️] Self-Improving Embodied Foundation Models, "Self-Improving Embodied Foundation Models".
- GigaBrain-0.5M*, GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning.
- RISE: "RISE: Self-Improving Robot Policy with Compositional World Model".
- GigaBrain-0, GigaBrain-0: A World Model-Powered Vision-Language-Action Model.
-
NinA: "NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows".
-
Ada-Diffuser: "Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-Making".
-
Steering Diffusion Policies: "Steering Diffusion Policies with Value-Guided Denoising".
-
SPUR: "SPUR: Scaling Reward Learning from Human Demonstrations".
-
A Smooth Sea Never Made a Skilled SAILOR: "A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search".
-
RADI: "RADI: LLMs as World Models for Robotic Action Decomposition and Imagination".
-
WMPO: "WMPO: World Model-based Policy Optimization for Vision-Language-Action Models".
-
Motus2: "Motus2: A Self-Evolving General World Model for Dexterous Manipulation".
6. World Models x Policy Learning
This subsection focuses on general policy learning methods in embodied intelligence via leveraging world models.
- [⭐️] LingBot-VA, "LingBot-VA: Causal video-action world model for generalist robot control".
- [⭐️] LingBot-VA 2.0, "Native Video-Action Pretraining for Generalizable Robot Control".
- [⭐️] UWM, "Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets".
- [⭐️] UVA, Unified Video Action Model.
- DiWA, "DiWA: Diffusion Policy Adaptation with World Models".
- [⭐️] Dreamerv4, "Training Agents Inside of Scalable World Models".
- LVP, "Large Video Planner Enables Generalizable Robot Control".
- [⭐️] LDA-1B, "LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion".
- ABot-M0.5, "ABot-M0.5: Unified Mobility-and-Manipulation World Action Model".
- EgoWAM: "EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data".
- WAM-TTT: "WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time".
- FlowWAM: "FlowWAM: Optical Flow as a Unified Action Representation for World Action Models".
- GigaWorld-Policy-0.5: "GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch".
- WorldScape Policy 2.0: "WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory".
- [⭐️] Dyna-2: "Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models".
-
Latent Action Learning Requires Supervision: "Latent Action Learning Requires Supervision in the Presence of Distractors".
-
Beyond Experience: "Beyond Experience: Fictive Learning as an Inherent Advantage of World Models".
-
Robotic World Model: "Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics".
-
Sim-to-Real Contact-Rich Pivoting: "Sim-to-Real Contact-Rich Pivoting via Optimization-Guided RL with Vision and Touch".
[![Websit