← Open Source
knightnemo

Awesome-World-Models

A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.

ListsPaper collectionsDomain-specific
Open on GitHub
Momentum
+3stars in 24 hours+0.1%
3.47k
Stars
167
Forks
+16
This week
38
Contributors
Created 2025-10-31 · Updated 2026-10-06 · #2505 today
Top developers
README

🌍 Awesome World Models

Awesome GitHub stars License PRs Welcome

📜 A Curated List of Amazing Works in World Modeling, spanning applications in Embodied AI, Autonomous Driving, Natural Language Processing and Agents. Based on Awesome-World-Model-for-Autonomous-Driving and Awesome-World-Model-for-Robotics.

Awesome World Models

Photo Credit: Gemini-Nano-Banana🍌.


🚩 News & Updates

Major updates and announcements are shown below. Scroll for full timeline.

🚀 [2025-11] 1k+ Stars ⭐️ Under 30 Days — 🌍 Awesome World Models reached 1k github stars within 30 days of initial release, let's go!!!

🗺️ [2025-10] Enhanced Visual Navigation — Introduced badge system for papers! All entries now display arXiv Website Code for quick access to resources.

🔥 [2025-10] Repository Launch — Awesome World Models is now live! We're building a comprehensive collection spanning Embodied AI, Autonomous Driving, NLP, and more. See CONTRIBUTING.md for how to contribute.

💡 [Ongoing] Community Contributions Welcome — Help us maintain the most up-to-date world models resource! Submit papers via PR or contact us at email.

⭐ [Ongoing] Support This Project — If you find this useful, please cite our work and give us a star. Share with your research community!


Overview


Aim of the Project

World Models have become a hot topic in both research and industry, attracting unprecedented attention from the AI community and beyond. However, due to the interdisciplinary nature of the field (and because the term "world model" simply sounds amazing), the concept has been used with varying definitions across different domains.

Awesome World Models

This repository aims to:

  • 🔍 Organize the rapidly growing body of world model research across multiple application domains
  • 🗺️ Provide a minimalist map of how world models are utilized in different fields (Embodied AI, Autonomous Driving, NLP, etc.)
  • 🤝 Bridge the gap between different communities working on world models with varying perspectives
  • 📚 Serve as a one-stop resource for researchers, practitioners, and enthusiasts interested in world modeling
  • 🚀 Track the latest developments and breakthroughs in this exciting field

Whether you're a researcher looking for related work, a practitioner seeking implementation references, or simply curious about world models, we hope this curated list helps you navigate the landscape!


Definition of World Models

While world models' outreach has been expanded again and again, it is widely adopted that the original sources of world models come from these two papers:

  • [⭐️] World Models, World Models. arXiv Website
  • [⭐️] Yann Lecun's Speech, "A Path Towards Autonomous Machine Intelligence". OpenReview

Some other great blogposts on world models include:

  • [⭐️] Towards Video World Models, "Towards Video World Models". Blog
  • Status of World Models in 2025, "Beyond the Hype: How I See World Models Evolving in 2025". Blog
  • [⭐️] Jim Fan's tweet. Blog
  • [⭐️] World Model Workshop at Montreal, "Keynote: Yoshua Bengio, Yann Lecun, Jurgen Schmidhuber, Sherry Yang, Shirley Ho, etc. (Youtube Streamning included)". Website Blog

Tutorials & Starter Resources

  • [⭐️] Nano World Models, "Nano World Models: A Minimalist Implementation of Future Video Prediction". arXiv Website Code
  • [⭐️] minWM, "minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models". arXiv Code
  • StableWM, "stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation". arXiv Website Code
  • WorldFoundry, "WorldFoundry: Unified World Model Inference & Evaluation Infrastructure". Website Code
  • world-models.io, "A structured knowledge platform for discovering and comparing AI world models across robotics, model-based reinforcement learning, simulation engines, embodied AI, and autonomous systems, with model profiles, comparisons, research, benchmarks, and a practical taxonomy." Website Blog

Surveys of World Models

1. World Models and Video Generation:

  • [⭐️] Is Sora a World Simulator, "Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond". arXiv Website
  • Physics Cognition in Video Generation, "Exploring the Evolution of Physics Cognition in Video Generation: A Survey". arXiv Website
  • From Generation to Simulation, "From Generation to Simulation: How Far Are World Models from Being True Simulators?". arXiv

2. World Models and 3D Generation:

  • [⭐️] 3D and 4D World Modeling: A Survey, "3D and 4D World Modeling: A Survey". arXiv
  • [⭐️] Understanding World or Predicting Future?, "Understanding World or Predicting Future? A Comprehensive Survey of World Models". arXiv
  • From 2D to 3D Cognition, "From 2D to 3D Cognition: A Brief Survey of General World Models". arXiv

3. World Models and Embodied Artificial Intelligence:

  • [⭐️] World Models for Embodied AI, "A Comprehensive Survey on World Models for Embodied AI". arXiv Website
  • World Models and Physical Simulation, "A Survey: Learning Embodied Intelligence from Physical Simulators and World Models". arXiv Website
  • Embodied AI Agents: Modeling the World, "Embodied AI Agents: Modeling the World". arXiv
  • Aligning Cyber Space with Physical World, "Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI". arXiv Website
  • Physical Grounding in World Models, "From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models". arXiv
  • [⭐️] World Action Models v.s. VLA, "Do World Action Models Generalize Better than VLAs? A Robustness Study". arXiv
  • WorldScape, "WorldScape: A Unified Real-time World Model Integrating Locomotion And Manipulation". Blog
  • [⭐️] World Model for Robot Learning: "World Model for Robot Learning: A Comprehensive Survey". arXiv Website Code

4. World Models for Autonomous Driving:

  • [⭐️] A Survey of World Models for Autonomous Driving, "A Survey of World Models for Autonomous Driving". arXiv
  • World Models for Autonomous Driving: An Initial Survey, "World Models for Autonomous Driving: An Initial Survey". arXiv
  • Interplay Between Video Generation and World Models in Autonomous Driving, "Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey". arXiv
  • Progressive Robustness-Aware World Models in Autonomous Driving: A Survey, "Progressive Robustness-Aware World Models in Autonomous Driving: A Survey". DOI Code

5. Other Good Surveys:

  • From Masks to Worlds, "From Masks to Worlds: A Hitchhiker's Guide to World Models". arXiv Website
  • The Safety Challenge of World Models, "The Safety Challenge of World Models for Embodied AI Agents: A Review". arXiv
  • World Models in AI: Like a Child, "World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a Child". arXiv
  • World Model Safety, "World Models: The Safety Perspective". arXiv
  • Model-based reinforcement learning: "A survey on model-based reinforcement learning". Website
  • On Memory: A comparison of memory mechanisms in world models: "On Memory: A comparison of memory mechanisms in world models". arXiv

World Models for Game Simulation

Pixel Space:

  • [⭐️] GameNGen, "Diffusion Models Are Real-Time Game Engines". arXiv
  • [⭐️] DIAMOND, "Diffusion for World Modeling: Visual Details Matter in Atari". arXiv Code
  • MineWorld, "MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft". arXiv Website
  • Oasis, "Oasis: A Universe in a Transformer". Website
  • AnimeGamer, "AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction". arXivWebsite
  • [⭐️] Matrix-Game, "Matrix-Game: Interactive World Foundation Model." arXiv Code
  • [⭐️] Matrix-Game 2.0, Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World Model. arXiv Website
  • [⭐️] Matrix-Game 3.0, "Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory". arXiv Website Code
  • Matrix-Game 3.5, "Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory". arXiv Website Code
  • RealPlay, "From Virtual Games to Real-World Play". arXiv Website Code
  • GameFactory, "GameFactory: Creating New Games with Generative Interactive Videos". arXiv Website Code
  • WORLDMEM, "Worldmem: Long-term Consistent World Simulation with Memory". arXiv Website Code
  • Waypoint-1, "The Path to Real-Time Worlds and Why It Matters". Blog
  • SCOPE, "SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models". arXiv Website Code
  • MIRA, "Multiplayer Interactive World Models with Representation Autoencoders". arXiv Website Code
  • Solaris, "Solaris: Building a Multiplayer Video World Model in Minecraft". arXiv Website Code
  • ABot-World-0: "ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU". arXiv Website Code
  • ForgeWM: "ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models". arXiv Website Code
  • Programmable World Model: "Programmable World Model". arXiv Website

3D Mesh Space:

  • [⭐️] HunyuanWorld 1.0, HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels. arXiv Website Code
  • [⭐️] Matrix-3D, Matrix-3D: Omnidirectional Explorable 3D World Generation. arXiv Website

World Models for Autonomous Driving

Refer to https://github.com/LMD0311/Awesome-World-Model for full list.

[!NOTE] 📢 [Call for Maintenance] The repo creator is no expert of autonomous driving, so this is a more-than-concise list of works without classification. We anticipate community effort on turning this section cleaner and more well-sorted.

  • [⭐️] Cosmos-Drive-Dreams, "Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models". arXiv Website
  • [⭐️] OmniDreams, "NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation". arXiv Website Code
  • [⭐️] GAIA-2, "GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving". arXiv Website
  • WorldLens, "WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World". arXiv Website
  • Copilot4D, "Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion". arXiv
  • OmniNWM: "OmniNWM: Omniscient Driving Navigation World Models". arXiv Website
  • GAIA-1, "Introducing GAIA-1: A Cutting-Edge Generative AI Model for Autonomy". arXiv Blog
  • PWM, "From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction". arXiv Code

  • Dream4Drive, "Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks". arXiv Website

  • SparseWorld, "SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries". arXiv Code

  • DriveVLA-W0: "DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving". arXiv Code

  • "Enhancing Physical Consistency in Lightweight World Models". arXiv

  • IRL-VLA: "IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model". arXiv Website Code

  • LiDARCrafter: "LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences". arXiv Website Code

  • FASTopoWM: "FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models". arXiv Code

  • Orbis: "Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models". arXiv Code

  • "World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving". arXiv

  • NRSeg: "NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models" arXiv Code

  • World4Drive: "World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model". arXiv Code

  • Epona: "Epona: Autoregressive Diffusion World Model for Autonomous Driving". arXiv Code

  • "Towards foundational LiDAR world models with efficient latent flow matching". arXiv

  • SceneDiffuser++: "SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model". arXiv

  • COME: "COME: Adding Scene-Centric Forecasting Control to Occupancy World Model" arXiv Code

  • STAGE: "STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation". arXiv

  • ReSim: "ReSim: Reliable World Simulation for Autonomous Driving". arXiv Code Website

  • "Ego-centric Learning of Communicative World Models for Autonomous Driving". arXiv

  • V2XCrafter: "V2XCrafter: Learning to Generate Driving Scene Across Agents". arXiv

  • Dreamland: "Dreamland: Controllable World Creation with Simulator and Generative Models". arXiv Website

  • LongDWM: "LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model". arXiv Website

  • GeoDrive: "GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control". arXiv Code

  • FutureSightDrive: "FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving". arXiv Code

  • Raw2Drive: "Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)". arXiv

  • VL-SAFE: "VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving". arXiv Website

  • PosePilot: "PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth". arXiv

  • "World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks". arXiv

  • "Learning to Drive from a World Model". arXiv

  • DriVerse: "DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment". arXiv

  • "End-to-End Driving with Online Trajectory Evaluation via BEV World Model". arXiv Code

  • "Knowledge Graphs as World Models for Semantic Material-Aware Obstacle Handling in Autonomous Vehicles". arXiv

  • MiLA: "MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving". arXiv Website

  • SimWorld: "SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model". arXiv Website

  • UniFuture: "Seeing the Future, Perceiving the Future: A Unified Driving World Model for Future Generation and Perception". arXiv Website

  • EOT-WM: "Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space". arXiv

  • "Temporal Triplane Transformers as Occupancy World Models". arXiv

  • InDRiVE: "InDRiVE: Intrinsic Disagreement based Reinforcement for Vehicle Exploration through Curiosity Driven Generalized World Model". arXiv

  • MaskGWM: "MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction". arXiv

  • Dream to Drive: "Dream to Drive: Model-Based Vehicle Control Using Analytic World Models". arXiv

  • "Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving". arXiv

  • "Dream to Drive with Predictive Individual World Model". arXiv Code

  • HERMES: "HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation". arXiv

  • AdaWM: "AdaWM: Adaptive World Model based Planning for Autonomous Driving". arXiv

  • AD-L-JEPA: "AD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR Data". arXiv

  • DrivingWorld: "DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT". arXiv Code Website

  • DrivingGPT: "DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers". arXiv Website

  • "An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training". arXiv

  • GEM: "GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control". arXiv Website

  • GaussianWorld: "GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction". arXiv Code

  • Doe-1: "Doe-1: Closed-Loop Autonomous Driving with Large World Model". arXiv Website Code

  • "Physical Informed Driving World Model". arXiv Website

  • InfiniCube: "InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models". arXiv Website

  • InfinityDrive: "InfinityDrive: Breaking Time Limits in Driving World Models". arXiv Website

  • ReconDreamer: "ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration". arXiv Website

  • Imagine-2-Drive: "Imagine-2-Drive: High-Fidelity World Modeling in CARLA for Autonomous Vehicles". arXiv Website

  • DynamicCity: "DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes". arXiv Website Code

  • DriveDreamer4D: "World Models Are Effective Data Machines for 4D Driving Scene Representation". arXiv Website

  • DOME: "Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model". arXiv Website

  • SSR: "Does End-to-End Autonomous Driving Really Need Perception Tasks?". arXiv Code

  • "Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models". arXiv

  • LatentDriver: "Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving". arXiv Code

  • RenderWorld: "World Model with Self-Supervised 3D Label". arXiv

  • OccLLaMA: "An Occupancy-Language-Action Generative World Model for Autonomous Driving". arXiv

  • DriveGenVLM: "Real-world Video Generation for Vision Language Model based Autonomous Driving". arXiv

  • Drive-OccWorld: "Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving". arXiv

  • CarFormer: "Self-Driving with Learned Object-Centric Representations". arXiv Code

  • BEVWorld: "A Multimodal World Model for Autonomous Driving via Unified BEV Latent Space". arXiv Code

  • TOKEN: "Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving". arXiv

  • UMAD: "Unsupervised Mask-Level Anomaly Detection for Autonomous Driving". arXiv

  • SimGen: "Simulator-conditioned Driving Scene Generation". arXiv Code

  • AdaptiveDriver: "Planning with Adaptive World Models for Autonomous Driving". arXiv Code

  • UnO: "Unsupervised Occupancy Fields for Perception and Forecasting". arXiv Code

  • LAW: "Enhancing End-to-End Autonomous Driving with Latent World Model". arXiv Code

  • Delphi: "Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation". arXiv Code

  • OccSora: "4D Occupancy Generation Models as World Simulators for Autonomous Driving". arXiv Code

  • MagicDrive3D: "Controllable 3D Generation for Any-View Rendering in Street Scenes". arXiv Code

  • Vista: "A Generalizable Driving World Model with High Fidelity and Versatile Controllability". arXiv Code

  • CarDreamer: "Open-Source Learning Platform for World Model based Autonomous Driving". arXiv Code

  • DriveSim: "Probing Multimodal LLMs as World Models for Driving". arXiv Code

  • DriveWorld: "4D Pre-trained Scene Understanding via World Models for Autonomous Driving". arXiv

  • LidarDM: "Generative LiDAR Simulation in a Generated World". arXiv Code

  • SubjectDrive: "Scaling Generative Data in Autonomous Driving via Subject Control". arXiv Website

  • DriveDreamer-2: "LLM-Enhanced World Models for Diverse Driving Video Generation". arXiv Code

  • Think2Drive: "Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving". arXiv

  • MARL-CCE: "Modelling Competitive Behaviors in Autonomous Driving Under Generative World Model". arXiv Code

  • GenAD: "Generalized Predictive Model for Autonomous Driving". arXiv Website

  • GenAD: "Generative End-to-End Autonomous Driving". arXiv Code

  • NeMo: "Neural Volumetric World Models for Autonomous Driving". arXiv

  • MARL-CCE: "Modelling-Competitive-Behaviors-in-Autonomous-Driving-Under-Generative-World-Model". Code

  • ViDAR: "Visual Point Cloud Forecasting enables Scalable Autonomous Driving". arXiv Code

  • Drive-WM: "Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving". arXiv Code

  • Cam4DOCC: "Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications". arXiv Code

  • Panacea: "Panoramic and Controllable Video Generation for Autonomous Driving". arXiv Code

  • OccWorld: "Learning a 3D Occupancy World Model for Autonomous Driving". arXiv Code

  • DrivingDiffusion: "Layout-Guided multi-view driving scene video generation with latent diffusion model". arXiv Code

  • SafeDreamer: "Safe Reinforcement Learning with World Models". arXiv Code

  • MagicDrive: "Street View Generation with Diverse 3D Geometry Control". arXiv Code

  • DriveDreamer: "Towards Real-world-driven World Models for Autonomous Driving". arXiv Code

  • SEM2: "Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model". arXiv

  • COMPARATIVE STUDY OF WORLD MODELS: "COMPARATIVE STUDY OF WORLD MODELS, NVAE- BASED HIERARCHICAL MODELS, AND NOISYNET- AUGMENTED MODELS IN CARRACING-V2". OpenReview Website

  • Knowledge Graphs as World Models: "Knowledge Graphs as World Models for Material-Aware Obstacle Handling in Autonomous Vehicles". OpenReview Website

  • Uncertainty Modeling: "Uncertainty Modeling in Autonomous Vehicle Trajectory Prediction: A Comprehensive Survey". OpenReview Website

  • Divide and Merge: "Divide and Merge: Motion and Semantic Learning in End-to-End Autonomous Driving". OpenReview Website

  • RDAR: "RDAR: Reward-Driven Agent Relevance Estimation for Autonomous Driving". OpenReview Website

World Models for Embodied AI

1. Foundation Embodied World Models

  • [⭐️] Genie Envisioner: "Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation". arXiv Website Code
  • [⭐️] WoW, "WoW: Towards a World omniscient World model Through Embodied Interaction". arXiv Website Code
  • UnifoLM-WMA-0, "UnifoLM-WMA-0: A World-Model-Action (WMA) Framework under UnifoLM Family". Website Code
  • [⭐️] iVideoGPT, "iVideoGPT: Interactive VideoGPTs are Scalable World Models". arXivWebsite Code
  • Direct Robot Configuration Space Construction: "Direct Robot Configuration Space Construction using Convolutional Encoder-Decoders". OpenReview Website
  • ViPRA: "ViPRA: Video Prediction for Robot Actions". arXiv Website Code
  • ROPES: "ROPES: Robotic Pose Estimation via Score-based Causal Representation Learning". OpenReview arXiv Website
  • [⭐️] Qwen-RobotWorld, "Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation". arXiv Blog
  • [⭐️] Xiaomi-Robotics-U0: "Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model". arXiv Website Code
  • Masked Visual Actions: "Masked Visual Actions for Unified World Modeling". arXiv Website Code
  • Hydra-0: "Hydra-0: Action Flow for Generalist World Modeling and Control". arXiv Website
  • Zero-WAM: "Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization". arXiv
  • WALL-SS: "WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression". arXiv
  • Riemann-1.0: "Riemann-1.0: An Embodied World Action Model for Physical AI". arXiv
  • ZimaBlue: "ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training". arXiv Website Code
  • [⭐️] OpenWAM: "OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining". arXiv Website Code
  • SyncWorld: "SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators". arXiv Website Code
  • LeWAM: "Latent evolving World Action Model". arXiv Code
  • InternW0-Δ: "InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data". arXiv Website Code

2. World Models for Manipulation

  • [⭐️] PointWorld, arXiv Website
  • [⭐️] Dex-WM, "". arXiv Website
  • [⭐️] FLARE, "FLARE: Robot Learning with Implicit World Modeling". arXiv Website
  • [⭐️] Enerverse, "EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation". arXiv Website
  • [⭐️] AgiBot-World, "AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems". arXiv Website Code
  • [⭐️] DyWA: "DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation" arXiv Website Code
  • [⭐️] TesserAct, "TesserAct: Learning 4D Embodied World Models". arXiv Website Code
  • [⭐️] DreamGen: "DreamGen: Unlocking Generalization in Robot Learning through Video World Models". arXiv Website Code
  • [⭐️] HiP, "Compositional Foundation Models for Hierarchical Planning". arXiv Website Code
  • [⭐️] VLA-JEPA, "VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model". arXiv Website Code
  • PAR: "Physical Autoregressive Model for Robotic Manipulation without Action Pretraining". arXiv Website Code
  • iMoWM: "iMoWM: Taming Interactive Multi-Modal World Model for Robotic Manipulation". arXiv Website
  • WristWorld: "WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation". arXiv Website Code
  • "A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models". arXiv
  • EMMA: "EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer". arXiv Website
  • PhysTwin, "PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos". arXiv Website Code
  • [⭐️] KeyWorld: "KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models". arXiv
  • World4RL: "World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation". arXiv Website
  • [⭐️] SAMPO: "SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models". arXiv
  • PhysicalAgent: "PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models". arXiv
  • "Empowering Multi-Robot Cooperation via Sequential World Models". arXiv
  • [⭐️] "Learning Primitive Embodied World Models: Towards Scalable Robotic Learning". arXiv Website
  • [⭐️] GWM: "GWM: Towards Scalable Gaussian World Models for Robotic Manipulation". arXiv Website Code
  • [⭐️] Flow-as-Action, "Latent Policy Steering with Embodiment-Agnostic Pretrained World Models". arXiv
  • EmbodieDreamer: "EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling". arXiv Website Code
  • WEAVER: "WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation". arXiv Website Code
  • PhysisForcing: "PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation". arXiv Website Code
  • RynnWorld-4D: "RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation". arXiv Website Code
  • RynnWorld-Teleop: "RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation". arXiv Website Code
  • RoboScape: "RoboScape: Physics-informed Embodied World Model". arXiv Code
  • FWM, "Factored World Models for Zero-Shot Generalization in Robotic Manipulation". arXiv
  • [⭐️] ParticleFormer: "ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation". arXiv Website
  • ManiGaussian++: "ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model". arXiv Code
  • ReOI: "Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control". arXiv
  • GAF: "GAF: Gaussian Action Field as a Dynamic World Model for Robotic Manipulation". arXiv Website
  • "Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins". arXiv Website Code
  • "Time-Aware World Model for Adaptive Prediction and Control". arXiv Website Code
  • [⭐️] 3DFlowAction: "3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model". arXiv Code
  • [⭐️] ORV: "ORV: 4D Occupancy-centric Robot Video Generation". arXiv Code Website
  • [⭐️] WoMAP: "WoMAP: World Models For Embodied Open-Vocabulary Object Localization". arXiv Website
  • "Sparse Imagination for Efficient Visual World Model Planning". arXiv
  • [⭐️] OSVI-WM: "OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation". arXiv
  • [⭐️] LaDi-WM: "LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation". arXiv Website Code
  • FlowDreamer: "FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation". arXiv Website Code
  • PIN-WM: "PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation". arXiv Website Code
  • RoboMaster, "Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control". arXiv Website Code
  • ManipDreamer: "ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance". arXiv Website
  • [⭐️] AdaWorld: "AdaWorld: Learning Adaptable World Models with Latent Actions" arXiv Website Code
  • "Towards Suturing World Models: Learning Predictive Models for Robotic Surgical Tasks" arXiv Website
  • [⭐️] EVA: "EVA: An Embodied World Model for Future Video Anticipation". arXiv Website Code
  • "Representing Positional Information in Generative World Models for Object Manipulation". arXiv
  • DexSim2Real$^2$: "DexSim2Real$^2: Building Explicit World Model for Precise Articulated Object Dexterous Manipulation". arXiv Code
  • "Physically Embodied Gaussian Splatting: A Realtime Correctable World Model for Robotics". arXiv Website Code
  • [⭐️] LUMOS: "LUMOS: Language-Conditioned Imitation Learning with World Models". arXiv Website Code
  • [⭐️] "Object-Centric World Model for Language-Guided Manipulation" arXiv
  • [⭐️] DEMO^3: "Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning" arXiv Website Code
  • "Strengthening Generative Robot Policies through Predictive World Modeling". arXiv Website
  • RoboHorizon: "RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation. arXiv
  • Dream to Manipulate: "Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination". arXiv Website
  • [⭐️] RoboDreamer: "RoboDreamer: Learning Compositional World Models for Robot Imagination". arXiv Code Website
  • [⭐️] Vidar: "Vidar: Embodied Video Diffusion Model for Generalist Manipulation". arXiv
  • ManiGaussian: "ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation". arXiv Code Website
  • [⭐️] WHALE: "WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making". arXiv
  • [⭐️] VisualPredicator: "VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning". arXiv
  • [⭐️] "Multi-Task Interactive Robot Fleet Learning with Visual World Models". arXiv Code Website
  • PIVOT-R: "PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation". arXiv Website Code
  • Video2Action, "Grounding Video Models to Actions through Goal Conditioned Exploration". arXiv Website Code
  • Diffuser, "Planning with Diffusion for Flexible Behavior Synthesis". arXiv Website Code
  • Decision Diffuser, "Is Conditional Generative Modeling all you need for Decision-Making?". arXiv Code
  • Potential Based Diffusion Motion Planning, "Potential Based Diffusion Motion Planning". arXiv Website Code
  • GRIM: "GRIM: Task-Oriented Grasping with Conditioning on Generative Examples". OpenReview Website

  • World4Omni: "World4Omni: A Zero-Shot Framework from Image Generation World Model to Robotic Manipulation". OpenReview Website

  • In-Context Policy Iteration: "In-Context Policy Iteration for Dynamic Manipulation". OpenReview Website

  • HDFlow: "HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon Robotic Assembly". OpenReview Website

  • Mobile Manipulation with Active Inference: "Mobile Manipulation with Active Inference for Long-Horizon Rearrangement Tasks". OpenReview Website

  • [⭐️] Hierarchical World Model, "H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model". arXiv
  • Cosmos Policy: "Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning". arXiv Website
  • DiT4DiT: "DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control". arXiv Website
  • GEM-4D, "GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation". arXiv Website
  • GaussianDream, "GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation". arXiv
  • Demo-JEPA, "Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation". arXiv
  • Action-State Consistency, "Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models". arXiv
  • DriftWorld: "DriftWorld: Fast World Modeling through Drifting". arXiv Website
  • Robot-Factored World Models: "Robot-Factored World Models via Robot Rendering". arXiv Website
  • DreamX-Phi 1.0: "DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation". arXiv Code
  • EgoGenesis: "EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE". arXiv
  • N0-TWAM, "N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation". arXiv Website Code

3. World Models for Navigation

  • [⭐️] NWM, "Navigation World Models". arXiv Website
  • EfficientNWM, "An Efficient and Multi-Modal Navigation System with One-Step World Model". arXiv Website Code
  • [⭐️] MindJourney: "MindJourney: Test-Time Scaling with World Models for Spatial Reasoning". arXiv Website
  • Test-Time Scaling: "Test-Time Scaling with World Models for Spatial Reasoning". arXiv OpenReview Website

  • Scaling Inference-Time Search: "Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension". OpenReview Website

  • FalconWing: "FalconWing: An Ultra-Light Fixed-Wing Platform for Indoor Aerial Applications". OpenReview Website

  • Foundation Models as World Models: "Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds". OpenReview Website

  • Geosteering Through the Lens of Decision Transformers: "Geosteering Through the Lens of Decision Transformers: Toward Embodied Sequence Decision-Making". OpenReview Website

  • Latent Weight Diffusion: "Latent Weight Diffusion: Generating reactive policies instead of trajectories". OpenReview Website

  • Abstract Sim2Real: "Abstract Sim2Real through Approximate Information States". OpenReview Website

  • FLAM: "FLAM: Scaling Latent Action Models with Factorization". OpenReview Website

  • NavMorph: "NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments". arXiv Code
  • Unified World Models: "Unified World Models: Memory-Augmented Planning and Foresight for Visual Navigation". arXiv [code]
  • RECON, "Rapid Exploration for Open-World Navigation with Latent Goal Models". arXiv Website
  • WMNav: "WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation". arXiv Website
  • NavCoT, "NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning". arXiv Code
  • NaVi-WM, "Deductive Chain-of-Thought Augmented Socially-aware Robot Navigation World Model". arXiv Website
  • AIF, "Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation". arXiv
  • "Kinodynamic Motion Planning for Mobile Robot Navigation across Inconsistent World Models". arXiv
  • "World Model Implanting for Test-time Adaptation of Embodied Agents". arXiv
  • "Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation". arXiv
  • [⭐️] Persistent Embodied World Models, "Learning 3D Persistent Embodied World Models". arXiv
  • "Perspective-Shifted Neuro-Symbolic World Models: A Framework for Socially-Aware Robot Navigation" arXiv
  • X-MOBILITY: "X-MOBILITY: End-To-End Generalizable Navigation via World Modeling". arXiv
  • MWM, "Masked World Models for Visual Control". arXiv Website Code

4. World Models for Locomotion

Locomotion:

  • [⭐️] Ego-VCP, "Ego-Vision World Model for Humanoid Contact Planning". arXiv Website Code
  • [⭐️] RWM-O, "Offline Robotic World Model: Learning Robotic Policies without a Physics Simulator". arXiv
  • [⭐️] DWL: "Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning". arXiv
  • HRSSM: "Learning Latent Dynamic Robust Representations for World Models". arXiv Code
  • WMP: "World Model-based Perception for Visual Legged Locomotion". arXiv Website
  • TrajWorld, "Trajectory World Models for Heterogeneous Environments". arXiv Code
  • Puppeteer: "Hierarchical World Models as Visual Whole-Body Humanoid Controllers". arXiv Code
  • ProTerrain: "ProTerrain: Probabilistic Physics-Informed Rough Terrain World Modeling". arXiv
  • Occupancy World Model, "Occupancy World Model for Robots". arXiv
  • [⭐️] "Accelerating Model-Based Reinforcement Learning with State-Space World Models". arXiv
  • [⭐️] "Learning Humanoid Locomotion with World Model Reconstruction". arXiv
  • [⭐️] Robotic World Model: "Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics. arXiv

Loco-Manipulation:

  • [⭐️] 1X World Model, 1X World Model. Blog
  • [⭐️] GROOT-Dreams, "Dream Come True — NVIDIA Isaac GR00T-Dreams Advances Robot Training With Synthetic Data and Neural Simulation". Blog
  • Humanoid World Models: "Humanoid World Models: Open World Foundation Models for Humanoid Robotics". arXiv
  • Ego-Agent, "EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds". arXiv
  • D^2PO, "World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning" arXiv
  • COMBO: "COMBO: Compositional World Models for Embodied Multi-Agent Cooperation. arXiv Website Code
  • Scalable Humanoid Whole-Body Control: "Scalable Humanoid Whole-Body Control via Differentiable Neural Network Dynamics". OpenReview Website

  • HuWo: "HuWo: Building Physical Interaction World Models for Humanoid Robot Locomotion". OpenReview Website

  • Bridging the Sim-to-Real Gap: "Bridging the Sim-to-Real Gap in Humanoid Dynamics via Learned Nonlinear Operators". OpenReview Website

  • GigaBrain-WBC-0.5: "GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction". arXiv

5. World Models x VLAs

Unifying World Models and VLAs in one model:

  • [⭐️] CoT-VLA: "CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models". arXiv Website
  • [⭐️] UP-VLA, "UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent". arXiv Code
  • [⭐️] VPP, "Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations". arXiv Website
  • [⭐️] FLARE: "FLARE: Robot Learning with Implicit World Modeling". arXiv Code Website
  • [⭐️] MinD: "MinD: Unified Visual Imagination and Control via Hierarchical World Models". arXiv Website
  • [⭐️] DreamVLA, "DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge". arXiv Code Website
  • [⭐️] WorldVLA: "WorldVLA: Towards Autoregressive Action World Model". arXiv Code
  • 3D-VLA: "3D-VLA: A 3D Vision-Language-Action Generative World Model". arXiv
  • LAWM: "Latent Action Pretraining Through World Modeling". arXiv Code
  • [⭐️] UniVLA: "UniVLA: Unified Vision-Language-Action Model". arXiv Code
  • [⭐️] dVLA, "dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought". arXiv
  • [⭐️] Vidar, "Vidar: Embodied Video Diffusion Model for Generalist Manipulation". arXiv
  • [⭐️] UD-VLA, "Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process". arXiv Code Website
  • Goal-VLA: "Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation". arXiv Website
  • Vidarc: "Vidarc: Embodied Video Diffusion Model for Closed-loop Control". arXiv
  • [⭐️] VideoVLA: "VideoVLA: Video Generators Can Be Generalizable Robot Manipulators". arXiv Website
  • [⭐️] Motus: "Motus: A Unified Latent Action World Model". arXiv Website
  • [⭐️] mimic-video: "mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs". arXiv Website
  • InternVLA-A1.5, "InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization". arXiv Website Code

Combining World Models and VLAs:

  • [⭐️] Ctrl-World: "Ctrl-World: A Controllable Generative World Model for Robot Manipulation". arXiv Website Code
  • VLA-RFT: "VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators". arXiv
  • World-Env: "World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training". arXiv
  • [⭐️] Self-Improving Embodied Foundation Models, "Self-Improving Embodied Foundation Models". arXiv
  • GigaBrain-0.5M*, GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning. arXiv WebsiteCode
  • RISE: "RISE: Self-Improving Robot Policy with Compositional World Model". arXiv Website
  • GigaBrain-0, GigaBrain-0: A World Model-Powered Vision-Language-Action Model. arXiv Website
  • NinA: "NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows". OpenReview Website

  • Ada-Diffuser: "Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-Making". OpenReview Website

  • Steering Diffusion Policies: "Steering Diffusion Policies with Value-Guided Denoising". OpenReview Website

  • SPUR: "SPUR: Scaling Reward Learning from Human Demonstrations". OpenReview Website

  • A Smooth Sea Never Made a Skilled SAILOR: "A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search". OpenReview Website

  • RADI: "RADI: LLMs as World Models for Robotic Action Decomposition and Imagination". OpenReview Website

  • WMPO: "WMPO: World Model-based Policy Optimization for Vision-Language-Action Models". arXiv Website

  • Motus2: "Motus2: A Self-Evolving General World Model for Dexterous Manipulation". arXiv

6. World Models x Policy Learning

This subsection focuses on general policy learning methods in embodied intelligence via leveraging world models.

  • [⭐️] LingBot-VA, "LingBot-VA: Causal video-action world model for generalist robot control". arXiv Website Code
  • [⭐️] LingBot-VA 2.0, "Native Video-Action Pretraining for Generalizable Robot Control". arXiv Website
  • [⭐️] UWM, "Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets". arXiv Website
  • [⭐️] UVA, Unified Video Action Model. arXiv Website Code
  • DiWA, "DiWA: Diffusion Policy Adaptation with World Models". arXiv Code
  • [⭐️] Dreamerv4, "Training Agents Inside of Scalable World Models". arXiv Website
  • LVP, "Large Video Planner Enables Generalizable Robot Control". arXiv Website
  • [⭐️] LDA-1B, "LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion". arXiv Website Code
  • ABot-M0.5, "ABot-M0.5: Unified Mobility-and-Manipulation World Action Model". arXiv Code
  • EgoWAM: "EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data". arXiv Website
  • WAM-TTT: "WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time". arXiv
  • FlowWAM: "FlowWAM: Optical Flow as a Unified Action Representation for World Action Models". arXiv Website Code
  • GigaWorld-Policy-0.5: "GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch". arXiv Website Code
  • WorldScape Policy 2.0: "WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory". arXiv Website Code
  • [⭐️] Dyna-2: "Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models". Blog
  • Latent Action Learning Requires Supervision: "Latent Action Learning Requires Supervision in the Presence of Distractors". OpenReview Website

  • Beyond Experience: "Beyond Experience: Fictive Learning as an Inherent Advantage of World Models". OpenReview Website

  • Robotic World Model: "Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics". OpenReview Website

  • Sim-to-Real Contact-Rich Pivoting: "Sim-to-Real Contact-Rich Pivoting via Optimization-Guided RL with Vision and Touch". OpenReview [![Websit