:sunglasses: Awesome 3D and 4D World Models
![]() |
|---|
This survey reviews state-of-the-art 3D and 4D world models - systems that learn, predict, and simulate the geometry and dynamics of real environments from multi-modal signals.
We unify terminology, scope, and evaluations, and organize the space into three complementary paradigms by representation:
| Learn generative or predictive models from sequential video streams with geometric and temporal constraints. VideoGen focuses on long-horizon consistency, controllability, and scene-level generation, enabling agents to imagine or forecast plausible video rollouts. | |
| Model 3D/4D occupancy grids that encode geometry and semantics in voxel space. OccGen provides a physics-consistent scaffold for robust perception, forecasting, and simulation, bridging low-level sensor data and high-level reasoning. | |
| Leverage point cloud sequences from LiDAR sensors to generate or predict geometry-grounded scenes. LiDARGen emphasizes high-fidelity 3D structure, robustness to environment changes, and applications in safety-critical domains such as autonomous driving. | |
For more details, kindly refer to our paper and project page. :rocket:
:books: Citation
If you find this work helpful for your research, please kindly consider citing our papers:
@article{survey_3d_4d_world_models,
title = {{3D} and {4D} World Modeling: A Survey},
author = {Lingdong Kong and Yu Yang and Jianbiao Mei and Youquan Liu and Ao Liang and Dekai Zhu and Dongyue Lu and Wei Yin and Xiaotao Hu and Mingkai Jia and Junyuan Deng and Kaiwen Zhang and Yang Wu and Tianyi Yan and Shenyuan Gao and Song Wang and Linfeng Li and Liang Pan and Yong Liu and Jianke Zhu and Wei Tsang Ooi and Steven C. H. Hoi and Ziwei Liu},
journal = {arXiv preprint arXiv:2509.07996},
year = {2025}
}
@inproceedings{worldlens,
title = {{WorldLens}: Full-Spectrum Evaluations of Driving World Models in Real World},
author = {Ao Liang and Lingdong Kong and Tianyi Yan and Hongsi Liu and Wesley Yang and Ziqi Huang and Wei Yin and Jialong Zuo and Yixuan Hu and Dekai Zhu and Dongyue Lu and Youquan Liu and Guangfeng Jiang and Linfeng Li and Xiangtai Li and Long Zhuo and Lai Xing Ng and Benoit R. Cottereau and Changxin Gao and Liang Pan and Wei Tsang Ooi and Ziwei Liu},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
pages = {36385-36399},
year = {2026}
}
Table of Contents
- Background
- 1. Benchmarks & Datasets
- 2. World Modeling from Video Generation
- 3. World Modeling from Occupancy Generation
- 4. World Modeling from LiDAR Generation
- 5. Applications
- 6. Other Resources
- 7. Acknowledgements
Background
![]() |
World modeling has become a cornerstone of modern AI, enabling agents to understand, represent, and predict dynamic environments. While prior research has focused primarily on 2D images and videos, the rapid emergence of native 3D and 4D representations (e.g., RGB-D, occupancy grids, LiDAR point clouds) calls for a dedicated study. |
What Are Native 3D Representations?
Unlike 2D projections, native 3D/4D signals directly encode metric geometry, visibility, and motion in the physical coordinates where agents act. Examples include:
- RGB-D imagery (2D images with depth channels)
- Occupancy grids (voxelized maps of free vs. occupied space)
- LiDAR point clouds (3D coordinates from active sensing)
- Neural fields (e.g., NeRF, Gaussian Splatting)
What Are World Models in 3D and 4D?
A 3D/4D world model is an internal representation that allows an agent to imagine, forecast, and interact with its environment in the 3D space.
| Generative World Models: | |
| synthesize plausible 3D/4D worlds under conditions (e.g., text prompts, trajectories). | |
| Predictive World Models: | |
| anticipate the future evolution of 3D/4D scenes given past observations and actions. | |
Together, these models provide the foundation for simulation, planning, and embodied intelligence in complex environments.
![]() |
|---|
1. Benchmarks & Datasets
Benchmarks
![]() |
![]() |
![]() |
|---|---|---|
| WorldLens | VBench | WorldScore |
Workshops
| Theme | Venue | Date | Location | Recording |
|---|---|---|---|---|
| Workshop on 4D World Models: Bridging Generation and Reconstruction | CVPR 2026 | TBD | Denver | - |
| The 2nd Workshop on World Models | ICLR 2026 | April 23, 2026 | Rio de Janeiro | - |
| Workshop on World Modeling | - | February 4-6, 2026 | Montréal | - |
| Workshop on Embodied World Models for Decision Making | NeurIPS 2025 | December 6, 2025 | San Diego | - |
| Workshop on Reliable and Interactable World Models: Geometry, Physics, Interactivity and Real-World Generalization | ICCV 2025 | October 19, 2025 | Hawai'i | - |
| Workshop on Building Physically Plausible World Models | ICML 2025 | July 19, 2025 | Vancouver | - |
| Workshop on Assessing World Models | ICML 2025 | July 18, 2025 | Vancouver | - |
| Workshop on Benchmarking World Models | CVPR 2025 | June 12, 2025 | Nashville | - |
| Workshop on World Models: Understanding, Modelling and Scaling | ICLR 2025 | April 28, 2025 | Singapore | - |
| Workshop on Foundation Models for Autonomous Systems | CVPR 2024 | June 17, 2025 | Seattle | [YouTube] |
Datasets
:timer_clock: In chronological order, from the earliest to the latest.
2. World Modeling from Video Generation
:one: Data Engines
:timer_clock: In chronological order, from the earliest to the latest.
| Model | Paper | Venue | Website | GitHub |
|---|---|---|---|---|
HelloWorld (Driving) |
||||
| HelloWorld: Towards Practical Applications of Generative Driving World Models | arXiv 2026 | Project | - | |
4Director |
||||
| 4Director: Controlling Video World Models with Rigid 3D Geometry | arXiv 2026 | Project | - | |
MIVIFI |
||||
| MIVIFI: Bridging Perspective and Fisheye Domains for Training Multi-View Fisheye Image Generation Models | arXiv 2026 | - | - | |
SpatialCrafter |
||||
| SpatialCrafter: Single Image World Modeling with Generative 3D Proxies | arXiv 2026 | Project | - | |
4DStreamCtrl |
||||
| 4DStreamCtrl: Interactive Video Generation with Online 4D Control | arXiv 2026 | Project | - | |
ReWorld |
||||
| ReWorld: An Interactive World Model with Long-Horizon Memory | arXiv 2026 | - | - | |
NeoWorld-Pro |
||||
| NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation | arXiv 2026 | - | - | |
MiniWorld |
||||
| MiniWorld: Democratizing the Training of Video World Models from Scratch | arXiv 2026 | - | - | |
RealWeather |
||||
| RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models | arXiv 2026 | - | - | |
muSync |
||||
| muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards | arXiv 2026 | - | - | |
ABot |
||||
| ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space | arXiv 2026 | - | - | |
Sekai2 |
||||
| Sekai2: From World Exploration to Interactive World Modeling | arXiv 2026 | - | - | |
Xiaomi |
||||
| Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model | arXiv 2026 | - | - | |
EmbodiedVAE |
||||
| EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation | arXiv 2026 | - | - | |
H2R |
||||
| H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models | arXiv 2026 | - | - | |
BEVControl |
||||
| BEVControl: Accurately Controlling Street-View Elements with Multi-Perspective Consistency via BEV Sketch Layout | arXiv 2023 | - | - | |
BEVGen |
||||
| Street-View Image Generation from a Bird's-Eye View Layout | RA-L 2024 | |||
MagicDrive |
||||
| MagicDrive: Street View Generation with Diverse 3D Geometry Control | ICLR 2024 | |||
Panacea |
||||
| Panacea: Panoramic and Controllable Video Generation for Autonomous Driving | CVPR 2024 | |||
DrivingDiffusion |
||||
| DrivingDiffusion: Layout-Guided Multi-View Driving Scene Video Generation with Latent Diffusion Model | ECCV 2024 | |||
WoVoGen |
||||
| WoVoGen: World Volume-Aware Diffusion for Controllable Multi-Camera Driving Scene Generation | ECCV 2024 | - | ||
Delphi |
||||
| Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation | arXiv 2024 | |||
SimGen |
||||
| SimGen: Simulator-Conditioned Driving Scene Generation | NeurIPS 2024 | |||
BEVWorld |
||||
| BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents | arXiv 2024 | - | - | |
Panacea+ |
||||
| Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving | arXiv 2024 | - | ||
DiVE |
||||
| DiVE: DiT-Based Video Generation with Enhanced Control | arXiv 2024 | |||
MyGo |
||||
| MyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control | arXiv 2024 | - | - | |
SyntheOcc |
||||
| SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs | arXiv 2024 | |||
HoloDrive |
||||
| HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving | arXiv 2024 | - | - | |
CogDriving |
||||
| Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention | arXiv 2024 | - | ||
UniMLVG |
||||
| UniMLVG: Unified Framework for Multi-View Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving | arXiv 2024 | - | ||
DrivePhysica |
||||
| Physical Informed Driving World Model | arXiv 2024 | - | ||
DriveDreamer-2 |
||||
| DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation | AAAI 2025 | |||
SubjectDrive |
||||
| SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control | AAAI 2025 | - | ||
Glad |
||||
| Glad: A Streaming Scene Generator for Autonomous Driving | ICLR 2025 | - | ||
DualDiff |
||||
| DualDiff: Dual-Branch Diffusion Model for Autonomous Driving with Semantic Fusion | ICRA 2025 | - | ||
UniScene |
||||
| UniScene: Unified Occupancy-Centric Driving Scene Generation | CVPR 2025 | |||
DriveScape |
||||
| DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation | CVPR 2025 | - | ||
PerLDiff |
||||
| PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Models | ICCV 2025 | |||
MagicDrive-V2 |
||||
| MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control | ICCV 2025 | - | ||
DINO-Foresight |
||||
| DINO-Foresight: Looking into the Future with DINO | NeurIPS 2025 | |||
Cosmos-Transfer1 |
||||
| Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control | arXiv 2025 | |||
DualDiff+ |
||||
| DualDiff+: Dual-Branch Diffusion for High-Fidelity Video Generation with Reward Guidance | arXiv 2025 | - | ||
CoGen |
||||
| CoGen: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Driving | arXiv 2025 | - | ||
NoiseController |
||||
| NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration | arXiv 2025 | - | - | |
STAGE |
||||
| STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation | arXiv 2025 | - | - | |
WMReward |
||||
| Inference-time Physics Alignment of Video Generative Models with Latent World Models | arXiv 2026 | - | - | |
AutoScape |
||||
| AutoScape: Geometry-Consistent Long-Horizon Scene Generation | arXiv 2025 | - | - | |
Rethinking-DWM |
||||
| Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks | arXiv 2025 | - | - | |
OmniDrive |
||||
| OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation | arXiv 2026 | - | - | |
OpenLongTail |
||||
| OpenLongTail: Generative Scaling of Long-Tail Driving Data | arXiv 2026 | - | - | |
OmniSCS |
||||
| OmniSCS: Omni Safety-Critical Scenario Synthesis for Autonomous Driving via a Fully Editable Driving World | arXiv 2026 | - | - | |
DriveWeaver |
||||
| DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation | arXiv 2026 | - | - | |
:two: Action Interpreters
:timer_clock: In chronological order, from the earliest to the latest.
| Model | Paper | Venue | Website | GitHub |
|---|---|---|---|---|
CoDrive |
||||
| CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving | arXiv 2026 | Project | - | |
SA-WAM |
||||
| Spatially Aware World Action Model via Geometric Latent Diffusion | arXiv 2026 | Project | - | |
UniDynamics |
||||
| UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation | arXiv 2026 | - | - | |
GeoWAM |
||||
| GeoWAM: Visual Geometry World Action Models for Autonomous Driving | arXiv 2026 | - | - | |
4DGS-WAM |
||||
| 4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting | arXiv 2026 | - | - | |
WALL-SS |
||||
| WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression | arXiv 2026 | - | - | |
Cycle |
||||
| Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency | arXiv 2026 | - | - | |
Adaptive |
||||
| Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features | arXiv 2026 | - | - | |
ForgeWM |
||||
| ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models | arXiv 2026 | - | - | |
DriveCache |
||||
| DriveCache: Action-Aware Caching for Driving World Model Inference | arXiv 2026 | - | - | |
ContactFlow |
||||
| ContactFlow: A video action conditioning that transfers across embodiments | arXiv 2026 | - | - | |
GeniWorld |
||||
| GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions | arXiv 2026 | - | - | |
DreamX |
||||
| DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation | arXiv 2026 | - | - | |
GAIA-1 |
||||
| GAIA-1: A Generative World Model for Autonomous Driving | arXiv 2023 | - | ||
ADriver-I |
||||
| ADriver-I: A General World Model for Autonomous Driving | arXiv 2023 | - | - | |
Drive-WM |
||||
| Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving | CVPR 2024 | |||
DriveDreamer |
||||
| DriveDreamer: Towards Real-World-Driven World Models for Autonomous Driving | ECCV 2024 | |||
GenAD |
||||
| GenAD: Generalized Predictive Model for Autonomous Driving | CVPR 2024 | - | ||
GenAD (Gen-E2E) |
||||
| GenAD: Generative End-to-End Autonomous Driving | ECCV 2024 | - | - | |
Vista |
||||
| Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability | NeurIPS 2024 | |||
InfinityDrive |
||||
| InfinityDrive: Breaking Time Limits in Driving World Models | arXiv 2024 | - | ||
DrivingGPT |
||||
| DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive Transformers | arXiv 2024 | - | ||
DrivingWorld |
||||
| DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT | arXiv 2024 | |||
GEM |
||||
| GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control | CVPR 2025 | |||
MaskGWM |
||||
| MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction | CVPR 2025 | - | ||
Epona |
||||
| Epona: Autoregressive Diffusion World Model for Autonomous Driving | ICCV 2025 | |||
VaViM & VaVAM |
||||
| VaViM and VaVAM: Autonomous Driving through Video Generative Modeling | arXiv 2025 | |||
MiLA |
||||
| MiLA: Multi-View Intensive-Fidelity Long-Term Video Generation World Model for Autonomous Driving | arXiv 2025 | - | ||
GAIA-2 |
||||
| GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving | arXiv 2025 | - | ||
Ego-Other-WM |
||||
| Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space | arXiv 2025 | - | - | |
DriVerse |
||||
| DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment | arXiv 2025 | - | - | |
PosePilot |
||||
| PosePilot: Steering Camera Pose for Generative World Models with Self-Supervised Depth | arXiv 2025 | - | - | |
ProphetDWM |
||||
| ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos | arXiv 2025 | - | - | |
LongDWM |
||||
| LongDWM: Cross-Granularity Distillation for Building A Long-Term Driving World Model | arXiv 2025 | |||
UniDrive-WM |
||||
| UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving | arXiv 2026 | - | ||
DriveVA |
||||
| DriveVA: Video Action Models are Zero-Shot Drivers | ECCV 2026 | - | ||
DeepSight |
||||
| DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving | arXiv 2026 | - | - | |
:three: Neural Simulators
:timer_clock: In chronological order, from the earliest to the latest.
:four: Scene Reconstructors
:timer_clock: In chronological order, from the earliest to the latest.
3. World Modeling from Occupancy Generation
:one: Scene Representors
:timer_clock: In chronological order, from the earliest to the latest.
:two: Occupancy Forecasters
:timer_clock: In chronological order, from the earliest to the latest.
:three: Autoregressive Simulators
:timer_clock: In chronological order, from the earliest to the latest.
4. World Modeling from LiDAR Generation
:one: Data Engines
:timer_clock: In chronological order, from the earliest to the latest.
| Model | Paper | Venue | Website | GitHub |
|---|---|---|---|---|
LiDAR4D-Anno |
||||
| Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models | arXiv 2026 | - | - | |
DUSty |
||||
| Learning to Drop Points for LiDAR Scan Synthesis | IROS 2021 | |||
LiDARGen |
||||
| Learning to Generate Realistic LiDAR Point Clouds | ECCV 2022 | - | ||
DUSty v2 |
||||
| Generative Range Imaging for Learning Scene Priors of 3D LiDAR Data | WACV 2023 | |||
UltraLiDAR |
||||
| UltraLiDAR: Learning Compact Representations for LiDAR Completion and Generation | CVPR 2023 | - | ||
Copilot4D |
||||
| Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion | ICLR 2024 | - | ||
R2DM |
||||
| LiDAR Data Synthesis with Denoising Diffusion Probabilistic Models | ICRA 2024 | |||
ViDAR |
||||
| Visual Point Cloud Forecasting enables Scalable Autonomous Driving | CVPR 2024 | - | ||
LiDiff |
||||
| Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion | CVPR 2024 | - | ||
LiDM |
||||
| Towards Realistic Scene Generation with LiDAR Diffusion Models | CVPR 2024 | - | ||
RangeLDM |
||||
| RangeLDM: Fast Realistic LiDAR Point Cloud Generation | ECCV 2024 | - | ||
Text2LiDAR |
||||
| Text2LiDAR: Text-Guided LiDAR Point Cloud Generation via Equirectangular Transformer | ECCV 2024 | - | ||
LiDARGRIT |
||||
| Taming Transformers for Realistic LiDAR Point Cloud Generation | arXiv 2024 | - | ||
BEVWorld |
||||
| BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents | arXiv 2024 | - | ||
SDS |
||||
| Simultaneous Diffusion Sampling for Conditional LiDAR Generation | arXiv 2024 | - | - | |
DiffSSC |
||||
| DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic Models | IROS 2025 | - | - | |
HoloDrive |
||||
| HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving | arXiv 2024 | - | - | |
LOGen |
||||
| LOGen: Toward LiDAR Object Generation by Point Diffusion | arXiv 2024 | [ |





