← 开源
Thinklab-SJTU

Awesome-LLM4AD

A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually updated)

ListsPaper collectionsDomain-specific
在 GitHub 打开
增长势头
+-124 小时新增 Star-0.1%
1.89k
Star
113
Fork
+0
本周
6
贡献者
创建于 2023-10-06 · 更新于 2026-10-06 · 今日第 18315 名
主要开发者
README

Awesome-LLM-for-Autonomous-Driving-Resources

AwesomeGitHub stars GitHub forks GitHub license

This is a collection of research papers about LLM-for-Autonomous-Driving(LLM4AD). The repository will be continuously updated to track the frontier of LLM4AD (Large Language Models for Autonomous Driving), which encompasses VLM4AD (Vision-Language Models for AD) and VLA4AD (Vision-Language-Action models for AD) as integral components of this unified paradigm. Maintained by SJTU-ReThinklab.

Welcome to follow and star! If you find any related materials could be helpful, feel free to contact us ([email protected] or [email protected]) or make a PR.

Citation

Our survey paper is at https://arxiv.org/abs/2311.01043 which includes more detailed discussions and will be continuously updated.

If you find our repo is helpful, please consider cite it.

@misc{yang2023survey,
      title={LLM4Drive: A Survey of Large Language Models for Autonomous Driving}, 
      author={Zhenjie Yang and Xiaosong Jia and Hongyang Li and Junchi Yan},
      year={2023},
      eprint={2311.01043},
      archivePrefix={arXiv},
      primaryClass={cs.AI}
}

Table of Contents

Overview of LLM4AD

LLM-for-Autonomous-Driving (LLM4AD) refers to the application of Large Language Models(LLMs) in autonomous driving. We divide existing works based on the perspective of applying LLMs: planning, perception, question answering, and generation.

image info

Motivation of LLM4AD

The orange circle represents the ideal level of driving competence, akin to that possessed by an experienced human driver. There are two main methods to acquire such proficiency: one, through learning-based techniques within simulated environments; and two, by learning from offline data through similar methodologies. It’s important to note that due to discrepancies between simulations and the real-world, these two domains are not fully the same, i.e. sim2real gap. Concurrently, offline data serves as a subset of real-world data since it’s collected directly from actual surroundings. However, it is difficult to fully cover the distribution as well due to the notorious long-tailed nature of autonomous driving tasks. The final goal of autonomous driving is to elevate driving abilities from a basic green stage to a more advanced blue level through extensive data collection and deep learning.

image info

Papers

Toggle

format:
- [title](paper link) [links]
  - author1, author2, and author3...
  - publisher
  - task
  - keyword
  - code or project page
  - datasets or environment or simulator
  - publish date
  - summary
  - metrics
  • RefineDrive: Reliable Failure-Guided Learning for Vision-Language-Action Driving

    • Zhe Sun, Ziyi Luo, Yehao Lu, Lei Zhou, Xi Li
    • Publish Date: 2026.09.28
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • RefineDrive is a failure-guided post-training framework for Vision-Language-Action autonomous driving that learns from self-generated failures through reliable diagnosis, minimum-correction target retrieval, and safety-aware reinforcement learning.
      • It derives structured, verifiable feedback on collisions and drivable-area violations from simulator states, retrieves hard-safety-constrained corrections from a clustered human trajectory bank, and uses correction SFT followed by GRPO with a Safety-Layered Reward.
      • On NAVSIM v1, RefineDrive improves a 4B base SFT policy from 87.7 to 91.7 PDMS and achieves 89.4 EPDMS on original NAVTEST scenes under NAVSIM v2 extended metrics without additional training.
  • CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving

    • Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Shi Fan, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue
    • Publish Date: 2026.09.28
    • Code: CAR-VLA
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • CAR-VLA is a unified driving VLA model that jointly considers scene complexity and dynamic risk to guide reasoning depth, urgency, and focus.
      • It maps four complexity-risk categories to three reasoning modes: Fast Intuition, Slow Thinking, and Reflex Response, with high-risk scenarios centered on the most critical hazard and immediate safe response.
      • Trained via progressive supervised learning linking scene assessment, reasoning-mode selection, and trajectory generation, followed by reasoning-augmented reinforcement learning, achieving competitive results on NAVSIM v1, NAVSIM v2, and Navhard.
  • AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

    • Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska
    • Publish Date: 2026.09.28
    • Code: AD-E2E-JEPA
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • Systematically evaluates existing action-conditioned JEPA world models (LeWM, DINO-WM, JEPA-WM) for end-to-end autonomous driving using a goal-conditioned zero-shot planning setting.
      • Proposes AD-E2E-JEPA with a SIGReg-regularized learnable projector that reduces planning patches by 16× and embedding dimension by 4×, achieving a 100× inference speedup while retaining planning performance.
      • Achieves 67.3/72.9 EPDMS on NAVSIMv2 in goal-conditioned zero-shot planning, and the self-supervised pretrained projector improves downstream imitation learning from 80.2 to 85.4 EPDMS.
  • MomWorld: Momentum-Aware Latent World Model for Long-Horizon Autonomous Driving

    • Ziying Song, Shengkai Zhang, Lei Yang, Haozhuang Chi, Yuchen Liu, Jiangtao Su, Lin Liu, Ziyang Liu, Chen Lv
    • Publish Date: 2026.09.27
    • Task: Planning
    • Datasets: NAVSIM, nuScenes, Bench2Drive
    • Summary:
      • MomWorld is a momentum-aware latent world model for long-horizon planning that extracts scene motion trends from historical-to-current observations and propagates latent momentum into future horizons.
      • It jointly predicts future configuration and momentum states with a learnable momentum persistence mechanism, scene-conditioned momentum updates, and a scene-adaptive reset gate to suppress stale momentum under abrupt changes.
      • It introduces MoFlow, a momentum-conditioned flow-matching module with horizon-aware residual fusion, improving long-horizon planning consistency and reducing the average collision rate by 12.2% relative to MomAD over a 6-second planning horizon on NAVSIM, nuScenes, and Bench2Drive.
  • VehDyn: A Driving World Model Benchmark for Vehicle Dynamics

    • Tianyi Wang, Wangsheng Du, Jiazhou Chen, Tianyi Zeng, Xiangyu Li, Jiseop Byeon, Yujin Wang, Yiming Xu, Yangyang Wang, Bingzhao Gao, Sikai Chen, Zhaomiao Guo, Junfeng Jiao, Christian Claudel, Alexandre Bayen
    • Publish Date: 2026.09.27
    • Task: Evaluation
    • Summary:
      • VehDyn is a driving world model benchmark for vehicle dynamics, built on a CARLA-CarSim co-simulation platform that couples photorealistic rendering with a validated multi-body dynamics model.
      • It contains 10,080 configurations from a full factorial design over vehicle types, tire-road friction coefficients, maneuvers, target speeds, scenes, and illuminations, each paired with synchronized position, velocity, and attitude sequences.
      • It introduces a hierarchical evaluation framework for trajectory alignment, kinematic consistency, and dynamic consistency, benchmarks 12 state-of-the-art video world models, and shows that trajectory metrics are nearly saturated while dynamic consistency remains challenging and visual quality weakly correlates with vehicle-dynamics fidelity.
  • RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving

    • Lianqing Zheng, Xiaokai Bai, Yixuan Luo, Runwei Guan, Minghao Liu, Zhiqiang Wei, Hui-liang Shen, Xichan Zhu, Zhixiong Ma
    • Publish Date: 2026.09.26
    • Task: Planning
    • Summary:
      • RCVLA is a radar-camera vision-language-action framework with a radar-grounded semantic reasoning stage (RCVLA-Sem) and a trajectory arbitration stage (RCVLA-Phys).
      • It constructs Cap4DR with 86,016 radar-image-text samples for alignment pretraining and OmniHD-QA with 520,161 question-answer pairs for instruction tuning across scene description, key-object reasoning, occupancy understanding, and trajectory planning.
      • RCVLA-Sem improves CIDEr by 9.92 points and reduces key-object velocity error by 21.9% over OmniDrive; RCVLA-Phys reduces average L2 error from 0.348 to 0.259 m and average open-loop collision rate from 0.576% to 0.175%.
  • From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving

    • Yiyao Wang, Pei Liu, Fangzhou Liu, Jun Ma
    • Publish Date: 2026.09.26
    • Task: VQA
    • Datasets: NuScenes-QA, nuScenes
    • Summary:
      • Introduces a query-adaptive neuro-symbolic reasoning framework for autonomous-driving question answering, with a hierarchical Spatiotemporal Scene Graph (STSG) that separates persistent object identities from frame-specific states and represents spatial relations and temporal transitions as explicit directed structures.
      • Uses a symbolic executor to resolve queries through exact graph operations and invokes an LLM only when symbolic execution abstains, with query-conditioned graph retrieval and evidence filtering preserving relation direction, temporal locality, and object semantics.
      • Achieves 80.63% overall accuracy on 5,916 NuScenes-QA questions under oracle perception, improving over the LLM-only configuration by 5.48 points with GPT-5.4-mini and 6.64 points with DeepSeek-V4-Flash, with largest gains on counting questions.
  • OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes

    • Boseong Jeon, Junhyeop Lee, Juhan Cha, Hayoung Kim
    • Publish Date: 2026.09.26
    • Project Page: OneFixer
    • Task: Perception
    • Datasets: Waymo
    • Summary:
      • OneFixer is a one-step autoregressive video-diffusion fixer for 3D Gaussian Splatting in driving scenes, trained in a single task-specific adaptation stage.
      • It uses a deployment-matched shared rollout where one-step predictions serve as causal context for flow matching while receiving direct pixel-space perceptual supervision, jointly learning fidelity and autoregressive robustness.
      • It exploits lane geometry and dynamic-agent states to improve geometric fidelity, achieving the lowest FVD, LPIPS, and DISTS at one step on Waymo and proprietary driving scenes, and reducing collision rate by one third in closed-loop simulation.
  • CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving

    • Narendiran Chembu, Navvrat Rao, Shreedhar Shreeshail Kodate, Gayatri Srujana Banda, Arko Sarkar, Abhinav Khanna, Rajarshee Das, Umesh Kanala, Siddarth Khandelwal, Kumar Aman, Aish Dubey, Kaustubh Beedkar, Arjun Jain
    • Publish Date: 2026.09.26
    • Task: Reasoning
    • Datasets: nuScenes
    • Summary:
      • CausalDriveBench is an evaluation framework grounded in Pearl's Causal Hierarchy that tests causal reasoning in driving-specific VLAs through structured visual question answering and alternative-trajectory prediction.
      • It constructs causal scene graphs over nuScenes to distinguish causally active, dormant, and distractor entities, covering all four rungs of PCH and containing 7,285 verified causal QA pairs and 1,000 counterfactual trajectories.
      • Evaluating 10 driving-specific VLAs and 3 general-purpose VLMs shows that causal QA accuracy remains limited, driving fine-tuning incurs variable costs on causal QA, and causal QA and trajectory accuracy are statistically uncorrelated across models.
  • DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution

    • Chengkai Xu, Jiaqi Liu, Yicheng Guo, Peng Hang, Jian Sun
    • Publish Date: 2026.09.25
    • Code: DriveHierarchy
    • Task: Evaluation
    • Summary:
      • DriveHierarchy is a hierarchical benchmark that organizes VLM-based autonomous driving into four ranks: perceptual grounding, contextual memory, mental reasoning, and closed-loop execution.
      • It integrates multiple open-source autonomous-driving datasets into a unified open-loop benchmark with 76,798 question-answer pairs over 84,279 frames, and develops a closed-loop simulation platform with 100 interactive scenarios.
      • Experiments on 15 VLMs show that DriveHierarchy captures structured non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization.
  • Evaluation Is All You Need for Multi-Modal Autonomous Driving

    • Zeyu He, Shiqi Liu, Ke Chen, Yun Yan, Jinzi Wu, Dianqiao Lei, Sirui Wang, ShuRui Peng, Tao Chen, Zhuo Huang, Yu Wu, Yadong Shao, Zhichao Li, Ke Sun, Yang Guan, Keqiang Li, Shengbo Eben Li
    • Publish Date: 2026.09.25
    • Task: Evaluation
    • Datasets: NAVSIM
    • Summary:
      • iDriveVLA is a multi-modal planning framework that improves the candidate trajectory space while enabling reliable and context-aware trajectory evaluation.
      • It introduces a unified trajectory evaluator with a Safety-aware Scorer for quality and risk estimation and a VLM-guided Modulator for scene-adaptive criterion weighting.
      • An oracle-aligned progressive training strategy combines candidate imitation pretraining, candidate space refinement, and semantic ranking alignment, achieving 94.95 PDMS on NAVSIM v1 and surpassing human-expert reference.
  • WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

    • Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
    • Publish Date: 2026.09.24
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • WALT learns a compact generative trajectory latent space by transferring information from a frozen pretrained driving world model, aligning raw waypoints with action-relevant semantics without modifying the world model.
      • It uses a dual-branch trajectory autoencoder and systematically studies JEPA-based latent learning and REPA-style feature alignment for trajectory-only representation learning.
      • On NAVSIM, WALT improves PDMS from 89.4 to 89.8 on NAVSIMv1 and EPDMS from 87.3 to 87.9 on NAVSIMv2 while reducing trajectory planner FLOPs by 30.5%.
  • Sim-to-Real Aware End-to-End Learning Environment for Micromobility

    • Shouma Amano, Takuya Azumi
    • Publish Date: 2026.09.24
    • Task: End-to-End
    • Summary:
      • Proposes a sim-to-real-aware, vehicle-specific end-to-end learning environment for the WHILL Model CR on AWSIM and ROS 2.
      • Optimizes physical parameters via Bayesian optimization using real-world data to reduce the sim-to-real gap and trajectory errors across various driving scenarios.
      • Introduces a synchronized architecture for stable training of world model-based agents, with a DreamerV3 policy achieving obstacle avoidance and direct sim-to-real transfer to the physical vehicle without fine-tuning.
  • Less Language, More Latents: Annotation-Efficient VLAs for Driving

    • Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania
    • Publish Date: 2026.09.23
    • Task: End-to-End
    • Datasets: Bench2Drive
    • Summary:
      • Introduces Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into training data for language-conditioned control.
      • Uses a latent action model with a vector-quantised bottleneck to build a compact codebook of high-level vehicle intents, and trains a vision-language translator on a small language-annotated subset.
      • Trains a driving VLA on observation-latent-action pairs over the full unlabelled corpus, achieving 87.98 Driving Score and 70.46% Success Rate on Bench2Drive with fewer than 5% language annotations.
  • ForeDrive: Foresight-Guided End-to-End Autonomous Driving with a Planning-Relevant Latent World Model

    • Sinuo Wang, Zichong Gu, Yuhan Huang, Wenxin Wen, Xun Yang, Yiqing Zhang, Xingyu Zhang, Ningyu Che, Jie Ling, Qiankun Yu, Wei Liu, Jing Xu, Xinggang Wang
    • Publish Date: 2026.09.22
    • Task: End-to-End
    • Datasets: NAVSIM
    • Summary:
      • ForeDrive learns a planning-relevant latent future representation with a JEPA-style world model and couples it asymmetrically to a Diffusion Transformer (DiT) planner, so that predicted futures act as direct conditioning signals for trajectory generation rather than mere pretraining or auxiliary supervision.
      • The framework uses gated visual fusion, future-status injection, and Trajectory-Adaptive Bias (TAB) to inject future latents as guidance without overriding the current observation, addressing varying future reliability across horizons and the misalignment between BEV trajectories and image tokens.
      • Trained with pure imitation learning and using only the current front-view image as visual input at inference, ForeDrive attains 89.9 PDMS on NAVSIM v1 and 90.0 one-stage EPDMS on NAVSIM v2, without reinforcement learning or an external trajectory scorer.
  • Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving

    • Yuqi Ye, Shangkun Sun, Junhong Lin, Jiayi Zhao, Changhao Peng, Wei Zheng, Guoqing Liu, Tiesong Zhao, Wei Gao
    • Publish Date: 2026.09.22
    • Task: Planning
    • Datasets: NAVSIM, nuScenes
    • Summary:
      • Run-then-Walk is a simple yet effective two-stage reward scheduling strategy for GRPO that explicitly separates progress discovery from safety repair, achieving both better performance and faster convergence.
      • The Run phase focuses on progress to escape conservative bias and discover high-progress modes, while the subsequent Walk phase introduces endpoint and safety strategies to repair unsafe behaviors, overcoming the conservatism of Walk-first methods and the unsafe progress-seeking of joint optimization.
      • Validated on NAVSIMv1, NAVSIMv2, Navhard, and nuScenes with various VLM-based planners, demonstrating improved driving performance while requiring 40--50% fewer RL training epochs than baselines.
  • Relationally Grounded Latent World Models for Autonomous Driving

    • Fabian Schmidt, Markus Enzweiler, Abhinav Valada
    • Publish Date: 2026.09.21
    • Task: Prediction
    • Datasets: nuScenes
    • Summary:
      • Investigates whether traffic scene graphs can serve as privileged semantic supervision for latent world representations in autonomous driving.
      • Constructs actor-centric scene graphs from nuScenes 3D annotations, encodes their serialized relational structure with a frozen text embedding model, and aligns visual latent representations with this semantic target during training.
      • Removes the supervision branch at inference and, on nuScenes, reduces average trajectory L2 error from 0.661 to 0.622 and collision rate from 0.456 to 0.217, outperforming an unstructured caption-style semantic target.
  • PRIME: Perception Feedback with Situational Memory Embeddings in VLA Models

    • Erik Deinzer, Naya Baslan, Luca Paparusso, Narunas Vaskevicius, Peter Knott, Luigi Palmieri
    • Publish Date: 2026.09.18
    • Task: End-to-End
    • Datasets: Bench2Drive, Think2Drive
    • Summary:
      • PRIME introduces a learned feedback mechanism that conditions VLA perceptual queries on a novel Situational Memory, enabling intent-driven perceptual attention at minimal computational cost.
      • It aggregates latent representations of past perception, reasoning, navigation goals, and predicted behaviors across an L-step window via cross-attention, adding at most 29.7M parameters (0.41% of the 7.3B-parameter base model).
      • On Bench2Drive, PRIME achieves a state-of-the-art Driving Score of 82.47 (+4.73 over ORION) and a Success Rate of 60.00% (+5.38 pp), the highest reported among published VLAs trained on Think2Drive demonstrations.
  • ZYT-World: A Real-Time Controllable World Model for Closed-Loop Autonomous-Driving Simulation

    • Boni Hu, Xiong Wei, Haoming Huang, Yong Huang, Chenbo Wang, Yi Yang, Jiancheng Wang, Ruicheng Zhu, Zhimin Yang, Guanglai Liu, Qiaowan Jin, Dongzhuo Wang, Haiwei Kuang, Jiajun Fan, Yue Wu, Jiaxin Wei, Hao Sun, Feihong Yan, Yuyao Zhou, Wei Bi, Kaixuan Wang, Zichao Guo, Xiaozhi Chen
    • Publish Date: 2026.09.18
    • Task: Generation
    • Summary:
      • ZYT-World is a real-time controllable world model that natively generates four fisheye views with field of view > 180° and three pinhole views for closed-loop autonomous-driving simulation.
      • It combines projection-specific Plücker adapters, ego-motion adaptive layer normalization, and a lightweight pixel-aligned layout to control global motion and instance-level traffic participants and signals.
      • A distillation pipeline converts a 40-step bidirectional teacher into a one-step per-latent streaming generator, while TinyVAE, W8A8 quantization, and RigCritic enable over 90% PSNR/SSIM retention and 107.7× faster generation.
  • Astronex-World 1.0: Real-Time Interactive World Model Foundation

    • Xin Zhou, Cong Miao
    • Publish Date: 2026.09.17
    • Task: Prediction
    • Summary:
      • Astronex-World 1.0 is an open controllable video world-model foundation that predicts future visual states from text prompts or initial observations under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at specified rollout positions.
      • The family includes a bidirectional full-context generation model and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior; PRoPE injects camera intrinsics/extrinsics while a 64-dimensional action stream modulates every Transformer layer.
      • A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching; the causal model generates 832x480 video at 24 fps and streams in real time on one NVIDIA L20, scoring 73.5 on WBench Navi and 70.0 on WBench Full.
  • REARL: A Closed-loop Autonomous Driving Simulation Enhancement Framework with Real Traffic Data and Large Language Models

    • Xiaojun Bi, Jun Jiang, Yiwen Sun, Quanyi Ou, Yizhi Ma, Ke Cheng, Mingjie Bi, Yexin Li, Bowen Du
    • Publish Date: 2026.09.17
    • Task: Planning
    • Datasets: HighD
    • Summary:
      • REARL is a closed-loop simulation enhancement framework that integrates real traffic data with LLMs, using clustered real traffic data as representative scenarios and a timed sliding-window detector to monitor speed distribution and mean spacing discrepancies.
      • When a metric exceeds a threshold, the LLM adjusts vehicle decision-making; otherwise the existing controller is retained, and the LLM also selects matching real vehicle actions from a traffic snapshot to modulate simulated vehicles.
      • In a controlled HighD highway setting, REARL reduces the Hellinger distance for speed distributions to 0.3067 and the MAPE for mean spacing to 0.8371, while achieving a time headway of 22.8575 and a lane change rate of 0.0708.
  • FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

    • Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao, Puneet K. Dokania
    • Publisher: Five AI, University of Oxford
    • Publish Date: 2026.09.16
    • Task: Planning
    • Datasets: Bench2Drive, NVIDIA Physical AI AV Dataset
    • Summary:
      • FIVE-VLA introduces an efficient vision encoder that processes 448×896 images with only 98 tokens (over 5× fewer than existing approaches) and bypasses text generation for single-pass trajectory prediction.
      • It proposes Recurrent Action Memory (RAM), a lightweight module that conditions action prediction on previous action tokens to provide temporal context for manoeuvres such as overtaking and emergency braking.
      • With 641M parameters, FIVE-VLA achieves ~10% more routes without traffic rule infractions than the previous state-of-the-art VLA on Bench2Drive, reduces collision-violation rates on the NVIDIA Physical AI AV dataset, and runs at ~30 fps on an A100 and ~4 fps on a T4 (8–30× speedup).
  • RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving

    • Dogun Kim, Yongjae Lee, Joonhee Lim, Yeina Lee, Junhyeok Park, Moogeun Park, Dongsuk Kum
    • Publish Date: 2026.09.15
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • RAF-VLA (Representation Alignment with the Future) is a VLA-based autonomous driving framework that aligns the policy's hidden states with future-frame representations from a pretrained world encoder, avoiding explicit future generation.
      • It employs Future-Aligned Supervised Fine-Tuning, a straightforward regularization that shapes planning-relevant internal representations while learning driving actions, eliminating the training burden and inference latency of world modeling.
      • Experiments on the NAVSIM benchmark show competitive planning performance against state-of-the-art VLA planners with substantially fewer training samples, while incurring only 3.8% training overhead and negligible 1 ms inference overhead.
  • EditWM: Event-Decomposed World Modeling with Incremental Correction for End-to-End Autonomous Driving

    • Junjie Yang, Qingwei Zeng, Youyou Li, Zicheng Ding, Ziyi Shi, Shuqi Shen, Hongliang Lu, Hai Yang
    • Publish Date: 2026.09.15
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • EditWM decomposes future prediction into normal evolution and event-driven incremental correction in compact visual feature space.
      • It freezes a trajectory-conditioned normal predictor and trains a correction decoder with a learned gate to produce bounded feature updates, using corrected future features for candidate-specific trajectory scoring.
      • Evaluated on all 12,146 NAVSIM navtest scenes, it reduces future-feature MSE by 5.35% over Normal with improvements in 83.54% of scenes and achieves 91.05 EPDMS.
  • GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving

    • Xiao Liu, Haoyu Li, Jianghao Leng, Lin Wang, Chao Sun
    • Publish Date: 2026.09.14
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • GRAVA is a framework built around Grounded Reasoning-to-Action (GRA), unifying grounding, reasoning, and action generation in a single autoregressive stream.
      • It introduces an agentic GRA data construction pipeline that combines forward scene grounding with backward trajectory anchoring to build GR-NavSim with 2.2M grounded question-answer pairs and 70K GRA reasoning traces.
      • A progressive training strategy pre-trains grounded cognition, establishes the reasoning-to-action interface through imitation, and improves driving behavior via reinforcement learning; GRAVA-8B achieves state-of-the-art performance among purely autoregressive driving models on the full NAVSIM benchmark and improves long-tail key-object compliance and Closed-loop Driving Score by 19.3% and 20.5% over action-only prediction.
  • READ: Learning Risk-Informed Fields for End-to-End Autonomous Driving

    • Zhiyuan Liu, Yuanxin Tian, Zehong Ke, Jinhao Li, Hao Cheng, Zhenhua Xu, Wenhao Yu, Jianqiang Wang
    • Publish Date: 2026.09.11
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • READ learns an explicit, planning-aligned risk representation from complementary geometric and behavioral constraints for end-to-end autonomous driving.
      • It instantiates this representation as a continuous spatiotemporal field, enabling differentiable queries along candidate trajectories for trajectory evaluation and refinement.
      • READ integrates with both end-to-end planners and Vision-Language-Action models, showing consistent gains on NAVSIM and competitive results on NAVSIM v2.
  • PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

    • Yuan Gao, Sebastian Müller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Schäfer, Qunying Song, Johannes Betz
    • Publish Date: 2026.09.08
    • Task: Planning
    • Summary:
      • PlannerForge is a unified LLM-agent framework that extends all scenario-based testing stages for ADSs, from Scenario Generation to ADS Assessment, and adds ADS Enhancement and ADS Benchmarking.
      • It evaluates 10 off-the-shelf LLMs across Generation, Selection, Modification, Module Routing, Planner Testing, and Enhancement under 5 prompt conditions, with best-per-task scores ranging from 0.88 to 1.00; open-source 20-35B models match commercial APIs on most tasks.
      • End-to-end chaining retains 83% / 78% of seed queries (commercial / open), outperforms Scenario Factory 2.0 on natural-language generation and BM25/From-Words-to-Collisions on selection and physically valid edits, and at N=400 cost-tuning lifts planner success from 50.4% to 70.2% while cutting collisions from 19.0% to 8.4%.
  • Drive by Hindsight and Foresight: Tool-Grounded Synergistic Reasoning over Hierarchical Memory for Autonomous Driving

    • Baojie Chen, Zijun Jia, Jing Zhong
    • Publish Date: 2026.09.08
    • Task: Reasoning
    • Summary:
      • Introduces a synergistic framework that tightly couples hierarchical memory with proactive tool invocation in a closed reasoning loop for autonomous driving, addressing hallucination, weak spatio-temporal perception, and limited generalization in VLMs.
      • Proposes Hierarchical Driving Memory with scene-level short-term memory for dynamic scene state and evolving long-term memory for reusable experience and tool strategies, plus a Memory-Tool Synergistic Reasoning Framework for adaptive tool invocation and offline consolidation.
      • Uses verified memory-tool trajectories from multi-step teacher rollout for SFT and GRPO training; the 7B model achieves 80.03 overall reasoning score and 79.09% MCQ accuracy on DriveLMM-o1, surpassing the strongest baseline by 7.74 MCQ points, with ablations validating short-term memory (+24.2 STSBench accuracy) and long-term memory consolidation (+3.57 MCQ with frozen parameters).
  • CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving

    • Zhe Huang, Zhaoxin Fan, Shuo Wang, Wenjun Wu, Xuan Zhao, Min Liu
    • Publish Date: 2026.09.04
    • Task: Planning
    • Datasets: CARLA
    • Summary:
      • CoLMIN is an LLM-based multi-decision path negotiation framework for cooperative autonomous driving, achieving stable decision consensus through multi-decision path negotiation and reflective reasoning.
      • It consists of three key components: an LLM-based Multi-Intent Negotiation module (LMin), an Evaluation-based Shallow Reflection Module (ESRM), and an LLM-based Deep Reflection Module (LDRM).
      • Experimental results in the CARLA simulation environment demonstrate that CoLMIN significantly outperforms existing methods in challenging interactive driving scenarios.
  • Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving

    • Ruoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu, Zewei Yang, Yipeng Zhu, Xiaolong Wang, Jun Ma
    • Publish Date: 2026.09.03
    • Task: Planning
    • Datasets: nuScenes
    • Summary:
      • LaPla is a unified Vision-Language-Action framework with latent-aligned planning that grounds semantic understanding in precise motion execution for end-to-end autonomous driving.
      • It uses a residual VQ-VAE action tokenizer and repurposes the representation as a physical prior, projecting hidden states into a pretrained VQ-VAE latent space to avoid quantization errors and bypass autoregressive generation.
      • LaPla reduces long-horizon L2 error by 15.52% on nuScenes and improves closed-loop success rate by 33.34 percentage points on NVIDIA AlpaSim with significantly reduced inference latency.
  • SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

    • Jinyang Wang, Shiwei Li, Junjian Wang, Zhiqiang Deng, Jianbin Gao, Yihang Zhao, Liu Liu, Yongjia Zhao, Jinlong Chen, Huirui Xu, Yifeng Pan, Kangwei Liu, Fan Ren, Ji Tao, Minghao Yang
    • Publish Date: 2026.09.03
    • Task: Planning
    • Datasets: NAVSIMv2, nuScenes
    • Summary:
      • SV-WAM is an efficient surround-view world-action model that preserves full six-camera observations while using future-video prediction as dense training supervision rather than as an inference-time output.
      • It introduces an action-centered causal mask for joint action-video denoising, enabling the video branch to be discarded at deployment for efficient action-only planning.
      • A differentiable drivable-area compliance regularizer improves planning safety and boundary awareness, and experiments on NAVSIMv2 and nuScenes show state-of-the-art planning with low inference latency and competitive zero-shot transfer.
  • Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

    • Zhaoxin Fan, Tianbao Zhang, Wenjun Wu, Xiaofeng Wang, Yeying Jin, Jian Zhao, Zheng Zhu, Shuicheng Yan
    • Publish Date: 2026.09.03
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • Drive-HWM is a hierarchical slow–fast world modeling framework for autonomous driving that organizes future representation prediction and action generation at complementary temporal scales.
      • It introduces Dynamic-Aware Latents learned through optical-flow prediction to model driving motion dynamics; the slow world model predicts multi-step future representations, while the fast model jointly predicts the next frame and immediate action guided by these representations.
      • Extensive experiments on NAVSIM v1 and v2 demonstrate strong driving performance, and ablations validate the hierarchical slow–fast design, dynamics-aware future representations, and joint next-frame and action prediction.
  • Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving

    • Yuxuan Han, Kunyuan Wu, Liyunong Yang, Zilu Wang, Cansen Jiang, Yi Xiao, Liang Hu
    • Publish Date: 2026.09.03
    • Task: End-to-End
    • Datasets: Bench2Drive
    • Summary:
      • StyleDrive, a world-model-based learning framework for end-to-end autonomous driving that jointly enforces long-horizon consistency, explicitly disentangles interactive traffic states, and supports multi-style policy optimization.
      • Introduces temporal consistency regularization with gated cross-attention, an explicit state disentanglement module separating ego-relevant from ego-irrelevant interactive states, and Group Relative Policy Optimization for multi-style driving without retraining.
      • Achieves 88.44 driving score (+17.08) and 66.82 success rate (+16.58) on Bench2Drive, and demonstrates promising sim-to-real transfer on a real automated guided vehicle platform.
  • From Proxy Learning to Driving Decisions: A Transfer-Based Framework for Evaluating Future-Aware Autonomous Driving Planners

    • Yikai Wu
    • Publish Date: 2026.09.02
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • Proposes the Proxy-to-Decision Transfer (PDT) Framework, an analysis framework for evaluating when learned future information supports a reliable driving-performance improvement claim.
      • Introduces a Decision-Transfer Decomposition Module for localizing value loss and a Reliability-Constrained Validation Module for requiring exact pairing, meaningful effect, scale-expanded confirmation, safety non-compensation, sequential comparability, and family-level robustness.
      • Evaluates a future-aware planner on NAVSIM-v1, showing proxy BCE improves while held selected PDM decreases, subset gains fail scale and support intervals, and none of 432 screened configurations passes the robustness gate.
  • Towards Zero-Shot Transfer Across Embodiments For Driving VLAs

    • Caio Azevedo, Stefano Sabatini, Sascha Hornauer, Fabien Moutarde
    • Publish Date: 2026.09.02
    • Task: End-to-End
    • Summary:
      • Studies multi-dataset training for driving VLAs and proposes BEV-Forcing, an auxiliary objective that transfers ground-plane object-layout information from a specialized Bird's-Eye-View model into the VLA backbone.
      • Shows that BEV-Forcing can improve both in-distribution and out-of-distribution performance when training on a small number of camera rigs.
      • Finds that benefits of the auxiliary task diminish as the number of training embodiments increases, motivating evaluation that accounts for data scaling.
  • CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation

    • Shucheng Zhang, Yuang Zhang, Bingzhang Wang, Muhammad Monjurul Karim, Kehua Chen, Yinhai Wang
    • Publish Date: 2026.09.02
    • Task: Generation
    • Datasets: WOMD
    • Summary:
      • CrashDiffuser is a closed-loop VLM-guided diffusion framework for fine-grained safety-critical traffic scenario generation, decoupling semantic collision reasoning from continuous trajectory synthesis.
      • It uses a hierarchical collision-intent interface from the requested target contact region, where the VLM extracts reusable scene-level context and predicts structured action tuples to condition a diffusion model for executable adversarial trajectories.
      • On WOMD-derived closed-loop scenarios, it achieves 50.33% single-attempt and 67.98% three-attempt target-collision rates, 40.05% contact-region control success, and competitive trajectory naturalness.
  • CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction

    • Zhengxu Tang, Guofeng Cui, Ziyu Gong, Xiaozhou Zhang, Ruifeng Deng, Chengzhi Qi, Ke Chen, Sachin Patil, Tianjun Xiao, Langechuan Liu, Pichao Wang
    • Publish Date: 2026.08.31
    • Code: CoLT-Drive
    • Task: Prediction
    • Datasets: CoLT-Drive
    • Summary:
      • CoLT-Drive, a 3,536-sample counterfactual long-tail benchmark for decision-level driving affordance prediction that inserts rare objects into otherwise fixed driving scenes and measures whether models predict acceptable longitudinal--lateral meta-action pairs.
      • KPA, a knowledge-preserving adaptation framework combining structured perception-to-decision prompting, SLERP-based expert merging, and RegMoE, a regime-aware LoRA mixture-of-experts module.
      • KPA achieves 60.8% pair accuracy on CoLT-Drive, outperforming the pretrained Qwen3-VL-2B baseline (50.3%) and LoRA SFT (32.4%) while maintaining competitive in-domain accuracy.
  • Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

    • Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
    • Publish Date: 2026.08.31
    • Task: Planning
    • Summary:
      • Qwen-Drive-1.0 is an initial vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning within a single framework.
      • It integrates an external BEV perception head for 3D object detection, semantic occupancy prediction, and BEV map segmentation, providing an explicit and inspectable interface to 3D scene structure.
      • A Planning Expert generates future ego trajectories from shared VLM representations, while a staged training recipe combines driving supervision with general-purpose vision-language data to preserve broad visual understanding and instruction-following capabilities.
  • Rethinking Language's Role in Efficient VLA for Autonomous Vehicles: Toward Smarter, Trustworthy Driving

    • Tongfei Guo, Lili Su
    • Publish Date: 2026.08.31
    • Task: End-to-End
    • Datasets: nuScenes, NAVSIM, Bench2Drive
    • Summary:
      • Reframes language's role in efficient Vision-Language-Action (VLA) models for autonomous driving by introducing the Language Residue taxonomy (L1–L4) based on inference-time use of language.
      • Reviews representative methods across deployment axes (latency, parameters, memory, FLOPs, tokens) and analyzes them on open- and closed-loop driving benchmarks such as nuScenes, NAVSIM, and Bench2Drive.
      • Traces how efficient NLP/LLM methods are adapted to autonomous driving, identifying constraints and motivations, with a continuously updated repository to be available on GitHub.
  • Self-Aware Active Learning Enables Continual Improvement in Autonomous Driving

    • Dong Hu, Chao Huang, Carman K. M. Lee, Dimitrios Kanoulas
    • Publish Date: 2026.08.30
    • Task: Planning
    • Datasets: Waymo
    • Summary:
      • SAGE is an active learning framework for post-training adaptation in autonomous driving that uses a predictive world model to generate fear and curiosity signals.
      • Fear estimates short-horizon predictive risk and model uncertainty, while curiosity measures novelty through prediction error and adaptively calibrates the intervention threshold; when fear exceeds it, control transfers to an expert or fallback policy for focused imitation learning.
      • Evaluated in simulated route-transfer tasks, Waymo-based logged driving, CARLA occlusion hazards, and real-world mobile robot navigation, SAGE improves robustness, reduces safety violations, and maintains task performance.
  • GeoWAM: Visual Geometry World Action Models for Autonomous Driving

    • Yiren Lu, Xin Ye, Jiaming Liu, Philip Jacobson, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, Min Cai, Tom Lampo, Yu Yin, Danhua Guo, Burhan Yaman
    • Publish Date: 2026.08.24
    • Task: Planning
    • Summary:
      • GeoWAM introduces a visual geometry world action model for autonomous driving that uses point-cloud geometry as a more natural state space than pixels.
      • The model is pretrained to forecast future scene geometry, learning representations that jointly encode spatial structure and temporal evolution.
      • A geometry-conditioned action head leverages these learned geometric dynamics to predict future ego trajectories, yielding stronger driving policies than image-based alternatives in open-loop and closed-loop evaluations.
  • MomADv2: Reliable Temporal Memory for End-to-End Autonomous Driving

    • Ziying Song, Shengkai Zhang, Lin Liu, Peiliang Wu, Lei Yang, Dongyang Xu, Bin Sun, Li Wang, Shaoqing Xu, Caiyan Jia, Yadan Luo
    • Publish Date: 2026.08.24
    • Task: Planning
    • Datasets: NAVSIM, Bench2Drive, nuScenes
    • Summary:
      • Proposes MomADv2, a reliable state-space memory framework for long-horizon end-to-end autonomous driving that addresses command-inconsistent memory.
      • Introduces a Selective State-Space Planning Memory Query Module to filter historical planning queries and a Flow-Matching Trajectory Residual Refiner for fine-grained trajectory refinement.
      • Experiments on NAVSIM, Bench2Drive, and nuScenes show improved planning consistency and a 15.6% reduction in average collision rate over MomAD under 6-second planning.
  • Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning

    • Yitao Xu, Tong Wu, Yiyan Wu, Guoji Xu, Yanbo Jiang, Jiahao Wang, Zehong Ke, Junkai Jiang, Fang Zhang, Jianqiang Wang
    • Publish Date: 2026.08.21
    • Task: End-to-End
    • Summary:
      • Introduces V2XBench, a V2X simulation platform with synchronized ego-roadside sensing and closed-loop evaluation, and Chat-V2XBench, a progressively structured VQA dataset for cooperative reasoning.
      • Proposes AURORA, an end-to-end cooperative driving framework with dual-view perception and a Query-level Cross-View Query Alignment and Fusion (CQAF) module to align ego and roadside viewpoints.
      • Achieves state-of-the-art closed-loop performance on V2XBench in heavily occluded scenarios, with 98.21% Route Completion and 76.02 Driving Score at low roadside communication bandwidth.
  • WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

    • Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
    • Publish Date: 2026.08.21
    • Code: WA-JEPA
    • Task: Planning
    • Datasets: NAVSIM, nuPlan
    • Summary:
      • WA-JEPA is a V-JEPA-native world-action model designed for autonomous driving planning, rethinking the V-JEPA paradigm to overcome its limitations with random-mask completion and deterministic regression.
      • It employs hybrid future-masked pre-training and conditional flow matching over latent futures, and introduces a joint future-action predictor that denoises future scene tokens and ego trajectories together in a unified spatiotemporal latent space.
      • Pre-trained on nuPlan videos and fine-tuned on NAVSIM, WA-JEPA achieves 91.7 EPDMS on NAVSIM-v2 and attains the best HD-Score of 0.4462 on HUGSIM without HUGSIM-specific fine-tuning.
  • A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving

    • Jingtao Sun, Xiaohai He, Yike Zhang, Dong Huang, Yaonan Wang, Ajmal Mian, Mike Zheng Shou
    • Publish Date: 2026.08.21
    • Task: Planning
    • Summary:
      • Proposes a robust VLA-based end-to-end autonomous driving system that combines multi-modality interaction with multi-trajectory planning and optimization for more reliable, interpretable, and safer driving decisions.
      • Introduces three core components: Affinity-Guided Optimal Transport for main-auxiliary modality two-way interaction; Distribution-Consistent Modality Transfer for heterogeneous modality distribution transfer and cross-modal interaction; and Multi-modal Multi-Trajectory Planning with Perception-Oriented Trajectory Refinement.
      • Experimental results on open-loop and closed-loop datasets demonstrate improvements in safety, long-horizon driving reasoning, and road scene perception over existing driving systems.
  • Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving

    • Mehdi Azarafza, Faezeh Pasandideh, Ali Ehteshami Bejnordi, Stefan Henkler, Achim Rettberg
    • Publish Date: 2026.08.20
    • Task: Planning
    • Datasets: CARLA
    • Summary:
      • Proposes a hybrid multi-agent orchestration framework that uses LLM common-sense reasoning to coordinate PPO-trained reinforcement learning and PID control for autonomous driving.
      • Applies LLM reasoning iteratively to refine the RL reward function for dynamic driving environments, evaluated in highly randomized CARLA scenarios while retaining structured control and safety mechanisms.
  • Planning-Oriented End-to-End Autonomous Driving: Architectures, Evaluation, and Emerging Paradigms

    • Yanchen Guan, Xingcheng Liu, Bin Rao, Chengyue Wang, Guofa Li, Yunjian Li, Lishengsa Yue, Zhiyong Cui, Chengzhong Xu, Zhenning Li
    • Publish Date: 2026.08.20
    • Task: Planning
    • Summary:
      • Reviews the transition from camera-to-control regression to planning-oriented end-to-end autonomous driving, covering behavior cloning, conditional imitation learning, privileged distillation, BEV and vectorized planning, unified perception-prediction-planning architectures, world-model-based planners, and vision-language-action systems.
      • Organizes existing methods along four axes: input representation, planning output, supervision signal, and evaluation protocol, and examines the benchmark shift from open-loop trajectory matching to closed-loop simulation, non-reactive real-log evaluation, long-tail testing, and human-preference-aware metrics.
      • Argues that architectural progress requires benchmark-consistent evaluation and identifies open challenges in uncertainty-aware planning, learner-expert mismatch, runtime safety assurance, language-action grounding, world-model validation, and reproducible benchmarking.
  • DA-WAM: Decision-Aligned Future Latents for Driving World Models

    • Ruiguo Zhong, Benshan Ma, Xiaolong Chen, Lang Zhang, Mingyue Feng, Yaonong Wang, Pei Liu, Jun Ma
    • Publish Date: 2026.08.19
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • DA-WAM is a framework that unifies predictive representation learning, action-conditioned future modeling, and trajectory scoring under a single decision-making objective for autonomous driving.
      • It maintains predictive supervision throughout planner optimization via an online encoder and a stable momentum target, allowing future representations to co-evolve with the driving task.
      • An action-conditioned predictor generates a distinct future latent state per trajectory candidate, which is evaluated by a future-latent-conditioned factorized scorer with supervision from expert-matched trajectories and safety-critical hard negatives.
      • Extensive experiments on NAVSIM-v1 and NAVSIM-v2 demonstrate state-of-the-art performance.
  • GAPL: Grounded Action-effect Policy Learning for LLM-Based Trajectory Planning

    • Zhihong Cui, Hengyu Liu, Zhangkai Wu, Yushuai Li, Tianyi Li, Peiyuan Guan, Amir Taherkordi, Tor Skeie
    • Publish Date: 2026.08.18
    • Task: Planning
    • Datasets: Highway-env
    • Summary:
      • GAPL is a unified framework for LLM-based trajectory planning that integrates LLM-based effect estimation, simulation-based effect grounding, and policy optimization into a closed-loop system.
      • It consists of an LLM-based Effect Evaluator for structured multi-dimensional action-effect estimation, a Simulation-based Effect Grounder for dynamics-consistent effects, and an Effect-Aware Decision Maker that distills effect estimates to guide PPO-based policy learning.
      • Experiments on four Highway-env scenarios show consistent improvements over baselines, with average reductions in collision rate, ADE, and FDE of 0.76, 0.86, and 2.00, and an average reward gain of 1.44.
  • Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving

    • Zongzheng Zhang, Jijun Wang, Saining Zhang, Shuo Wang, Yiru Wang, Hai Yang, Yang Chen, Yuwen Heng, Hao Sun, Anqing Jiang, Hao Zhao
    • Publish Date: 2026.08.18
    • Task: End-to-End
    • Datasets: nuScenes, NAVSIM-v1/NAVSIM-v2, Bench2Drive
    • Summary:
      • Presents the first systematic investigation of traffic element awareness for end-to-end autonomous driving.
      • Constructs a unified research infrastructure by augmenting multiple public driving datasets with comprehensive traffic-element annotations.
      • Introduces a minimal plug-and-play integration design that consistently improves driving performance across modern paradigms and datasets, establishing a new state of the art on NAVSIM-v2.
  • Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics

    • Ran Chen, Jiaxing Ren, Zhikun Zhang, Yunhao Hou, Junbao Zhuo, Bochao Zou
    • Publish Date: 2026.08.18
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • Geo-VLA is a plug-and-play framework that enhances VLA models by learning geometry-aware visual representations, internalizing geometric map semantics during training while requiring no HD maps or additional lane information at inference.
      • Introduces Geo-QA, a geometry-focused question-answering dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning.
      • Achieves 92.1 PDMS on NAVSIM v1 and establishes a new state-of-the-art among single-camera VLA planners.
  • DriveCache: Action-Aware Caching for Driving World Model Inference

    • Jianchun Yang, Jian Liang, Xianda Guo, Pinhan Fu, Yanlun Peng, Conglang Zhang, Wenke Huang, Mang Ye
    • Publish Date: 2026.08.17
    • Task: Prediction
    • Summary:
      • DriveCache is a training-free, action-aware caching controller for driving world model inference that uses planned motion to allocate feature reuse across scenes.
      • It uses dynamic programming to place reuse across denoising steps under a calibrated response budget, with a causal drift check to refresh features and replan when generation departs from calibration.
      • Across three generator configurations, DriveCache improves the overall fidelity-efficiency trade-off over evaluated cache methods; code will be publicly available.
  • Risk-Adaptive Edge--Cloud Visual Reasoning for Communication-Efficient Autonomous Driving

    • Meng Ma, Shuyang Li, Naigang Wang, Ruimin Ke
    • Publish Date: 2026.08.15
    • Task: Planning
    • Summary:
      • Presents a risk-adaptive edge–cloud architecture where onboard traffic assessment triggers selective cloud VLM reasoning, reducing unnecessary visual uploads.
      • Uses an onboard VLM and lightweight detector for temporal traffic and path-relative hazard assessment, while keeping validation, vehicle control, and automatic emergency braking local.
      • In CARLA experiments, matches periodic cloud access success while reducing cloud requests by 54.1% and fewer AEB activations, and continues reducing cloud traffic across emulated network profiles.
  • SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models

    • Haojie Feng, Peizhi Zhang, Xinrui Zhang, Zhuoren Li, Junpeng Huang, Xiurong Wang, Dongxiao Yin, Yuxiang Zhang, Junfan Zhu, Lu Xiong
    • Publish Date: 2026.08.14
    • Task: End-to-End
    • Summary:
      • SSP (Synthetic-Simulation-Physical) is an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison of autonomous driving VLA models to the same safety-critical interaction.
      • It starts from a synthetic long-tail video, builds a validated event specification preserving road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases, and constructs platform-specific realizations in CARLA and on a closed proving ground with transfer audits.
      • SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response; across Cut-in and vulnerable-road-user crossing cases, macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in Synthetic, Simulation, and Physical domains, respectively.
  • LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles

    • Md Wasiul Haque, Sagar Dasgupta, Mizanur Rahman, Md Rayhanur Rahman
    • Publish Date: 2026.08.13
    • Task: Reasoning
    • Summary:
      • Investigates whether large language models (LLMs) can automate dynamic confirmation of attacker-reachable software weaknesses in Autoware, an open-source autonomous-driving stack.
      • Performs compiler-precise static analysis across 185 packages, identifying 1,375 decision rules, 2,274 validation checks, and 482 input-to-safety-output flows, then generates and compiles 3,700 artifact sets using two local open-weight LLMs.
      • Finds that build integration is the primary barrier: 80% of first-shot compilation failures stem from dependency wiring, the reasoning model compiled 64% of harnesses on the first attempt versus 6% for the code-specialized model, and no candidate weakness was dynamically confirmed within budget.
  • FIRE-VLA: Failure-Informed Self-Evolution for Vision-Language-Action Models in Autonomous Driving

    • Hao Dou
    • Publish Date: 2026.08.13
    • Task: Planning
    • Datasets: nuScenes
    • Summary:
      • FIRE-VLA introduces a failure-informed self-evolution framework for autonomous-driving VLA models, converting low-reward, low-diversity groups into privileged supervision via self-distillation from a frozen round-start copy of the same model, with only the teacher observing the hidden future trajectory.
      • It keeps GRPO active for every group while routing unresolved failures into answer-token supervision conditioned on the student's generated prefix; the updated policy then supplies the next round's teacher without requiring a larger external teacher.
      • Starting from the same Qwen2.5-VL-3B SFT checkpoint, FIRE-VLA retains comparable single-sample planning while reducing G=4 mean L2 from 1.848 to 1.500 m and lowering evaluation-persistent failure prevalence from 13.03% to 11.20% on 6,019 examples from 150 held-out nuScenes scenes.
  • FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

    • Zekai Li, Yihao Liang, Hongfei Zhang, Jian Chen, Yesheng Liang, Zhijian Liu
    • Publish Date: 2026.08.13
    • Task: End-to-End
    • Summary:
      • FlashDrive is an algorithm-system co-design framework that accelerates Vision-Language-Action inference for autonomous driving by simultaneously targeting four bottlenecks: visual encoding, language-model prefill, reasoning token generation, and flow-matching denoising.
      • It exploits algorithmic shortcuts—streaming KV-cache reuse across frames, a non-autoregressive diffusion drafter for speculative decoding, and adaptive step caching guided by the velocity field's structure—layered on CUDA Graph compilation and kernel fusion.
      • Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive cuts end-to-end latency from 717ms to 151ms (4.7x) with essentially unchanged accuracy, raising throughput from 1.4 Hz to 6.6 Hz on a single GPU.
  • BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

    • Bing Zhan, Shuyao Shang, Shuo Lu, Yuan Xu, Zhao Wang, Yida Wang, Xueyang Zhang, Kun Zhan, Jiahao Gu
    • Publish Date: 2026.08.13
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • BrainWAM is a structured action-space coordination framework for autonomous driving that unifies semantic priors from VLA models and predictive dynamics from WAMs by aligning two specialized action-oriented pathways at compact action representations.
      • It addresses the attention-allocation mismatch in naive token-level joint attention, where semantic shortcuts dominate the shared attention space and suppress predictive dynamics.
      • It introduces asynchronous rectified-flow inference with decoupled video and action denoising, achieving state-of-the-art performance on NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS).
  • Do Not Forget the Obvious - RISC: A Risk-Informed Slice-Coverage Protocol for Safe Autonomous Driving

    • Fabian Hüger
    • Publish Date: 2026.08.12
    • Task: Perception
    • Datasets: Zenseact Open Dataset
    • Summary:
      • RISC (Risk-Informed Slice Coverage) is a practical, model-agnostic protocol for risk-guided stress testing and coverage-qualified evaluation, translating safety concerns into machine-readable risk slices and selecting compact audit sets by risk.
      • The protocol uses lightweight signals to tag candidate data and can optionally leverage an LLM to surface relevant but overlooked conditions during test planning, helping engineers not to forget the obvious.
      • Instantiated for monocular pedestrian perception using 1,000 frames from the Zenseact Open Dataset, image statistics, and a YOLO-based detector proxy, risk-guided selection raises critical failure discovery from 34.0% under random sampling to 98.5%.
  • RiskWorld: Object-Centric Latent World Modeling for Autonomous Driving Risk Identification

    • Jingzheng Li, Yufei Ge, Qianren Mao, Zhijun Chen, Bing Li, Xingyu Peng, Baochang Zhang, Xianglong Liu
    • Publish Date: 2026.08.12
    • Task: Prediction
    • Summary:
      • RiskWorld is an object-centric latent world model that identifies risk by imagining the evolution of each candidate object relative to the ego vehicle.
      • It combines pretrained predictive video representations with structured ego-object histories, contextualizes observed interactions, and rolls relation-aware object states into the future using RSSM-style latent dynamics.
      • It decodes rollouts into object-level risk scores with auxiliary future-relation and temporal-risk predictions, achieving 63.0% F1 and 2.1% false-alarm rate on RiskBench.
  • XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving

    • Foundation Model Team, XPeng Inc
    • Publisher: XPeng Inc
    • Publish Date: 2026.08.11
    • Task: Planning
    • Summary:
      • XCoT-VLA replaces verbose natural-language Chain-of-Thought with compact executable CoT tokens learned from automatically constructed Reason-Action supervision, connecting scene understanding, semantic reasoning, and trajectory generation.
      • The method conditions fixed trajectory queries via shared multimodal self-attention, with deterministic token-function routing applying the Reason FFN to XCoT tokens and the Control FFN to trajectory queries for flow-matching trajectory generation.
      • XCoT Policy Optimization (XCPO) is introduced as an optional refinement extension, and XCoT-VLA reduces longitudinal ADE from 1.645 to 1.323 and lateral FDE from 1.616 to 0.648 using only 2-6 executable XCoT tokens within the real-time planning budget.
  • Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

    • Kiet T. Nguyen, Hanbo Shim, Jinwoo Kim, Seunghoon Hong
    • Publish Date: 2026.08.11
    • Task: Reasoning
    • Summary:
      • MVRD distills patch-wise cosine similarities across views instead of teacher features to improve visual-spatial reasoning while preserving language alignment.
      • It outperforms supervised fine-tuning and feature distillation, approaching feature fusion methods with considerably fewer added parameters and lower latency.
      • MVRD generalizes to 3D scene understanding tasks such as object grounding, dense captioning, and question answering.
  • DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

    • Zebin Xing, Yupeng Zheng, Qiang Chen, Linbo Wang, Yichen Zhang, Pengxuan Yang, Junli Wang, Deheng Qian, Xiaoqing Ye, Junyu Han, Yifeng Pan, Qichao Zhang, Dongbin Zhao
    • Publish Date: 2026.08.11
    • Code: DriveVLA-M0
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • DriveVLA-M0 is a retrieval-augmented Vision-Language-Action model with failure-aware latent memory for end-to-end autonomous driving.
      • It constructs a latent memory pool storing failure cases with structured scene representations and expert trajectory labels, and uses a Retrieve Model that decouples static road structure and dynamic agent interactions for structurally grounded retrieval.
      • Retrieved cases are injected at inference via lightweight decoupled LoRA-based test-time training (TTT) for targeted scenario-specific correction without modifying the backbone, achieving 94.1 PDMS on Navtest and 47.0 EPDMS on Navhard with only 26.44 ms TTT backward latency overhead.
  • Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

    • Jiazhuo Li, Linjiang Cao, Qi Liu, Xi Xiong
    • Publish Date: 2026.08.11
    • Task: Planning
    • Summary:
      • Dreamer-SAC integrates a recurrent state-space world model with an off-policy soft actor-critic algorithm trained directly in latent space for sample-efficient autonomous driving.
      • The framework combines real interactions with short-horizon generated trajectories, n-step target estimation, and multi-objective supervision.
      • It outperforms DreamerV3, SAC, and PPO with substantially fewer real environment interactions, and reveals an inverted-U relationship between rollout horizon and policy performance.
  • Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

    • Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian, Richard Shi, Jinjun Shan, Amir Rasouli, Dongfeng Bai
    • Publish Date: 2026.08.10
    • Task: Reasoning
    • Summary:
      • Space Tokens is a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without additional inference-time modules.
      • It distills scene-level 3D geometry and object-centric spatial attributes into continuous latent tokens, enabling their direct incorporation into chain-of-thought reasoning to improve spatial reasoning.
      • Experiments on VSI-Bench improve Qwen3-VL-8B by 4.3% and SenseNova-SI-1.3 by 1.3%, while achieving state-of-the-art performance on object size (79.2%) and room size estimation (75.7%).
  • 4D-WAM: 4D Consistent World Modeling for Autonomous Driving

    • Jiacheng Fu, Yibo Yuan, Meng Tian, Yue Li, Jiangtong Zhu, Jianhua Han, Yueyi Zhang, Jianwu Fang, Jianru Xue, Hang Xu, Zhiwei Xiong
    • Publish Date: 2026.08.10
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • 4D-WAM leverages geometric foundation models for training-time supervision, feeding WAM-predicted future frames into a geometric foundation model and using 4D-aware responses to define a 4D consistency loss, enabling physically consistent 4D scene understanding and prediction without additional inference cost.
      • It identifies an early-decision phenomenon in WAMs and proposes a decision-oriented timestep sampling strategy that emphasizes supervision at early, high-noise stages where driving decisions are primarily formed, further improving trajectory planning.
      • Extensive experiments show that 4D-WAM effectively models 4D consistent scene evolution and achieves state-of-the-art performance on the challenging NAVSIM-v1 and NAVSIM-v2 benchmarks.
  • FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving

    • Guolei Huang, Tengfei She, Yuxuan Lu, Yao Huang, Yuqi Ye, Yongjun Shen
    • Publish Date: 2026.08.10
    • Task: Planning
    • Datasets: nuScenes, NAVSIM
    • Summary:
      • FactorDrive, an end-to-end autonomous driving framework for adaptive multi-step reasoning driven by planning-critical factors (PCFs), grounding planning reasoning in trajectory-relevant spatial-physical evidence.
      • Introduces PCF-CoT, a chain-of-thought dataset organizing reasoning around scene-specific PCFs so that composition and depth of reasoning paths adapt to different planning demands.
      • Proposes Quality Search-Guided Group Relative Policy Optimization (QS-GRPO), which uses MCTS guided by trajectory-level planning rewards to discover higher-quality reasoning paths and optimizes the policy via GRPO, achieving state-of-the-art planning performance on nuScenes and NAVSIM.
  • DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving

    • Ziyi Song, Chen Xia, Hang Yu, Sheng Zhou, Zhisheng Niu
    • Publish Date: 2026.08.10
    • Task: Planning
    • Summary:
      • DH-VLM is a dual-horizon cooperative latent reasoning framework that enables asymmetric semantic cooperation between infrastructure and ego vehicle.
      • It integrates infrastructure-driven latent evolution to refine conditional latents, allowing the ego vehicle to leverage long-range contextual understanding while preserving autonomous local planning.
      • The framework includes a cooperation-oriented QA dataset and achieves state-of-the-art planning performance, reducing L2 error by 14.6%, collision rate by 26.9%, communication cost by 57.3%, and GPU memory usage by 25.5%.
  • CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving

    • Junyao Wang, Yulin Xu, Yu Li, Pramod Khargonekar, Mohammad Abdullah Al Faruque
    • Publisher: University of California, Irvine
    • Publish Date: 2026.08.10
    • Task: Perception
    • Summary:
      • CRUISE is a novel uncertainty-aware cross-modal sensor fusion framework for robust autonomous driving.
      • It integrates a vision-language model (VLM)-guided uncertainty quantification module that generates fine-grained, pixel-level uncertainty estimates.
      • A dynamic adaptive mechanism explicitly models cross-modal dependencies to exploit the complementary nature of multi-sensor inputs, improving generalization in out-of-distribution scenarios.
  • UnsDrive: Towards Robust End-to-End Autonomous Driving in Unstructured Scenes

    • Nanxin Zeng, Ruiqi Song, Xiangyu Guo, Baiyong Ding, Yunfeng Ai
    • Publish Date: 2026.08.10
    • Task: Planning
    • Summary:
      • UnsDrive is an end-to-end planner for unstructured mining scenes that builds an unknown-aware occupancy representation using multi-frame visibility cues and conditions a flow-matching planner to generate multimodal future trajectories.
      • It introduces an occupancy trajectory consistency loss and an uncertainty-aware trajectory scorer to penalize trajectories entering non-traversable or unobserved regions, improving safety under partial observability.
      • The paper also presents MineLoop, a mining-oriented closed-loop simulator, and shows UnsDrive outperforms strong baselines in trajectory accuracy, collision avoidance, and long-horizon driving robustness.
  • SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

    • Zongchuang Zhao, Xin Zhou, Tianyang Xu, Zhengyang Sun, Kaixuan Zhou, Yu Wu, Honglin Li, Dingkang Liang, Xiang Bai
    • Publisher: Huazhong University of Science and Technology
    • Publish Date: 2026.08.07
    • Code: SimWAM
    • Task: Planning
    • Datasets: NAVSIM, nuScenes
    • Summary:
      • SimWAM is a simple yet effective World-Action Model that leverages future-video prediction as a training-time supervision signal, co-training a pretrained video expert and a lightweight action expert with joint flow matching.
      • An isolated attention mask keeps action prediction independent of future frames, enabling trajectory prediction without explicit future-frame generation at inference while allowing the video backbone to be replaced and the action expert scaled independently.
      • Reinforcement learning is further applied to optimize a compositional driving reward beyond trajectory imitation, achieving 91.5 PDMS on NAVSIM with substantially lower latency and zero-shot transfer to nuScenes.
  • CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

    • Hsu-kuang Chiu, Stephen F. Smith
    • Publisher: Carnegie Mellon University
    • Publish Date: 2026.08.07
    • Task: End-to-End
    • Summary:
      • CMU-Drive is a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles operating in safety-critical driving scenarios with background traffic participants.
      • V2V-VLA is a cooperative Vision-Language-Action model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies.
      • Experiments establish the first benchmark and baseline for cooperative VLA driving, providing a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving.
  • Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features

    • Sining Ang, Yuguang Yang, Yan Wang
    • Publish Date: 2026.08.06
    • Task: Planning
    • Datasets: NAVSIM, nuScenes
    • Summary:
      • Adaptive-WAM is a quality-aware multi-exit planner built on a Wan2.2-5B backbone, with trajectory diffusion heads attached to selected DiT blocks.
      • It uses a lightweight trajectory-quality scorer to terminate inference once the best decoded trajectory satisfies a quality threshold, otherwise continuing from cached hidden states to deeper exits.
      • It avoids iterative classifier-free denoising and VAE decoding for future-video synthesis, achieving 90.8 PDMS with adaptive single-trajectory planning, 92.6 PDMS with a fixed-exit variant, 89.9 EPDMS on NAVSIM v2, and transfer to nuScenes without target-domain fine-tuning.
  • muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards

    • Yang Chen, Yicheng Zhu, Tao Li, Zilin Bian
    • Publish Date: 2026.08.05
    • Task: Generation
    • Summary:
      • muSync-GS is a physics-synchronized framework for driving video synthesis under adverse-weather and road-elevation hazards, jointly coupling road-surface condition with tire friction and road-elevation profile with axle excitation.
      • It uses a calibrated vehicle model to predict speed, slip ratio, normal loads, and pitch for constructing ego-camera trajectories and synchronized physical annotations.
      • On 12 held-out CarSim cases, it achieves mean case-wise RMSEs of 0.0273 m/s for speed, 0.0590 degrees for pitch, 0.0101 for slip ratio, and 26.61 N for per-wheel normal load, while synchronizing controllable scene edits with vehicle responses.
  • Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

    • Jiaju Han, Xuemeng Sun, Qike Zhang, Xiang Chen, Luwei Yang, Jiahuan Long, Yiwei Wei, Jiujiang Guo, Chengyin Hu
    • Publish Date: 2026.08.04
    • Task: Reasoning
    • Datasets: K-Radar
    • Summary:
      • Radar4D-VLM is a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input.
      • It extracts geometrically grounded object proposals and organizes radar evidence into object, scene, and kinematic tokens, which are mapped into frozen language backbones with a parameter-efficient projector and auditable prediction heads.
      • On sequence-isolated K-Radar development validation, Top-64 proposal recall reaches 98.13% at 4 m, and the radar-token interface remains compatible across Qwen, Phi, Mistral, Llama, and Gemma backbones, while matched language-control runs show sensor dependence but no stable direct-head gain from aligned language supervision.
  • The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk

    • Francis Heylighen
    • Publish Date: 2026.08.04
    • Summary:
      • Traces the evolutionary origin of value in biological organisms, arguing that values emerge from autopoiesis and natural selection's vicarious selectors.
      • Contrasts LLMs as allopoietic and allotelic systems that lack autonomous drives for self-preservation, dominance, resource competition, or embodied suffering.
      • Concludes that the real alignment challenge is ensuring LLMs intelligently apply learned ethical values, and that the orthogonality and instrumental convergence theses do not apply to them.
  • RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models

    • Yuwei Ning, Liangzhi Wang, Yi Xiao, Zhenhua Wu, Yun Pang, Mingkun Chang, Jichang Li, Guanbin Li
    • Publisher: Huazhong University of Science and Technology
    • Publish Date: 2026.08.03
    • Project Page: RealWeather
    • Task: Perception
    • Summary:
      • RealWeather is a driving world model for realistic and scene-faithful weather translation, learning authentic weather dynamics directly from real-world videos.
      • It employs Progressive Realism Bootstrapping with a Pseudo-Clear Generation pipeline to bridge the pseudo-to-real domain gap and support bidirectional clear-adverse translation.
      • Scene-Fidelity RL Optimization penalizes alterations to safety-critical driving elements, improving visual realism, structural preservation, long-tail weather generation, and zero-shot out-of-distribution generalization.
  • MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

    • Ambarish Govindarajulu Kaliamurthi, Kaikai Liu
    • Publish Date: 2026.08.03
    • Task: Reasoning
    • Datasets: nuScenes
    • Summary:
      • MoRAL is a two-stage fine-tuning pipeline that teaches Cosmos-Reason2-2B to read a physics-encoded BEV representation and reason over it for driving decisions.
      • It externalizes spatial perception into the input image by encoding LiDAR metric distance, object class, and radar Doppler velocity, avoiding a learned 3D backbone at inference.
      • Stage 1 trains the vision encoder on 60,000 grounding records; Stage 2 trains the full model on 57,696 chain-of-thought records, achieving gains on seven of eight nuScenes question types and improving emergency braking recall.
  • TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving

    • Mengmeng Zhu, Yuxuan Sun, Wei Chen, Bo Ai
    • Publish Date: 2026.08.03
    • Task: Reasoning
    • Datasets: nuScenes
    • Summary:
      • TALSC is a timeliness-aware large-small VLM collaboration framework for infrastructure-assisted autonomous driving, addressing on-board compute constraints and latency-induced timeliness degradation.
      • It models AoI evolution for VLM inference and derives a timeliness metric coupling AoI, token length, and task performance, then proposes an online scheduling algorithm with guaranteed performance.
      • Simulations on nuScenes show TALSC outperforms baselines, achieving up to a 12.6% normalized improvement in Micro-F1.
  • Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

    • Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Xiaozhi Chen, Yikun Ban, Deqing Wang
    • Publish Date: 2026.08.03
    • Task: Reasoning
    • Summary:
      • Identifies trajectory anchoring bias in chain-of-thought annotations for autonomous driving VLAs when teacher models are exposed to logged ground-truth future trajectories, leading to less causally faithful reasoning and more hallucinations.
      • Introduces AD-MCQ, which reformulates planning as multiple-choice selection among explicit trajectory candidates, enabling verifiable trajectory-level decisions without open-ended geometric synthesis.
      • Proposes DEFT-RLVR, deferring future trajectory exposure from pre-decision anchors to post-decision verification targets, improving AD reasoning while preserving general visual capabilities.
  • STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision

    • Pou-Chun Kung, Aryaman Rao, Utkrisht Sahai, Hemanth Murali, Yi Liu, Rui-Yu Lin, Katherine A. Skinner
    • Publish Date: 2026.08.02
    • Task: Reasoning
    • Summary:
      • STAR-VLM is an automotive radar-supervised framework that enhances spatiotemporal vision-language models with motion reasoning and metric velocity estimation for autonomous driving.
      • By leveraging radar range and Doppler measurements as label-free ground truth during training, STAR-VLM improves metric spatiotemporal reasoning without complex preprocessing or expensive human annotations.
      • Experiments demonstrate state-of-the-art performance on motion classification and metric velocity estimation, outperforming task-specific methods.
  • PRISM: Privileged Probabilistic Latent Supervision for End-to-End Autonomous Driving Motion Planning

    • Volodymyr Havrylov, Faris Janjoš, Andreas Look, Jürgen Mathes, Andreas Geiger
    • Publish Date: 2026.08.02
    • Task: Planning
    • Datasets: nuScenes
    • Summary:
      • Investigates VLM latent supervision for end-to-end autonomous driving, finding that performance gains stem from latent connections to ground-truth data rather than VLM reasoning.
      • Proposes a probabilistic deep supervision framework that regularizes intermediate latent representations directly from GT data by treating model latents as reparameterizable distributions and optimizing via the Evidence Lower Bound.
      • On nuScenes, supervising trajectory-related latents with future GT paths reduces planning L2 error by 8% and collision rate by 3% compared with vectorized baselines, with negligible computational overhead.
  • WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

    • Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He, Yaoyi Li, Jianhua Han, Hang Xu, Siyu Zhu
    • Publish Date: 2026.08.02
    • Task: End-to-End
    • Summary:
      • WAM-Diff2 is a multi-task discrete diffusion VLA framework that transforms pre-trained autoregressive generalists into parallel diffusion models using a three-stage hierarchical distillation strategy.
      • It bridges mismatched causal/bidirectional attention patterns and divergent optimization objectives through progressive block-wise adaptation, block-wise distillation, and model-wise cross-scale distillation, preserving semantic foundations while accelerating inference.
      • Evaluations across driving understanding, perception, and planning benchmarks show reduced exposure bias, performance parity with autoregressive baselines, and up to 15.1x decoding acceleration with system-level optimizations including FlashInfer and CUDA Graphs.
  • WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation

    • Jianxun Cui, Ping Wu, Stanisa Peric, Marko Milojkovic, Vladan Devedzic
    • Publish Date: 2026.07.31
    • Task: Planning
    • Summary:
      • WM-Cov formulates interactive world-model-style testing adequacy and introduces a provider-agnostic evaluation layer that converts raw provider outputs into requested, realized, and valid evidence.
      • It reports adequacy through coverage growth, valid-failure discovery, failure-mode diversity, realism, artifact suppression, duplicate accounting, and valid-evidence precision.
      • Studies on TeraSim/SUMO events, WM-like mixed trace pools, and a DriveArena matrix show that dangerous-looking events can include valid ADS failures, duplicates, partial realizations, and artifacts, supporting evaluation by convergence of valid interactive evidence under budget.
  • Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

    • Meibo Hu, Jiamian Wang, Pichao Wang, Zhiqiang Tao
    • Publish Date: 2026.07.31
    • Code: LCS
    • Task: Planning
    • Datasets: Bench2Drive, nuScenes
    • Summary:
      • Latent-Centroid Steering (LCS) addresses the command-following gap in VLM-based end-to-end autonomous driving, identified as conditional policy collapse.
      • Introduces a principled classifier-free guidance formulation for regression-based vision-language driving, interpreting CFG as isolating the instruction-induced residual in action space.
      • Proposes single-pass guidance using class-level latent shifts toward precomputed command-specific centroids, reducing inference latency by ~50% while improving command adherence on Bench2Drive and nuScenes.
  • LLM-Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality

    • Noura Fady, Farah Khaled, Catherine M. Elias
    • Publish Date: 2026.07.31
    • Task: Perception
    • Datasets: nuScenes
    • Summary:
      • Introduces an automated pipeline that transforms safe driving scenes into safety-critical scenarios by combining computer vision, Large Language Models (LLMs), and Augmented Reality (AR).
      • Detects and tracks road users, extracts safety features including distance, velocity, motion direction, and Time-to-Collision (TTC), and assesses scene criticality; an LLM generates realistic collision-inducing objects and behaviors integrated into the original scene via AR.
      • Evaluated on the nuScenes dataset, achieving 97.52% safety classification accuracy and successfully generating realistic scenarios such as pedestrian crossings, rear overtaking vehicles, and sudden-stop events.
  • Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving

    • Zeyu Dong, Yimin Zhu, Yu Wu, Yu Sun
    • Publish Date: 2026.07.31
    • Task: End-to-End
    • Datasets: Waymo
    • Summary:
      • Proposes a teacher-student framework that integrates structured reasoning and geometric precision for end-to-end autonomous driving with Vision-Language Models.
      • Introduces reflective reasoning in the teacher model to refine logical explanations under ground-truth action supervision, and distills reasoning into a student model via supervised fine-tuning.
      • Designs a separate waypoint decoder to convert textual reasoning into continuous trajectories, improving zero-shot reasoning, waypoint accuracy, and inference efficiency on Waymo benchmarks.
  • Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

    • Jiwei Yang, Zhengxian Chen, Chaosheng Huang, Jun Li
    • Publish Date: 2026.07.31
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • Auto-JEPA is an action-oriented latent world model that learns continuous future driving intent through joint-embedding prediction, avoiding dense future-world reconstruction.
      • It predicts an intent embedding aligned with the latent representation of the future ego trajectory, retrieves executable trajectories from a fixed trajectory memory, and ranks them with a scene-conditioned candidate selection module.
      • It achieves 91.3 PDMS on NAVSIM v1 and 89.1 EPDMS on NAVSIM v2, and semantic occlusion experiments show it focuses on planning-relevant dynamic-agent regions.
  • Mitigating Compounding Error via Video Representation Regularization

    • Taiye Chen, Qi Zhang, Yisen Wang
    • Publish Date: 2026.07.29
    • Task: Prediction
    • Datasets: VBench
    • Summary:
      • Investigates the internal representation dynamics of video diffusion-based world models and finds that compounding error in sliding-window autoregressive inference is tightly coupled with dimensional collapse of hidden representations, with effective rank dropping sharply at the onset of generation drift.
      • Shows that pure training data scaling fails to improve resistance to error drift, contradicting mainstream scaling paradigms, and adopts erank as a quantitative metric for error accumulation.
      • Proposes video representation regularization, a lightweight training constraint that stabilizes latent representations and suppresses iterative error accumulation, improving VBench Aesthetic Quality from 38.65 to 55.56 and Imaging Quality from 44.37 to 72.08 over Diffusion Forcing.
  • Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

    • Yueru Luo, Xu Yan, Changqing Zhou, Yiming Yang, Chao Zhan, Shuqi Mei, Chao Zheng, Zhen Li
    • Publish Date: 2026.07.27
    • Task: Reasoning
    • Summary:
      • Proposes a reasoning-based framework that equips vision-language models with chain-of-thought capabilities for traffic regulation understanding and regulation-aware autonomous driving.
      • Introduces a scalable CoT curation pipeline that bootstraps rationales from a strong LLM via a two-round strategy and uses a VLM-based verifier to filter incorrect cases, producing high-quality (CoT, answer) pairs.
      • Adopts a two-stage training scheme with supervised fine-tuning followed by GRPO reinforcement learning using answer-grounded, fine-grained rewards; experiments on MapDR demonstrate improved interpretability and accuracy.
  • HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

    • Quanfu Yu, Xian Wu, Hao Xu, Liulong Ma
    • Publish Date: 2026.07.23
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • HyWorldVLA is a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning for end-to-end autonomous driving.
      • It pre-trains by predicting video latents encoded by a pre-trained video VAE while reconstructing video frames for pixel-level grounding, then co-fine-tunes to predict latent features fed into an action expert for trajectory generation.
      • Experiments on NAVSIM v1 and v2 show significant improvements over pixel-based and latent-based world model baselines, and present the first comprehensive qualitative and quantitative analysis of world model noise robustness.
  • PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

    • Yushan Liu, Tianxiong Lv, Bohua Wang, Hangqi Fan, Chenxu Zhao, He Zheng, Xuchang Zhong, Yifan Xie, Congyang Zhao, Zhihao Liao, Leigang Luo, Yang Cai, Xiao-Ping Zhang, Wenbo Ding
    • Publish Date: 2026.07.22
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • PerceptDrive is a perception prior world-action modeling framework with adaptive expert routing, addressing the prior-to-plan transfer problem for end-to-end autonomous driving.
      • It feeds teacher-distilled priors from a frozen driving-adapted provider and dense observation latents from a frozen self-supervised video encoder into a trainable expert-routed world-action model with expert-specific query branches, prior-retention objectives, and scene-conditioned soft-gate routing.
      • PerceptDrive achieves state-of-the-art performance with 90.4 PDMS on NAVSIM v1 and 90.2 EPDMS on NAVSIM v2, generating one trajectory per planning step with a single front-facing camera and no test-time scoring, reranking, or search.
  • SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

    • Jiangfan Liu, Zexuan Cui, Tianyuan Zhang, Zonglei Jing, Zonghao Ying, Yaoyuan Zhang, Jiakai Wang, Xiaoqi Jiang, Aishan Liu, Xianglong Liu
    • Publish Date: 2026.07.22
    • Task: End-to-End
    • Summary:
      • SafeGen is a goal-conditioned diffusion framework for safety-critical scenario generation in VLM-based autonomous driving, addressing the sim-to-real gap and the lack of realistic human-vehicle interaction dynamics in simulator-based pipelines.
      • It formulates scenario generation as a goal-conditioned diffusion process where a predefined catastrophic end-state serves as a strong supervisory signal, guiding temporally coherent video trajectories toward safety-critical outcomes.
      • It introduces Context Grounded End State Reasoning and End State Conditioned Video Evolution with depth-aware geometric projection and boundary-conditioned diffusion; experiments across 3 VLMADs show a 24.25% average increase in Judge Overall Score and 15.9% average improvement in real-world driving after fine-tuning.
  • Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency

    • Zhongyao Yang, Haoyu Li, Yu Yan, Zhuangxuan Yu, Jiangfeng Nan, Jinrui Nan
    • Publish Date: 2026.07.21
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • A cognitive dual-process planning framework represents planning-relevant scene knowledge in a machine-parsable structured chain-of-thought (S-CoT) schema, with an automated data engine that integrates perception foundation models, critical-path filtering, and an expert VLM to generate S-CoT supervision without manual annotation of individual rationales.
      • A lightweight visual Arbiter estimates scene complexity before language decoding and routes each input to either fast meta-action prediction or slow structured reasoning, while a deterministic rule-based validator verifies S-CoT consistency with the final meta-action and provides rewards for Group Relative Policy Optimization (GRPO).
      • Achieves 91.8% CoT accuracy and 98.5% Logical Consistency Score in a 195-scene audit, and 80.14% planning accuracy with 97.20% LCS on 574 manually verified NAVSIM test samples while reducing average latency by 17.39% relative to applying slow reasoning to every scene.
  • GeoWorldAD: Geometry World Action Model for Autonomous Driving

    • Songyan Zhang, Jinyuan Tian, Hanbing Li, Daqi Liu, Hao Chen, Wenhui Huang, Fang Li, Guang Chen, Hangjun Ye, Long Chen, Kuiyuan Yang, Chen Lv
    • Publish Date: 2026.07.20
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens.
      • It progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement.
      • Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.
  • WorkDrive: Roadwork Chain of Causation for Autonomous Driving

    • Tianyi Jiang, Wen Zhang, Sihan Yang, Ming Lu, Wentao Zhang
    • Publish Date: 2026.07.16
    • Task: Planning
    • Datasets: ROADWork
    • Summary:
      • WorkDrive constructs perception-grounded causal reasoning for roadwork zones by extracting structured scene facts and injecting them into a Chain-of-Causation annotation pipeline.
      • It uses supervised fine-tuning on reasoning labels followed by reinforcement learning with consistency between lateral meta-actions and the predicted trajectory as the sole reward.
      • On ROADWork, the roadwork CoC reduces trajectory average displacement error (ADE) by 9.0%, and consistency-based GRPO yields a further 3.0%, achieving progressive improvement over the trajectory-only baseline.
  • Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving

    • Yuan Gao, Wenting Miao, Mattia Piccinini, Haoyu Wang, Qunying Song, Johannes Betz
    • Publisher: Technical University of Munich
    • Publish Date: 2026.07.15
    • Code: Chat2Scenic
    • Task: Planning
    • Summary:
      • Chat2Scenic is the first iterative retrieval-augmented framework to generate scenario scripts in Domain Specific Language (DSL) for autonomous driving.
      • It provides a chatbot interface for interactive scenario refinement and integrates Retrieval-augmented Generation (RAG) to ground scenario generation in regulatory knowledge and DSL syntax.
      • It introduces an open benchmark of 123 scenarios from NHTSA and United Nations Vehicle Regulations, achieving 76.42% CSR and 58.17% FA, outperforming existing methods.
  • M4World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

    • Ke Cheng, Hanqiao Ye, Lei Shi, Yahui Liu, Yunhan Shen, Jingtao Dong, Zhenke Wang, Wenxuan Ao, Weixiang Xu, Kaining Huang, Shuhan Shen
    • Publish Date: 2026.07.15
    • Task: Prediction
    • Summary:
      • M4World is a multi-view multimodal driving world model that synthesizes future surround-view video and synchronized LiDAR scans while supporting interactive object manipulation and stable minute-long streaming.
      • It achieves fine-grained object manipulation through a flexible conditioning interface and stable minute-long streaming via a multi-stage training framework with online causal generation in only four denoising steps.
      • It introduces an automated VLM-based judging pipeline for controllability evaluation and demonstrates potential for controllable, scalable driving simulation.
  • S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving

    • Jianguo Yu, Rukang Wang, Duanfeng Chu, Chen Wang, Renju Feng, Liping Lu
    • Publish Date: 2026.07.15
    • Task: Planning
    • Datasets: NAVSIM
    • Summary:
      • S-squared-VLA explicitly decouples the semantic and spatial streams in Vision-Language-Action models, using hierarchical bridging for multi-scale VLM intent reasoning while an independent spatial stream bypasses the autoregressive language bottleneck to preserve uncompressed spatial features.
      • A dual-stream planning adapter fuses high-level semantic intent with precise spatial constraints via cascaded attention mechanisms, and auxiliary perception supervision equips the model with rich spatial and geometric priors.
      • On the NAVSIM closed-loop benchmark, it achieves a PDMS of 87.1, a new state-of-the-art for VLA models under purely supervised fine-tuning, with the highest No Collision rate of 98.4 among all evaluated methods.
  • Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Embodiment Adaptation

    • Zhidong Wang, Jingsong Liang, Zirui Li, Zhan Chen, Han Yu, Chen Lv
    • Publish Date: 2026.07.15
    • Task: End-to-End
    • Summary:
      • DynaDreamer introduces an ego-dynamics-augmented world model that conditions latent distributions on a physics-informed ego-dynamics context from a lateral dynamics model with a neural tire force formulation.
      • A neural-ODE encoder-decoder extracts this context online and identifies chassis parameters, removing ego-motion terms from transition entropy and prior-posterior KL divergence.
      • The framework enables zero-shot cross-embodiment adaptation without per-platform retraining, achieving 28% and 43% improvements in urban and highway success rates and up to 73% gain over the base Transformer WM on unseen chassis.
  • Technical Report on the CVPR 2026@AdvML Workshop Challenge

    • Tianyuan Zhang, Zonglei Jing, Jiangfan Liu, Ligong Zhang, Ke Ma, Chengzhi Sun, Xiaohai Xu, Zhirui Zhang, Qianqian Xu, Qingming Huang, Hanyu Fang, Junhua Liu, Zheng Wang, Xiaoliang Liu, Yuanbo Li, Shuai Gui, Bin Wang, Menghe Zheng, Jing Nie, Hanyang Meng, Zeyang Zhang, Xiang Zhang, Yongxuan Zhu, Rui Ding, Hainan Li, Yongkang Zhang, Zhilei Zhu, Xianglong Kong, Jin Hu, Zonghao Ying, Yisong Xiao, Lei Chen, Haotong Qin, Jiakai Wang, Aishan Liu, Ruikai Li, Julia Karbing, Yinpeng Dong, Zhenfei Yin, Shao Jing, Xia Hu, Jingyi Xu, Juntao Dai, Xinyun Chen, Vishal M. Patel, Xianglong Liu, Dawn Song, Alan Yuille, Philip H. S. Torr, Dacheng Tao
    • Publish Date: 2026.07.13
    • Task: VQA
    • Datasets: DriveLM
    • Summary:
      • This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving Vision-Language Agents (VLAs).
      • The challenge is built on DriveLM-style multi-view VQA with six synchronized camera images and structured question-answer pairs, requiring adversarial images and suffix-only textual perturbations that preserve image fidelity and limit textual cost.
      • It comprises two phases, with Phase II adding a hidden black-box model for transferability, and analyzes five leading submissions to identify recurring attack patterns and vulnerabilities in driving VLAs.
  • LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving

    • Yongzhi Liu, Yang Xiao, Zhong Cao, Zeng Kang, Sunan Zhang, Zhaozhi Dong, Guojun Yu, Weichao Zhuang
    • Publish Date: 2026.07.12
    • Task: Planning
    • Datasets: nuPlan
    • Summary:
      • LIDAR-AD is a decoder-free Latent-Interaction Dreamer with Action-Residual Chains for autonomous driving.
      • It replaces observation reconstruction with redundancy-reduced latent alignment and models vehicle control as residual action updates using residual-action sequence contrastive learning.
      • Experiments across simulated driving scenarios and nuPlan-derived log-reconstructed scenarios show consistent improvements over world-model baselines and strong transferability.
  • Large Language Model Enhanced Differentiable Trajectory Planning for IoT-Enabled Autonomous Driving

    • Shihao Zhang, Jing Yang, Ziyu Song, Zheng Lin, Sunil Prajapat, Zhaochen Xia, Hemant Ghayvat, Haitao Ding, Lip Yee Por, Ashok Kumar Das
    • Publish Date: 2026.07.11
    • Task: Planning
    • Datasets: nuPlan
    • Summary:
      • Proposes an LLM-enhanced differentiable trajectory planning framework for IoT-enabled autonomous driving, addressing imitation learning limitations in long-tail interactions, constrained refinement, and real-time semantic utilization.
      • Introduces surrounding-agent-centric data augmentation and a complexity-aware asynchronous LLM-based semantic enhancement module to improve training distribution and extract high-level scene semantics with controlled overhead.
      • Incorporates differentiable optimization with explicit residual penalties and gradient backpropagation to the upstream planner, achieving top nuPlan Hard20 scores and verifying CARLA-ROS deployment.
  • WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving

    • Xuerun Yan, Zhexi Lian, Nuoheng Zhang, Shiyu Fang, Haoran Wang, Chen Lv, Jia Hu, Binyang Song
    • Publish Date: 2026.07.09
    • Task: End-to-End
    • Datasets: NAVSIM
    • Summary:
      • WCog-VLA is a novel dual-level World-Cognitive Vision-Language-Action framework that bridges semantic world forecasting with generative world evolution to enable proactive autonomous driving.
      • At the semantic level, it unifies world cognition and reasoning by incorporating 3D spatial perception, injecting agent tokens to capture world dynamics, and enabling Game-theoretic Chain-of-Thought reasoning.
      • At the generative level, it introduces the Aligned Decoupled Diffusion Transformer (ADDT) for physically-plausible joint multi-agent trajectory synthesis, reduces denoising steps via scene representation alignment, constructs an 85k Game-CoT dataset, and achieves a SOTA PDMS score of 92.9 on NAVSIM.
  • CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

    • Kaicong Huang, Meng Ma, Ruimin Ke
    • Publish Date: 2026.07.08
    • Task: Planning
    • Datasets: Waymo Open Dataset
    • Summary:
      • CARLA-GS is a modular corner-case synthesis pipeline that decouples visual representation, semantic reasoning, and physics-based execution while maintaining tight cross-module coupling.
      • It reconstructs an editable Gaussian scene from real driving data with geometry-consistent constraints, uses a multi-agent LLM for scene-level reasoning and intent-level waypoint generation, and delegates low-level motion control to CARLA with a PID controller for kinematic and dynamic feasibility.
      • Experiments on the Waymo Open Dataset show controllable corner-case generation and photorealistic, spatiotemporally consistent videos aligned with semantic intent and physically feasible motion.
  • Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators

    • Christian Oefinger, Finn Rasmus Schäfer, Korbinian Moller, Mattia Piccinini, Johannes Betz
    • Publish Date: 2026.07.08
    • Task: Planning
    • Summary:
      • Proposes an admissibility ladder (L0-L4) that a World-Model simulator must climb before its closed-loop verdicts are accepted as assurance evidence.
      • Builds on safety-critical simulation practices including Verification, Validation & Accreditation (VV&A), Safety of the Intended Functionality (SOTIF), and scenario-based testing standards, with an embodiment-agnostic framework instantiated in autonomous driving.
      • Demonstrates on two driving World Models that higher visual generation quality (L0) does not predict action-following robustness (L1-L2), revealing a reversal and questioning visual fidelity as a trust proxy.
  • A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving

    • Heye Huang, Jingguang Li, Zhiyuan Zhou, Paul Liang, Mingyu Wu, Kitae Jang, Jianqiang Wang
    • Publish Date: 2026.07.08
    • Task: Reasoning
    • Summary:
      • K-Risk is a knowledge-augmented dataset that combines structured driving trajectories with LLM-generated semantic annotations for safety-critical driving scenarios.
      • It integrates 20 human-driven and autonomous-vehicle trajectory datasets from Europe, China, and the United States, covering highways, urban freeways, intersections, and roundabouts.
      • Using a unified risk-centric extraction pipeline, K-Risk curates 31,398 high-risk events and a 1,036-event extreme near-collision subset, with synchronized trajectory, metadata, and language triplet annotations including structured scenario descriptions, abnormal-behavior notifications, and validated causal risk analyses and action recommendations.
  • Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation

    • Seulbin Hwang, Kiyoung Om, Daejung Kim, Jinhan Lee
    • Publish Date: 2026.07.08
    • Project Page: Flow-ERD
    • Task: Prediction
    • Datasets: WOSAC
    • Summary:
      • Flow-ERD, a multi-agent traffic simulator that pursues realism and diversity jointly, built on an Agent-Type Aware Flow Matching (AFM) backbone that couples flow matching's multi-modal expressiveness with type-specific kinematic execution.
      • A second stage, Entropy-Regularized Distillation (ERD), fine-tunes the closed-loop rollout distribution with an entropy-regularized reverse-KL objective, mitigating covariate shift while explicitly preventing collapse onto high-density modes.
      • Flow-ERD ranks first on the WOSAC test benchmark and dominates the realism–diversity Pareto front among reproducible baselines, evaluated with a log-free diversity metric alongside standard realism scores.
  • VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving

    • Tianjia Yang, Ke Li, Ruwen Qin, Xianbiao Hu
    • Publish Date: 2026.07.06
    • Code: VLM-CASE
    • Task: Planning
    • Summary:
      • VLM-CASE is a framework that gives autonomous vehicles anticipatory capacity, using a LoRA fine-tuned vision-language model to reason about road surface and visibility conditions from front-camera images while motion stays bounded by a formal safety model.
      • The VLM output parametrizes a context-adaptive safety envelope (CASE) derived from physical limits and responsibility-sensitive safety guarantees, coupling braking and steering through a shared friction budget, with a model predictive controller driving freely within the envelope while the VLM runs asynchronously.
      • Closed-loop CARLA validation across weather, road-surface, and lighting conditions shows VLM-CASE-MPC completes all trials, outperforming a conventional MPC baseline and a state-of-the-art VLM-integrated controller, with ablations confirming that friction and visibility adaptations are complementary.
  • Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning

    • Matthew Foutter, Matteo Cercola, Lena Wild, Yunshan Wang, Michelle Li, Daniele Gammelli, Marco Pavone
    • Publish Date: 2026.07.06
    • Project Page: Pinocchio
    • Task: Reasoning
    • Summary:
      • Distinguishes functional reasoning, where reasoning improves task performance, from faithful reasoning, where reasoning reflects the policy's internal decision process, and argues that SoTA alignment is necessary but insufficient for faithfulness.
      • Uses human evaluation of a SoTA autonomous driving reasoning model to show inconsistent coupling between reasoning quality and downstream trajectory improvement.
      • Introduces Pinocchio, a learned critic scoring observation grounding and stepwise coherence, as a dense reward for RL post-training; it improves faithfulness by 4% and 18% over SoTA alignment and trajectory-error baselines, and improves responsiveness to rare counterfactual scenarios by 1.6x.
  • PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving

    • Pin Tang, Guoqing Wang, Xiangxuan Ren, Zhongdao Wang, Guodongfang Zhao, Bailan, Chao Ma
    • Publish Date: 2026.07.06
    • Task: Planning
    • Summary:
      • PixelPilot is a novel Vision-Language-Action (VLA) model for end-to-end autonomous driving that features a decoupled planning and lifting paradigm.
      • It reformulates scene understanding and trajectory prediction as sensor-agnostic 2D-to-2D tasks in the image plane for scalable training across diverse datasets, and deterministically lifts planned 2D trajectories to 3D during inference.
      • A knowledge-instilled policy learning strategy applies dense intermediate rewards via Group Relative Policy Optimization (GRPO) to enforce a rigorous causal chain from visual perception to spatial planning, achieving state-of-the-art open-loop and closed-loop performance.
  • RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

    • Shuangyu Xie, Kaiyuan Chen, Ziyang Chen, Simeon Adebola, Yixuan Huang, Zehan Ma, Tianshuang Qiu, Wentao Yuan, Dhruv Shah, Pannag R. Sanketi, Ken Goldberg
    • Publish Date: 2026.07.06
    • Task: VQA
    • Summary:
      • Proposes Robot Question Answering (RQA), a modular evaluation framework, and RoboVista, a benchmark curated from real robotic systems, research papers, and expert annotations.
      • RoboVista contains 474 VQA instances with human-annotated reasoning and covers 39 unique task types across agricultural, industrial, domestic, surgical robotics, autonomous driving, and open robot datasets.
      • Experiments show that state-of-the-art VLMs exhibit substantial gaps, and physical robot experiments suggest a strong correlation between RoboVista performance and real-world task execution.
  • CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining

    • Jingyu Song, Yi Liu, Katherine A. Skinner
    • Publish Date: 2026.07.05
    • Project Page: CRISP
    • Task: Perception
    • Datasets: nuScenes
    • Summary:
      • CRISP is a spatiotemporal camera-radar backbone pretrained through forecasting-based representation learning, predicting future LiDAR point clouds as privileged supervision while requiring only camera and radar at deployment.
      • It introduces an enhanced radar encoder, radar-enhanced temporal self-attention, and multimodal feature rendering with modality innovation gating to inject radar range and Doppler cues into BEV temporal propagation.
      • Experiments on nuScenes show improved long-horizon point cloud forecasting and effective transfer to downstream tasks including 3D detection, tracking, online mapping, motion forecasting, future occupancy prediction, and planning.
  • CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving

    • Zhaohong Liu, Hao Ye, Xianlin Zhang, Mengshi Qi
    • Publisher: Beijing University of Posts and Telecommunications
    • Publish Date: 2026.07.05
    • Code: CritiqueDriveVLM
    • Task: Reasoning
    • Summary:
      • CritiqueDriveVLM is a unified three-stage framework that internalizes reasoning directly into a VLM to resolve the reliability-efficiency trade-off in autonomous driving.
      • It introduces Critique-Driven Multi-Turn Reinforcement Learning guided by a multi-dimensional verifier, training a robust System-2 Teacher without fragile external tools.
      • Latent Thought Distillation compresses the Teacher's reasoning into a fast, CoT-free System-1 Student, achieving 76.54% MCQ on DriveLMM-01 and reducing inference latency by 88% (to 416 ms).
  • AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning

    • Qi Liu, Yabei Li, Hongsong Wang, Heng Zhang, Lei He
    • Publish Date: 2026.07.03
    • Task: Planning
    • Datasets: Bench2Drive
    • Summary:
      • AnchorVLA is a hierarchical decision-anchored VLA planning framework that uses trajectory-pattern anchors as an explicit interface between high-level VLA reasoning and continuous trajectory execution.
      • It introduces Decision-as-Anchor Representation, where anchor tokens encode behavior-level driving decisions as entire local motion patterns, and Decision-Anchored Residual Flow, which generates fine-grained continuous trajectories in the selected anchor-defined residual space.
      • Experiments on the Bench2Drive closed-loop benchmark show state-of-the-art Success Rate of 77.28 and a competitive Driving Score of 89.92.
  • CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving

    • Yunxiao Shi, Hong Cai, Mohammad Ghavamzadeh, Fatih Porikli
    • Publish Date: