Learning Vision-Language-Action World Models for Autonomous Driving

Guoqing Wang; Pin Tang; Xiangxuan Ren; Guodongfang Zhao; Bailan Feng; Chao Ma

arXiv:2604.09059·cs.CV·April 13, 2026

Learning Vision-Language-Action World Models for Autonomous Driving

Guoqing Wang, Pin Tang, Xiangxuan Ren, Guodongfang Zhao, Bailan Feng, Chao Ma

PDF

1 Repo

TL;DR

VLA-World introduces a unified vision-language-action world model that enhances autonomous driving by combining predictive imagination with reflective reasoning, leading to improved foresight and safety.

Contribution

The paper presents VLA-World, a novel VLA world model that integrates future scene generation with reasoning, supported by a new dataset and a three-stage training process.

Findings

01

VLA-World outperforms existing models on planning benchmarks.

02

The model achieves higher accuracy in future scene generation.

03

Reflective reasoning improves trajectory prediction quality.

Abstract

Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and control within a unified multimodal framework. However, they often lack explicit modeling of temporal dynamics and global world consistency, which limits their foresight and safety. In contrast, world models can simulate plausible future scenes but generally struggle to reason about or evaluate the imagined future they generate. In this work, we present VLA-World, a simple yet effective VLA world model that unifies predictive imagination with reflective reasoning to improve driving foresight. VLA-World first uses an action-derived feasible trajectory to guide the generation of the next-frame image, capturing rich spatial and temporal cues that describe how the surrounding environment evolves. The model then reasons over this…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

https://vlaworld.github.io
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.