Geo-VLA |
 Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics |
arXiv 2026 |
- |
- |
RedLight-VLA |
 RedLight-VLA: Models for traffic-rule grounding and behavioral emphasis in driving policies |
arXiv 2026 |
- |
- |
Collab-MMI |
 A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
Drive the Thoughts |
 Drive the Thoughts: Runtime Monitoring of VLA Reasoning-Trajectory Consistency |
arXiv 2026 |
- |
- |
|
|
|
|
|
LMDrive |
 LMDrive: Closed-Loop End-to-End Driving with Large Language Models |
CVPR 2024 |
 |
 |
BEVDriver |
 BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving |
IROS 2025 |
- |
- |
CoVLA-Agent |
 CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving |
WACV 2025 |
 |
- |
ORION |
 ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation |
ICCV 2025 |
 |
 |
SimLingo |
 SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment |
CVPR 2025 |
 |
 |
DriveGPT4-V2 |
 DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving |
CVPR 2025 |
- |
- |
AutoVLA |
 AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning |
NeurIPS 2025 |
 |
 |
DriveMoE |
 DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving |
arXiv 2025 |
 |
 |
DSDrive |
 DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning |
arXiv 2025 |
- |
- |
OccVLA |
 OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision. |
arXiv 2025 |
- |
- |
VDRive |
 VDRive: Leveraging Reinforced VLA and Diffusion Policy for End-to-End Autonomous Driving |
arXiv 2025 |
- |
- |
ReflectDrive |
 Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving |
arXiv 2025 |
- |
 |
E3AD |
 E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving |
arXiv 2025 |
- |
- |
LCDrive |
 Latent Chain-of-Thought World Modeling for End-to-End Driving |
arXiv 2025 |
- |
- |
Alpamayo-R1 |
 Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail |
arXiv 2025 |
- |
- |
UniUGP |
 UniUGP: Unifying understanding, generation, and planing for end-to-end autonomous driving. |
arXiv 2025 |
- |
- |
MindDrive |
 MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving |
arXiv 2025 |
- |
- |
AdaThinkDrive |
 AdaThinkDrive: Adaptive Thinking via Reinforcement Learning for Autonomous Driving |
arXiv 2025 |
- |
- |
Percept-WAM |
 Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving |
arXiv 2025 |
- |
- |
Reasoning-VLA |
 Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving |
arXiv 2025 |
- |
- |
SpaceDrive |
 SpaceDrive: Infusing Spatial Awareness into VLM-Based Autonomous Driving |
arXiv 2025 |
- |
- |
OpenDriveVLA |
 OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model |
AAAI 2026 |
 |
 |
WAM-Flow |
 WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving |
CVPR 2026 |
|
 |
ColaVLA |
 ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving |
CVPR 2026 |
 |
 |
AutoMoT |
 AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
OneDrive |
 OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models |
arXiv 2026 |
- |
- |
UniDriveVLA |
 UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving |
arXiv 2026 |
- |
- |
VLA-World |
 Learning Vision-Language-Action World Models for Autonomous Driving |
arXiv 2026 |
- |
- |
Metis |
 Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation |
arXiv 2026 |
- |
- |
MindVLA-U1 |
 MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving |
arXiv 2026 |
- |
- |
DVGT-2 |
 DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale |
arXiv 2026 |
- |
- |
VLGA |
 VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving |
arXiv 2026 |
- |
- |
VECTOR-Drive |
 VECTOR-Drive: Tightly Coupled Vision-Language and Trajectory Expert Routing for End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
ChainFlow-VLA |
 ChainFlow-VLA: Causal Flow Planning with Vision-Language Models |
arXiv 2026 |
- |
- |
LVDrive |
 LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model |
arXiv 2026 |
- |
- |
StyleVLA |
 StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving |
arXiv 2026 |
- |
- |
Masked-VLA-Diffusion |
 Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion |
arXiv 2026 |
- |
- |
Uni-World VLA |
 Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving |
arXiv 2026 |
- |
- |
Does |
 Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior? |
arXiv 2026 |
- |
- |
Reasoning |
 Reasoning About Traversability: Language-Guided Off-Road 3D Trajectory Planning |
arXiv 2026 |
- |
- |
SpanVLA |
 SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model |
arXiv 2026 |
- |
- |
Sim2Real-AD |
 Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving |
arXiv 2026 |
- |
- |
Drive My Way |
 Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving |
arXiv 2026 |
- |
- |
DriveVLM-RL |
 DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving |
arXiv 2026 |
- |
- |
Learning from Mistakes |
 Learning from Mistakes: Post-Training for Driving VLA with Takeover Data |
arXiv 2026 |
- |
- |
SAMoE-VLA |
 SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving |
arXiv 2026 |
- |
- |
LaST-VLA |
 LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving |
arXiv 2026 |
- |
- |
PixelPilot |
 PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
S-squared |
 S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving |
arXiv 2026 |
- |
- |
WAM |
 WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA |
arXiv 2026 |
- |
- |
AnchorVLA |
 AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning |
arXiv 2026 |
- |
- |
CLEAR |
 CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
Post |
 Post-Training in End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
Latent |
 Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving |
arXiv 2026 |
- |
- |
Inference |
 Inference-Time Attention Steering for Vision-Language-Action Driving Models |
arXiv 2026 |
- |
- |
Geo-VLA |
 Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics |
arXiv 2026 |
- |
- |
RedLight-VLA |
 RedLight-VLA: Models for traffic-rule grounding and behavioral emphasis in driving policies |
arXiv 2026 |
- |
- |
Collaborative Multi-Modality |
 A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
Qwen-Drive-1.0 |
 Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving |
arXiv 2026 |
 |
 |
LaPla |
 Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
GRAVA |
 GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving |
arXiv 2026 |
- |
 |
RAF-VLA |
 RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving |
arXiv 2026 |
- |
- |
FIVE-VLA |
 FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory |
arXiv 2026 |
- |
- |
Run-then-Walk |
 Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving |
arXiv 2026 |
- |
- |
CAR-VLA |
 CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving |
arXiv 2026 |
- |
 |
RefineDrive |
 RefineDrive: Reliable Failure-Guided Learning for Vision-Language-Action Driving |
arXiv 2026 |
- |
- |
AD-Memo |
 Vision-Language-Action Autonomous Driving Agent with Language-based Memory |
arXiv 2026 |
- |
- |
|
|
|
|
|