about
Qwen-Drive-1.0 (arxiv.org)
2 points by amancina79 30 days ago | hide | past | pdf | discuss on HN

In plain words: A general vision-language model gets a bird's-eye-view head that maps 3D objects, occupied space, and road layout, plus a planner that turns shared scene understanding into future driving paths. It nearly matched the best driving performance while keeping most of its general vision and language skills.

Abstract · Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, et al.
arXiv:2609.00111 · cs.CV · submitted Aug 31, 2026
abstract · pdf · html · Code will be available at https://github.com/QwenLM/Qwen-Drive-1.0

add comment on HN