SOURCE-LINKED INTELLIGENCE
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
We introduce Qwen-Drive-1.0, the first vision-language foundation model for autonomous driving that unifies 3D perception and visual question answering at the pretraining stage and further extends to motion planning, while keeping the pretrained VLM architecture entirely untouched. Built on the natively multimodal Qwen3.5-4B, it attaches two external modules. A BEV perception head serves as an exp
Read original source ↗ Open in workspace
- recordType
- page-entry
- evidenceStatus
- publisher-reported
- region
- Global
Evidence & attribution
- Qwen News · 2026-09-03T00:00:00.000Z
First collected: 2026-09-23T00:41:11.323Z. This is not the publication date.