Alibaba’s research division has introduced Qwen-Drive 1.0, an end-to-end AI model designed to combine 3D spatial perception, traffic Q&A, and route planning into a single framework. Built upon the Qwen3.5-4B base model, the architecture integrates a specialized 3D perception module to generate bird’s-eye-view maps alongside a dedicated Planning Expert module tuned through reinforcement learning.
The researchers noted that vision-language models do not natively grasp 3D space purely through image captioning data. Performance only improved when the underlying vision-language network was explicitly co-trained on spatial tasks alongside driving datasets. This combined approach preserved general world knowledge while reducing the spatial errors common in adapted vision models.
In testing, Qwen-Drive 1.0 surpassed its base model on traffic reasoning tasks and reduced simulated driving error rates. However, the study highlighted ongoing industry challenges, showing that a model’s natural language explanations for driving decisions do not always reliably align with its underlying physical maneuvers.
Why it matters
Highlighting disconnects between model explanations and physical actions emphasizes persistent safety risks in end-to-end AI driving systems.
Consolidating perception, dialogue, and route planning onto single chips offers cost reductions for automotive hardware design.
Source: the-decoder.com



