Alibaba’s research division has introduced Qwen-Drive 1.0, an end-to-end AI model designed to combine 3D spatial perception, traffic Q&A, and route planning into a single framework. Built upon the Qwen3.5-4B base model, the architecture integrates a specialized 3D perception module to generate bird’s-eye-view maps alongside a dedicated Planning Expert module tuned through reinforcement learning.

The researchers noted that vision-language models do not natively grasp 3D space purely through image captioning data. Performance only improved when the underlying vision-language network was explicitly co-trained on spatial tasks alongside driving datasets. This combined approach preserved general world knowledge while reducing the spatial errors common in adapted vision models.

In testing, Qwen-Drive 1.0 surpassed its base model on traffic reasoning tasks and reduced simulated driving error rates. However, the study highlighted ongoing industry challenges, showing that a model’s natural language explanations for driving decisions do not always reliably align with its underlying physical maneuvers.

Why it matters

  • Highlighting disconnects between model explanations and physical actions emphasizes persistent safety risks in end-to-end AI driving systems.

  • Consolidating perception, dialogue, and route planning onto single chips offers cost reductions for automotive hardware design.

Source: the-decoder.com