Alibaba’s Qwen team has launched Qwen3.8-Omni-Flash, a multimodal AI model engineered specifically for autonomous agentic workflows. Built to process simultaneous audio and video inputs over a 1 million-token context window, the model enables capabilities such as direct video editing, real-time translation, and autonomous tool usage.
Qwen claims the new model closely matches Google’s Gemini 3.8 Flash across key audio and video benchmarks while significantly undercutting its competitor on price. API rates are set at $0.15 per million input tokens and $0.47 per million output tokens, compared to Gemini’s introductory pricing of $0.75 input and $3.75 output.
The launch includes open-source agent plugins for Claude Code, Gemini CLI, and Qwen Code, alongside a real-time interaction harness. By undercutting established API pricing, Qwen aims to lower operational costs for developers building real-time multimodal agents.
Why it matters
Aggressive price competition from open-weights labs continues to drive down unit economics for multimodal API consumers.
Developers building video and audio AI agents gain significant cost reduction for real-time tool use and media processing.
Lower API costs put pressure on major cloud providers like Google to adjust developer pricing strategies.
Source: the-decoder.com



