Alibaba’s Qwen team has introduced Qwen3.8-Flash-Next, a multimodal mixture-of-experts model featuring 125 billion total parameters with only 6 billion activated per token. Serving as an architectural preview for Qwen4, the model introduces a 51-billion-parameter N-gram embedding layer designed to store common word groups in system RAM rather than GPU memory. The production API version, Qwen3.8-Flash, will deploy via QwenCloud priced at $0.16 per million input tokens and $0.47 per million output tokens.
According to the Qwen team, the model delivers improved results over Qwen3.7-Plus at approximately one-ninth the training cost. In benchmark tests, Flash-Next outperformed rival models DeepSeek-V4-Flash and Anthropic’s Claude Opus 4.6 (Max) across several agentic coding and office productivity evaluations, including scoring 62.5 on SWE-bench Pro and 73.9 on CoWorkBench.
The release underlines ongoing price and efficiency competition among top AI providers. Operating just below Alibaba’s flagship Qwen3.8-Max at roughly one-twelfth the cost, Flash-Next’s aggressive pricing structure increases competitive pressure on Western providers like OpenAI and Anthropic, who are adjusting pricing models to defend market share.
Why it matters
Offers developers high-performing agentic coding capabilities at a fraction of the cost of legacy frontier models.
Demonstrates architectural techniques like offloading phrase dictionaries to system RAM to reduce GPU memory bottlenecks.
Signals intensifying price competition that compresses profit margins for proprietary model providers like OpenAI and Anthropic.
Source: the-decoder.com



