Alibaba’s Qwen team has introduced Qwen3.8-Flash-Next, a multimodal mixture-of-experts model featuring 125 billion total parameters with only 6 billion activated per token. Serving as an architectural preview for Qwen4, the model introduces a 51-billion-parameter N-gram embedding layer designed to store common word groups in system RAM rather than GPU memory. The production API version, Qwen3.8-Flash, will deploy via QwenCloud priced at $0.16 per million input tokens and $0.47 per million output tokens.

According to the Qwen team, the model delivers improved results over Qwen3.7-Plus at approximately one-ninth the training cost. In benchmark tests, Flash-Next outperformed rival models DeepSeek-V4-Flash and Anthropic’s Claude Opus 4.6 (Max) across several agentic coding and office productivity evaluations, including scoring 62.5 on SWE-bench Pro and 73.9 on CoWorkBench.

The release underlines ongoing price and efficiency competition among top AI providers. Operating just below Alibaba’s flagship Qwen3.8-Max at roughly one-twelfth the cost, Flash-Next’s aggressive pricing structure increases competitive pressure on Western providers like OpenAI and Anthropic, who are adjusting pricing models to defend market share.

Why it matters

  • Offers developers high-performing agentic coding capabilities at a fraction of the cost of legacy frontier models.

  • Demonstrates architectural techniques like offloading phrase dictionaries to system RAM to reduce GPU memory bottlenecks.

  • Signals intensifying price competition that compresses profit margins for proprietary model providers like OpenAI and Anthropic.

Source: the-decoder.com