AI research lab Z.ai has released GLM-5.3-Flash, a 320-billion parameter open-weights multimodal model operating with 18 billion active parameters. Released under an MIT license with a 1-million-token context window, the model achieves a 57 on Artificial Analysis’s Intelligence Index at maximum reasoning effort, matching GPT-5.6 Terra while undercutting its larger sibling’s task cost by roughly 7.5 times at $0.09 per task.
On Z.ai’s API, pricing is set at $0.15 per million input tokens and $0.50 per million output tokens. According to benchmarks from Artificial Analysis and OpenCode, the model maintains strong performance on agentic benchmarks, matching Grok 4.6 on GDPval-AA v2. However, the model exhibits lower token efficiency, with approximately 90% of its output generation dedicated internal reasoning steps.
Crucially, Z.ai revealed that the model was served handling 100 trillion tokens daily during pre-launch testing entirely on non-Nvidia Chinese hardware. To achieve performance parity with standard GPU setups, Z.ai developed custom serving software built on SGLang to decouple processing stages, offering a rare real-world challenge to Nvidia’s established CUDA software ecosystem.
Why it matters
High-performing open-weights models continue driving aggressive price deflation for enterprise inference across standard and agentic tasks.
Demonstrated serving efficiency on non-Nvidia hardware shows viable pathways around CUDA lock-in for large-scale model deployments.
Chinese AI labs are closing the gap on frontier reasoning capability while aggressively optimizing operational expenditure.
Source: the-decoder.com



