Data from developer platform OpenRouter reveals that weekly token consumption surged over 25,000 percent from January 2025 to reach 126.2 trillion tokens. While the spike illustrates massive growth in API activity, industry analysts caution that raw token metrics no longer map directly to proportional increases in end-user adoption or software revenue.
The volume growth is largely driven by reasoning models and unoptimized agentic AI systems that generate high volumes of internal intermediate tokens prior to producing an output. OpenAI’s GPT 5.6 Luna accounted for a significant share of consumed tokens, while spending on Chinese models such as Kimi, GLM, and DeepSeek expanded tenfold off smaller baselines.
Despite the exponential token throughput, the trend highlights the divergence between raw compute output and economic value. As agentic loops and reasoning chains consume higher token volumes per user request, developer efficiency and context optimization become critical focus areas.
Why it matters
Demonstrates how reasoning models and agent loops inflate token volume without necessarily reflecting underlying user growth.
Highlights rapid adoption of Chinese models like DeepSeek and Kimi alongside dominant offerings from OpenAI.
Encourages developers to optimize agentic architectures to control escalating API costs tied to high internal thinking tokens.
Source: the-decoder.com



