
AI agents now out-consume humans on OpenRouter, 14x since February
AI is using more AI. February 6, 2026 may have been the last day humans consumed more tokens than AI agents on OpenRouter, OpenRouter analyst Peter Walker found, via The Decoder. Agentic token usage has grown 14x since then, while human usage is up only 2.8x over the same stretch.
OpenRouter works as a routing marketplace: developers and AI agents call dozens of different model providers through one unified API instead of integrating each one separately. That makes it a useful window into how token consumption splits between humans typing prompts and agents running tasks on their own.
Token consumption by AI agents on OpenRouter jumped from 0.51 trillion to 7.3 trillion tokens since February 2026, a roughly 14x increase. Human usage grew only 2.8x over the same period. Agents increasingly work on their own over longer stretches, spinning up additional AI processes along the way rather than waiting on a single human-to-model exchange.
- Agent token usage since Feb. 2026: 0.51 trillion to 7.3 trillion, up 14x
- Human token usage over the same period: up 2.8x
- Share of agent tokens from cached prompts: nearly 70%
- Cached prompts: billed at much lower rates than fresh tokens
- OpenRouter's model mix: skews toward less token-efficient open-weight models
The gap makes sense once you look at what an agent does with each task. A human chatting with a model sends one short message and reads one reply. An agent working a multi-step task loops through planning, tool calls, intermediate reasoning, and self-checks, generating tokens at every stage that a human user never directly types or sees. Each of those internal steps counts toward the running total even though only the final output ever reaches a person on the other end.
The raw numbers overstate the cost impact, though. Nearly 70% of agent token usage comes from cached prompts, which are billed at much lower rates than fresh tokens because the model reuses context it already processed instead of reprocessing it from scratch. Actual spending isn't rising nearly as fast as the token count alone suggests. That distinction matters for anyone trying to forecast AI infrastructure costs from usage charts alone: a 14x jump in raw tokens does not translate into a 14x jump in the bill, since the marginal cost of a cached token can run a fraction of a fresh one.
OpenRouter skews toward open-weight models, which tend to be less token-efficient than models from OpenAI or Anthropic, so the exact multiple may differ elsewhere. But the underlying trend likely looks similar across the major labs, and token inflation already started with reasoning models, which think longer before responding even in cases where the extra thinking isn't needed. The shift points to a broader pattern in how AI compute gets consumed, which is the same dynamic behind Meta paying Microsoft hundreds of millions for Azure AI access: a growing share of that spending goes toward AI systems coordinating with other AI systems, not toward a person typing a single prompt.
Nothing here should be taken as financial advice — just information to consider.

Comments (0)
No comments yet — be the first!
Related news
Most readTop 7
Silicon Valley Workers Are Wearing Noise-Cancelling Masks to Dictate AI Prompts
251AI





