AI agents use 5x more tokens than humans as cached prompts explode, headed for 10x
Oct 03, 2026 - 19:01
01
(Image credit: OpenRouter)
Futurum Group CEO Daniel Newman wrote in an X post that “AI is currently used by AI 5x more than it is used by humans. That number will accelerate to 10x and then higher and higher.” His data, an Andreessen Horowitz (a16z) chart of OpenRouter figures, shows agents at 7.3 trillion tokens versus humans’ 1.4 trillion as of August, six months after agent usage first surpassed humans. But the agents are mostly rereading what they’ve already seen: more than 85% of agent tokens come from cached prompts, a16z wrote, citing OpenRouter.
AI is currently used by AI 5x more than it is used by humans. That number will accelerate to 10x and then higher and higher.We keep speaking to human adoption when trying to determine ROI, but the utilization and scale is exponentially larger than that.September 30, 2026
OpenRouter is “a leading AI model gateway and routing platform.” A chart from its head of insights, Peter Walker, lists “7-day average token usage on OpenRouter split by type.” It states that, since the February crossover point, agents are using 14x more tokens while human usage is up 2.8x. The a16z chart in Newman’s post shows the same data. OpenRouter sorts each API key into one of three categories: agentic, mixed, or human, using a “7-signal weighted composite score that includes inputs such as tool call rate, turn count, gap timing, and others.”
(Image credit: OpenRouter)
The mixed category, possibly covering behavior that is part agent and part human, grew 4.7x over the same period, according to our math. Depending on how that traffic splits, the agents’ lead over humans may vary. The data also measures token volume, not spending. This is data from only one platform, and the trend isn’t completely consistent, with dips in April and July. However, agent use is growing elsewhere. In McKinsey’s 2026 State of AI survey, 40% of respondents from large organizations reported scaling AI agents, up from 27% a year earlier.
Cached tokens also account for nearly all of the relative growth in token usage, a16z wrote. They cost far less than processing a prompt from scratch, but they still have to be held in memory, and a16z, an OpenRouter investor, ties that to rising demand for high-bandwidth memory (HBM). Models keep that stored context in what’s called the KV cache, and “the KV cache is outgrowing GPU HBM capacity,” according to our reporting.
The same pattern shows up in the logs of a call center consultancy that tested DeepSeek on rented Nvidia H200s. In its own agents’ September usage on Claude Code, “96% of all input was re-reading old conversation.” Coupled with OpenRouter’s data, this suggests the token count overstates the bill, but the hardware cost for memory remains very real.
That memory is already scarce. Micron expects RAM and storage shortages to worsen in 2027 and 2028, with customers paying more than this year, while memory makers put HBM for AI data centers first. If Newman’s prediction that the 5x will become “10X, 20X, 30X” is true, PC buyers will be bidding for memory against even more agents.
Comments (0)