AI companies are now racing to the bottom — crashing token prices and competitive models push companies to cut costs
Aug 03, 2026 - 22:06
00
(Image credit: Malte Mueller via Getty Images)
We've entered a new phase of the AI industry's development, with all the major players heavily cutting costs and boosting the capabilities of their entry-level models in order to compete with new models from China, like Moonshot's Kimi K3 and DeepSeek's V4 Flash. OpenAI did so most recently, cutting the price of its base frontier model, ChatGPT 5.6 Luna, by 80% per million tokens, and its mid-range 5.6 Terra by 20%. This comes just over a week after Google introduced its more-affordable Gemini 3.6 Flash and 3.5 Flash-Lite models. Anthropic hasn't cut prices, but replaced its most-affordable Opus 4.8 model with a more capable Claude 5.0 at the same price point.
Intelligence is getting more affordable thanks to increased global competition, but this can come at the cost of margin for these major companies. This follows months of major AI businesses announcing cuts and limits on their use of the technology, even by major AI boosters like Elon Musk's xAI. Despite more workers using AI than ever before, productivity gains are reported to have been less than ideal.
Intensifying competition
The story of Chinese and American AI development efforts has been somewhat emblematic of the countries' historic strengths. While American firms burn through enormous amounts of money to push frontier technologies, Chinese developers have leveraged their industrial base to develop models that are cheaper, leaner, and almost as good at the top end.
DeepSeek gave Western AI developers a shock in 2025, and Kimi K3 did much the same in 2026. Alone, these events would cause concern for companies like OpenAI, Google, and Anthropic. Still, after months of companies that use AI heavily complaining about skyrocketing token costs, the news of an almost-as-good model at a much lower price really made a splash.
Now, the big AI developers can't just compete by throwing more parameters and training data at the problem. Now they're having to really compete on price, and to do it, OpenAI has massively reduced the price of its models. Not its most powerful and capable — the faster version of that is actually becoming more expensive — but models in its frontier range are now the cheapest they've ever been, and the timeline for this transition of intelligence and pricing is wild.
OpenAI launched ChatGPT 5.4 in March with powerful new agentic capabilities for $2.50 per million input tokens and $15 per million output tokens. GPT 5.6 Luna is now just $0.20 and $1.20, respectively. That's a less-than-four-month window for a frontier model to remain cutting-edge and priced accordingly.
GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price.July 30, 2026
These latest cuts bring Luna into the realm of DeepSeek V4, with its pro model costing $0.435 per million input tokens and $0.87 per million output tokens.
GPT 5.6 Terra is a more capable model, but after its 20% price cut, it's now $2.0 per million input tokens and $12.00 per million output tokens. That undercuts the headline-grabbing K3, which is $3.00 and $15.00, respectively.
Meanwhile, GPT 5.6 Sol remains $5 and $30 per million input/output tokens, and OpenAI has actually raised the price of its top model, with 5.6 Sol in Fast mode charging $10 and $60, respectively, to deliver the same kind of intelligence but at a lower latency — competing directly with other flagship frontier models like Claude Fable 5 and Mythos 5.
But is any of this actually going to make OpenAI any money?
Bills are coming due
After OpenAI announced that it was effectively abandoning its idea of owning first-party data centers earlier this year, the lease contracts it held with Neoclouds became more important than ever. Deals like the enormous $300 billion compute commitment with Oracle became paramount for the very existence of OpenAI's service as a company.
OpenAI has committed to some $600 billion in compute spend by 2030. Even if revenue is rising, it might not be rising anywhere near quickly enough to cover these kinds of bills, and cutting the price of the most popular, affordable models suggests margins will either shrink dramatically or disappear altogether.
This may be why there's also a lot of talk of Nvidia backstopping OpenAI with a $250 billion investment. OpenAI is far from alone here, either. Google spent around nine times its cloud revenue on AI infrastructure over the past year, while Anthropic has only been able to post profits on annualized revenue recently because of a limited cut-price deal with xAI to rent its Colossus data center.
AI is not suddenly cheaper to run or cheaper to build for, and yet companies are slashing prices and making faster, more capable models available for less. On the surface, the numbers just don't add up.
Betting on Jevons Paradox
The AI industry often cites the Jevons Paradox when it comes to accelerating AI adoption and mass-market use. Where in Jevons' time making more efficient coal-powered engines resulted in more coal use, rather than less of it, AI developers claim that as AI use becomes more efficient, greater uses for it will be found, leading to greater overall use.
That may be the future that the token cost-cutting may be hoping to rush us towards. If tokens are cheap, people will use more of them overall, leading to higher earnings. Throw in next-generation AI accelerators becoming more prevalent within AI data centers towards the end of the year, and we could have 10x more tokens per watt, making slimmer margins more profitable by volume.
Then there's Vera Rubin to look forward to, which Nvidia claims will deliver another 10x increase in token performance efficiency. It is certainly possible that the advantages of Blackwell and Vera Rubin platforms will make AI a more potentially profitable industry for inference servers. But even then, it's hard to imagine the big companies covering anything close to their enormous investments with direct AI earnings. Especially as increasing competition drives down token pricing.
Jon Martindale is a contributing writer for Tom's Hardware. For the past 20 years, he's been writing about PC components, emerging technologies, and the latest software advances. His deep and broad journalistic experience gives him unique insights into the most exciting technology trends of today and tomorrow.
Comments (0)