The Token Spend Crunch Finally Hits OpenAI and Anthropic
Frontier labs built trillion-dollar valuations on uncapped AI usage; now finance teams are cutting budgets and routing traffic to cheaper models.
June 30, 2026 | Reading time: 8 minutes | Issue #199
Lead
The era of "tokenmaxxing" is ending. Companies that flooded OpenAI and Anthropic with unlimited inference are now imposing spending caps, switching to Chinese open-weight models, and routing simpler tasks to smaller, cheaper alternatives. CNBC reported on June 26 that Lindy CEO Flo Crivello moved 100 percent of his startup's traffic from Anthropic's Claude to DeepSeek, cutting costs from unsustainable levels to "crash to the ground" territory. Uber capped some employee AI tools at $1,500 per month after CTO Praveen Neppalli Naga revealed in April that the company burned through its entire annual AI budget in four months.
The numbers explain the urgency. Anthropic's annualized run rate reached $47 billion in May, up from roughly $10 billion in revenue for all of 2025, according to CNBC. OpenAI's run rate was pacing near $25 billion earlier this year, up from $13.1 billion in 2025 revenue. Both labs filed confidential IPO paperwork in early June. D.A. Davidson analyst Gil Luria told CNBC that current growth rates are likely the fastest either will ever see, partly because basic math demands that growth slow, and partly because enterprise customers are reining in "out-of-control token spend."
The response from the frontier labs is threefold: cut prices, diversify revenue, and lock in government contracts before the correction deepens. Anthropic's half-price California government deal and OpenAI's Codex hardware teaser both fit the same pattern: find buyers with budgets that do not depend on venture math. The open question is whether lower prices can restore growth faster than cheaper competitors can steal share.
Compute Watch
Google told Meta around March that it could not supply the full Gemini capacity Meta wanted, the Financial Times reported, forcing Meta to delay some internal AI projects and push staff to use tokens more efficiently. The constraint is not unique to Meta; several other Google cloud customers were affected, though less severely. The incident is a reminder that even trillion-dollar cloud providers are capacity-bound on frontier inference.
At the same time, Taiwan is tightening the physical supply chain. Prosecutors in Keelung raided Super Micro's Taiwan offices on June 29, along with the homes of six individuals and three affiliated companies, widening an investigation into alleged smuggling of Nvidia chips into China using Super Micro servers. Taiwan does not currently criminalize AI chip exports to China, so prosecutors are using document-fraud and other existing laws; Taipei is now considering criminalizing the exports directly. The move follows U.S. pressure to stop chip diversion through Taiwanese intermediaries.
Eastern Front
DeepSeek is weaponizing cost and speed at the same time. TechNode reported on June 30 that DeepSeek V4 will launch officially in mid-July with a standard 1-million-token context window, stronger agentic task execution, math, and coding performance, and peak/off-peak API pricing for the first time. Peak hours, 9 a.m. to noon and 2 p.m. to 6 p.m., will cost twice the off-peak rate. The structure turns inference pricing into a utility model, rewarding off-hours batch work and punishing peak-load dependence on Chinese models.
DeepSeek also open-sourced DSpark, a speculative-decoding framework VentureBeat says can raise per-user generation speed by 60 to 85 percent on DeepSeek-V4-Flash and 57 to 78 percent on V4-Pro versus the company's prior MTP-1 baseline. The release includes model checkpoints, a technical paper, and the DeepSpec codebase under an MIT license. DSpark is not an automatic plug-in for every model; the draft module must be aligned to the target. But it gives self-hosting teams another lever to cut inference cost without waiting for a U.S. API provider to lower prices.
Rest of World reported this month that U.S. developers and small companies are increasingly choosing DeepSeek, Minimax, Moonshot's Kimi, and Xiaomi's MiMo for routine tasks. On OpenRouter, DeepSeek, Tencent, Minimax, and Xiaomi are the four most popular models; Vercel said DeepSeek's share of token usage jumped from under 1 percent to 17 percent in May, while its revenue share stayed near 1 percent. The gap between popularity and revenue is the problem for Chinese labs: they are winning usage but not yet converting it into dollars.
India Lens
India's role in the spending correction is as a test market for frugal AI. The June 29 CNBC story on tokenmaxxing opened with a photograph of Indian Prime Minister Narendra Modi flanked by Sam Altman and Dario Amodei at the AI Impact Summit in New Delhi in February. The symbolism is apt: India is the second-largest ChatGPT market after the U.S., and Western labs are racing to serve it while local challengers argue that smaller, sovereign models are cheaper and safer for government and enterprise work.
OpenAI appointed Prabhjeet Singh, former Uber India and South Asia president, as its first managing director for India, starting in September, TechCrunch reported on June 26. The role covers consumer growth, enterprise adoption, partnerships, and regulatory engagement. The move follows OpenAI's New Delhi office opened in August 2025 and planned offices in Mumbai and Bengaluru. Meanwhile, Sarvam AI continues to push Sarvam 2B Omni, which it calls India's first 2-billion-parameter multimodal voice model, and Sarvam Shakti, a domain-specific enterprise LLM family. The bet is that Indian-language, locally hosted models will win contracts that Western frontier APIs are too expensive or too foreign to capture.
Capital Flows
The money is still flowing, but it is moving away from model providers and toward infrastructure, tooling, and sovereign stacks. On June 29, Chamath Palihapitiya's 8090 Labs raised a $135 million Series A led by Salesforce Ventures for Software Factory, an AI coding agent aimed at corporate programming teams with audit trails and governance controls. Palihapitiya is taking the CEO role himself, comparing the moment to early Facebook.
Baz Technologies extended its seed funding to $17 million after launching Baz Planner, a tool that uses specialized agents to analyze code at the planning stage before vulnerabilities enter production, SiliconANGLE reported. Arena, the UC Berkeley-born leaderboard provider, said it reached $100 million in annualized run-rate revenue just eight months after launching its AI Evaluations commercial service, TechCrunch reported. The company clarified the revenue is consumption-based, not recurring.
In the Gulf, 1001 raised a $30 million Series A led by Lux Capital to build sovereign AI operating systems for aviation, ports, energy, and industrial infrastructure, Wamda reported. The company sits above existing operator systems and builds a live working model of the operation, with decisions owned and governed locally. The funding signals that Middle Eastern capital is willing to back applied, sovereign infrastructure AI rather than just buying frontier model access.
Policy & Power
California Governor Gavin Newsom signed a first-of-its-kind deal with Anthropic on June 29 to provide Claude tools to state agencies at roughly half price, Politico reported. The state already uses Claude for the Poppy digital assistant, public engagement platforms, DMV customer service, Medicaid workflows, and a cybersecurity partnership. Newsom's executive order in March was designed to let California separate its AI procurement from the Trump administration's approach. The new Anthropic deal is not framed as a rebuke of Washington, but the timing makes the separation visible.
Meta is also under scrutiny. WIRED reported on June 29 that hundreds of Meta contractors on a project called Cannes, managed by Covalen, posed as minors to prompt rival chatbots including ChatGPT, Gemini, and Character.AI with high-risk subjects including suicide, sex, eating disorders, drugs, and images of pills, knives, and nooses. A single round in August 2025 included more than 45,000 prompts; the rival companies were unaware of the testing. Meta said the project was for safety benchmarking. Regulators may view it differently, especially in the European Union, where child safety and transparency rules are tightening.
Builder's Corner
Meta is limiting how its engineers use Claude Code and OpenAI's Codex, The Information reported, citing internal documents and a follow-up by The Decoder. The concern is distillation: rival model outputs leaking into Meta's training data and accelerating Meta's own coding assistant, MetaCode. Meta has temporarily halted some work with the tools and requires human review. The episode shows that even AI labs that rely on rivals' APIs are now treating those APIs as a contamination risk, not just a cost line.
OpenAI is also reaching toward hardware, but not the Jony Ive device. The Verge reported on June 29 that OpenAI will launch a Codex-focused device on July 15 in partnership with Work Louder, a maker of mechanical keyboards and macro pads. The teaser shows a square device with programmable buttons and dials. It is a peripheral, not a phone, but it points to OpenAI's desire to make Codex a physical workflow layer, not just a chat interface.
The View
Three structural shifts are converging. First, the demand curve is flattening for frontier model APIs as finance teams impose budgets and developers route routine tasks to cheaper models. Second, the supply side is diversifying: Chinese open-weight models, speculative decoding, and sovereign infrastructure stacks are all ways to deliver capability without paying U.S. frontier prices. Third, governance is becoming a competitive dimension: California is using procurement to assert state-level AI independence, Taiwan is using law enforcement to control chip diversion, and Meta is learning that using rival models for safety research can look like surveillance.
The net effect is that the AI market is moving from a single-axis competition on capability to a multi-axis competition on cost, sovereignty, and compliance. OpenAI and Anthropic still lead on frontier benchmarks, but their pricing power is eroding. The labs that survive the correction will be the ones that can make governance and enterprise control a feature, not an afterthought.
The Miss
TIDAL announced on June 29 that it will tag fully AI-generated music with an "AI" badge, remove tracks that impersonate artists or are tied to fraud, and stop paying royalties on music it identifies as wholly AI-generated. The policy takes effect July 15. Music is a small slice of the AI economy, but it is a leading indicator: content platforms are starting to treat generative outputs as a liability class, not a growth category. The major labels will watch closely; if TIDAL's demonetization sticks, Spotify and Apple Music will face pressure to follow.
Pull Quotes
"We did it, and you could see that cost curve go down, like, crash to the ground." — Flo Crivello, Lindy, CNBC, June 26, 2026
"Current growth rates for Anthropic and OpenAI are the fastest they will ever be, which is mostly a matter of basic math." — Gil Luria, D.A. Davidson, CNBC, June 26, 2026
"The GCC runs some of the world's most important infrastructure... Business leaders here don't just want pilots. They want sovereign systems." — Bilal Abu-Ghazaleh, 1001, Wamda, June 30, 2026
"A lot of departments are going to switch their usage to this contract, and that's very much our intent." — Chris Given, California Department of Technology, Politico, June 29, 2026
Reads & Links
OpenAI and Anthropic face new AI reality as users shift from 'tokenmaxxing' to efficiency — CNBC on the spending correction. https://www.cnbc.com/2026/06/26/openai-anthropic-new-ai-spending-reality-as-users-shift-to-efficiency.html
DeepSeek to launch V4 in mid-July with new peak-time API pricing — TechNode on DeepSeek's July launch and pricing structure. https://technode.com/2026/06/30/deepseek-to-launch-v4-in-mid-july-with-new-peak-time-api-pricing/
DeepSeek open sources DSpark, a new framework to speed up LLM inference by up to 85% — VentureBeat on the speculative-decoding release. https://venturebeat.com/orchestration/deepseek-open-sources-dspark-a-new-framework-to-speed-up-llm-inference-by-up-to-85
Google limits Meta's use of its Gemini AI models, FT reports — CNBC on Google's capacity constraints. https://www.cnbc.com/2026/06/28/google-limits-metas-use-of-its-gemini-ai-models-ft-reports.html
Super Micro Office Raided as Taiwan Expands Nvidia Chip Smuggling Probe — Yahoo Finance/Bloomberg on the Taiwan raids. https://finance.yahoo.com/technology/ai/articles/super-micro-office-raided-taiwan-175946955.html
Newsom, Anthropic ink deal to expand government use — Politico on California's statewide Claude contract. https://www.politico.com/news/2026/06/29/exclusive-newsom-anthropic-ink-deal-to-expand-government-use-00979584
Meta Contractors Posed as Teens to Prompt Rival Chatbots About Suicide, Sex, and Drugs — WIRED on the Cannes project. https://www.wired.com/story/meta-contractors-pretending-to-be-teens-chatbot-testing/
OpenAI is teasing new hardware for Codex — The Verge on the July 15 Work Louder device. https://www.theverge.com/ai-artificial-intelligence/959174/openai-codex-hardware-work-louder
When Americans choose Chinese AI — Rest of World on U.S. developers switching to cheaper Chinese models. https://restofworld.org/2026/when-americans-choose-chinese-ai/
Chamath Palihapitiya raises $135M Series A for his AI coding startup, takes CEO role — TechCrunch on 8090 Labs. https://techcrunch.com/2026/06/29/chamath-palihapitiya-raises-135m-series-a-for-his-ai-coding-startup-takes-ceo-role/
Arena, the AI leaderboard everyone uses, is now a $100M business — TechCrunch on Arena's revenue milestone. https://techcrunch.com/2026/06/29/arena-the-ai-leaderboard-everyone-uses-is-now-a-100m-business/
TIDAL cracks down on AI music by cutting off monetization — TechCrunch on TIDAL's AI music policy. https://techcrunch.com/2026/06/29/tidal-cracks-down-on-ai-music-by-cutting-off-monetization/
1001 closes $30 million Series A to build sovereign AI in GCC — Wamda on Middle Eastern sovereign infrastructure AI. https://www.wamda.com/2026/06/1001-closes-30-million-series-a-build-sovereign-ai-gcc
The token spend crunch is reshaping who pays, who routes, and who governs AI.