Observed Signal · Apr 6, 2026 · Industry Analysis · Source: Exponential View · Impact: 3/5 · Sentiment: Negative
AI Compute Demand Sparks an 'AI Capacity Trap'
Cheaper AI inference has increased demand faster than supply can scale, creating a compute crunch. OpenAI’s API token throughput rose from 6 billion tokens/minute in October 2025 to 15 billion tokens/minute by April, a 2.5x increase in five months. Providers including OpenAI and Anthropic are racing to expand compute; Google reports full utilisation across seven generations of TPUs. Anthropic is tightening session limits for Pro users and, despite rising total revenue, is seeing price-per-token fall faster than revenue growth, increasing dependence on volume. Across major AI platforms, usage allowances and tiers tightened, causing customers to face stricter limits and occasional service adjustments. The piece frames this dynamic as an example of the Jevons paradox applied to AI: lower per-token costs spur demand that outstrips available compute capacity.
Rising AI compute demand and platform limits affect availability, pricing and product design for AI-enabled services used across marketing and adtech; impacts on cost, rate limits and service continuity matter to vendors and buyers in the ecosystem.
Track Anthropic Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- OpenAI’s APIs processed 6 billion tokens per minute in October 2025 and 15 billion tokens per minute by April (a 2.5x increase).
- Anthropic has adjusted session limits for its Pro users and is experiencing falling price-per-token even as total revenue grows.
- Google’s TPUs across seven generations are reported to be running at full utilisation.
- Major AI platforms tightened usage allowances last year—adding tiers, stricter limits, and changes often without notice.
Connected Companies & Entities
3 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
US Firms Ration AI Usage as Token Costs Soar
Several large US companies including Amazon, Meta Platforms, Uber and Microsoft are curbing employee use of generative AI tools because computing costs tied to AI 'tokens' have surged. Internal memos and public reporting show some firms exhausting annual token budgets within months, while Google reported processing more than 3.2 trillion AI tokens per month — roughly seven times year‑ago levels. Companies are introducing limits, encouraging cheaper tools, and removing internal usage leaderboards after examples of deliberate overuse (“tokenmaxxing”) and even autonomous bots inflating metrics. Industry observers warn that slower enterprise adoption and rationing could reduce growth for model providers such as Anthropic and OpenAI, while others stress adoption is still in an early phase. Executives and vendors are reassessing controls, budgets and tooling to manage rapidly rising inference costs.
AI Shrinkflation: Providers Quietly Dial Back Models
The article argues that AI providers are quietly reducing model quality, introducing peak/off-peak pricing, throttling capacity, and restricting third-party access as demand outstrips inference capacity and infrastructure costs rise. It cites an AMD AI group analysis that found a ~67% drop in reasoning depth in Claude Code after a February 2026 update and reports an injected consumer-side parameter (reasoning_effort=25) in Anthropic's Claude.ai. The piece links these changes to broader supply constraints (GPU memory shortages, data‑center power bottlenecks) and compares possible futures: consolidation, growth of local inference, or efficiency gains restoring capacity. The author recommends building hybrid cloud/local inference strategies, treating token budgets as real costs, and diversifying provider commitments. Publication date: 2026-06-06.
Newsletter: Tech Industry Faces an AI Compute Crunch
The Exponential View newsletter highlights an emerging "compute crunch" in the AI industry: large generative-AI workloads and cloud demand are creating capacity shortages that force firms to decline business. Examples cited include AWS losing a $10M Fortnite hosting contract due to capacity constraints, OpenAI's CFO saying some opportunities are being passed on for lack of compute, Anthropic tightening session limits (affecting ~7% of users), and H100 GPU rental prices reaching an 18‑month high. The piece also notes related product and commercial moves—Alibaba closed-sourced Qwen, Kuaishou’s Kling AI reported strong video revenue, and GitHub temporarily pulled a Copilot feature after developer backlash. The author connects compute scarcity to broader economic and productivity debates about AI-driven GDP growth and measurement challenges.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
