Observed Signal · Apr 6, 2026 · Industry Analysis · Source: Exponential View · Impact: 3/5 · Sentiment: Negative

AI Compute Demand Sparks an 'AI Capacity Trap'

Executive Signal Summary

Cheaper AI inference has increased demand faster than supply can scale, creating a compute crunch. OpenAI’s API token throughput rose from 6 billion tokens/minute in October 2025 to 15 billion tokens/minute by April, a 2.5x increase in five months. Providers including OpenAI and Anthropic are racing to expand compute; Google reports full utilisation across seven generations of TPUs. Anthropic is tightening session limits for Pro users and, despite rising total revenue, is seeing price-per-token fall faster than revenue growth, increasing dependence on volume. Across major AI platforms, usage allowances and tiers tightened, causing customers to face stricter limits and occasional service adjustments. The piece frames this dynamic as an example of the Jevons paradox applied to AI: lower per-token costs spur demand that outstrips available compute capacity.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Rising AI compute demand and platform limits affect availability, pricing and product design for AI-enabled services used across marketing and adtech; impacts on cost, rate limits and service continuity matter to vendors and buyers in the ecosystem.

SIGNAL RADAR

Track Anthropic Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • OpenAI’s APIs processed 6 billion tokens per minute in October 2025 and 15 billion tokens per minute by April (a 2.5x increase).
  • Anthropic has adjusted session limits for its Pro users and is experiencing falling price-per-token even as total revenue grows.
  • Google’s TPUs across seven generations are reported to be running at full utilisation.
  • Major AI platforms tightened usage allowances last year—adding tiers, stricter limits, and changes often without notice.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: Exponential View•Published: Apr 6, 2026
Original Coverage Title: “📈 Data to start your week: The AI capacity trap”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIMay 29, 2026

US Firms Ration AI Usage as Token Costs Soar

Several large US companies including Amazon, Meta Platforms, Uber and Microsoft are curbing employee use of generative AI tools because computing costs tied to AI 'tokens' have surged. Internal memos and public reporting show some firms exhausting annual token budgets within months, while Google reported processing more than 3.2 trillion AI tokens per month — roughly seven times year‑ago levels. Companies are introducing limits, encouraging cheaper tools, and removing internal usage leaderboards after examples of deliberate overuse (“tokenmaxxing”) and even autonomous bots inflating metrics. Industry observers warn that slower enterprise adoption and rationing could reduce growth for model providers such as Anthropic and OpenAI, while others stress adoption is still in an early phase. Executives and vendors are reassessing controls, budgets and tooling to manage rapidly rising inference costs.

Read assessment
Large Language Models (LLM) & AIJun 6, 2026

AI Shrinkflation: Providers Quietly Dial Back Models

The article argues that AI providers are quietly reducing model quality, introducing peak/off-peak pricing, throttling capacity, and restricting third-party access as demand outstrips inference capacity and infrastructure costs rise. It cites an AMD AI group analysis that found a ~67% drop in reasoning depth in Claude Code after a February 2026 update and reports an injected consumer-side parameter (reasoning_effort=25) in Anthropic's Claude.ai. The piece links these changes to broader supply constraints (GPU memory shortages, data‑center power bottlenecks) and compares possible futures: consolidation, growth of local inference, or efficiency gains restoring capacity. The author recommends building hybrid cloud/local inference strategies, treating token budgets as real costs, and diversifying provider commitments. Publication date: 2026-06-06.

Read assessment
Large Language Models & AI infrastructureApr 5, 2026

Newsletter: Tech Industry Faces an AI Compute Crunch

The Exponential View newsletter highlights an emerging "compute crunch" in the AI industry: large generative-AI workloads and cloud demand are creating capacity shortages that force firms to decline business. Examples cited include AWS losing a $10M Fortnite hosting contract due to capacity constraints, OpenAI's CFO saying some opportunities are being passed on for lack of compute, Anthropic tightening session limits (affecting ~7% of users), and H100 GPU rental prices reaching an 18‑month high. The piece also notes related product and commercial moves—Alibaba closed-sourced Qwen, Kuaishou’s Kling AI reported strong video revenue, and GitHub temporarily pulled a Copilot feature after developer backlash. The author connects compute scarcity to broader economic and productivity debates about AI-driven GDP growth and measurement challenges.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.