Observed Signal · May 29, 2026 · Industry Analysis · Source: CNBC Technology · Impact: 3/5 · Sentiment: Negative

Enterprise AI: Tokens vs. Humans Trade-off

Executive Signal Summary

CFOs at large U.S. companies are confronting a new budget dilemma as AI inference costs surge, forcing a choice between spending on model tokens or hiring staff. CNBC spoke with Arvind Jain (CEO of Glean) and Matan Grinberg (CEO of Factory AI), who described how many enterprises are exhausting annual AI budgets within months, with roughly 95% of usage still routed to the most expensive frontier models. Companies are moving from a phase of ‘tokenmaxxing’ to reassessing whether premium models are needed for every task; routing simpler work to cheaper model tiers could yield substantial savings. The article cautions that demand may be more price‑sensitive than market assumptions, with implications for valuations and revenue growth of premium model providers like OpenAI and Anthropic.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Rising inference/token costs can materially affect enterprise AI adoption, vendor revenues and valuations; enterprises' budget responses (rationing or routing to cheaper models) could alter demand dynamics across the AI model ecosystem.

SIGNAL RADAR

Track Anthropic Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Published on 2026-05-29.
  • Arvind Jain, CEO of Glean, said enterprise AI budgets are often exhausted within one to two months and that each new frontier model release is roughly twice as expensive per token as its predecessor.
  • Approximately 95% of enterprise AI usage still runs on the priciest frontier models, even for tasks that could use cheaper alternatives.
  • Routing simple tasks to cheaper model tiers can deliver large savings (Jain cited up to 10x savings with proper model routing).
  • Matan Grinberg, CEO of Factory AI, described the issue as a resource-allocation decision between headcount and AI spend per employee and outlined three adoption phases including 'tokenmaxxing' and subsequent cost reassessment.

Ontology Mapping & Concepts

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: CNBC Technology•Published: May 29, 2026
Original Coverage Title: “Tokens or humans? The new corporate trade-off”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIMay 29, 2026

US Firms Ration AI Usage as Token Costs Soar

Several large US companies including Amazon, Meta Platforms, Uber and Microsoft are curbing employee use of generative AI tools because computing costs tied to AI 'tokens' have surged. Internal memos and public reporting show some firms exhausting annual token budgets within months, while Google reported processing more than 3.2 trillion AI tokens per month — roughly seven times year‑ago levels. Companies are introducing limits, encouraging cheaper tools, and removing internal usage leaderboards after examples of deliberate overuse (“tokenmaxxing”) and even autonomous bots inflating metrics. Industry observers warn that slower enterprise adoption and rationing could reduce growth for model providers such as Anthropic and OpenAI, while others stress adoption is still in an early phase. Executives and vendors are reassessing controls, budgets and tooling to manage rapidly rising inference costs.

Read assessment
Large Language Models (LLM) & AIJun 26, 2026

OpenAI, Anthropic Face Shift From Tokenmaxxing to Efficiency

Enterprise customers are reining in runaway AI token spending and shifting from 'tokenmaxxing' toward cost-efficient model use, putting pressure on leading model providers OpenAI and Anthropic. Startups and enterprises are routing tasks to cheaper models, switching providers (one startup moved all traffic from Anthropic to Chinese firm DeepSeek), and implementing usage caps and analytics to control bills. The shift comes as both Anthropic and OpenAI report multibillion-dollar annualized run rates and weigh IPO timing; investors and analysts note urgency to list before corporate customers rationalize AI spend. Big cloud and platform vendors — Microsoft, Amazon and Google — are promoting lower-cost alternatives and model-routing features, intensifying competition. Vendors have added enterprise spend controls and analytics, while finance leaders and consultants urge proving ROI before large-scale deployments.

Read assessment
Large Language Models (LLM) & AIJul 5, 2026

Companies Cut AI Costs with 'Modelmaxxing' Strategy

The article reports a shift in corporate AI usage from indiscriminate high-cost model usage (“tokenmaxxing”) toward a more targeted approach called “modelmaxxing,” where teams pick models by task complexity to reduce inference expenses. Tokenmaxxing reportedly produced extreme consumption at some tech firms — The Information found an internal Meta leaderboard with about 60 trillion tokens in 30 days and a top user consuming ~280 billion tokens — potentially costing hundreds of thousands to millions of dollars. Companies such as Meta and Amazon helped popularize heavy token use. In response, firms and developers (e.g., Bold Metrics’ CTO Morgan Linton and developer Alejandra Thomas) are prescribing specific models for tasks. Model-routing startups have emerged and Ramp’s chief economist reports adoption rising from ~1% to ~5% of companies; a Bitkom survey found about one-third of German firms were surprised by AI costs. The trend aims to keep AI benefits while controlling spend.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.