Observed Signal · Sep 7, 2026 · Technical Release · Source: SemiAnalysis · Impact: 5/5 · Sentiment: Positive

Google TPUv7 Ironwood Delivers Up to 50% Better Performance per Dollar

Executive Signal Summary

SemiAnalysis, an independent analyst firm, published third-party inference benchmarks for Google's TPUv7 Ironwood, comparing it against NVIDIA's B200 and B300 GPUs. The results show Ironwood delivering up to 50% better performance per dollar in apples-to-apples comparisons, with a 19-34% lower cost per token at typical interactivity levels. Key factors include Google's co-designed hardware and software stack, the emerging TorchTPU external stack providing native PyTorch support, and optimizations across kernels and serving engines (vLLM, SGLang). Google is externalizing its TPU infrastructure, selling chips outright and on Google Cloud, with Anthropic as a major customer. The TorchTPU stack is in private beta, open-sourcing around mid-October, and future work includes speculative decoding, disaggregated serving, and KV-cache offloading, positioning TPUs as a strong competitor in the AI inference market.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Major technical release from Google showing TPUv7 Ironwood outperforming NVIDIA GPUs on performance per dollar, with significant implications for AI inference cost structure and the competitive landscape.

SIGNAL RADAR

Track Google Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Google's TPUv7 Ironwood delivers up to 50% better performance per dollar compared to NVIDIA B200/B300 GPUs.
  • Ironwood costs approximately $0.181 per million tokens at 100 tokens/s/user, vs $0.222 for B200 and $0.276 for B300.
  • Anthropic has committed to over one million TPUs, including 400k+ direct purchases and 600k+ rented via GCP.
  • The TorchTPU stack, enabling native PyTorch on TPUs, is in private beta and will be open-sourced around mid-October.
  • TPUv8i Boardfly, with native FP4 support, is expected to be competitive with NVIDIA's Rubin NVL72.

Connected Companies & Entities

7 Entities mapped

“Google has been running disaggregated serving internally and is externalizing its TPU stack....”

“Ironwood delivers up to 50% better performance per dollar compared to B200/B300 from NVIDIA....”

“SemiAnalysis published third-party inference results for TPUv7 Ironwood on InferenceX....”

“Anthropic is the biggest user of TPUs, surpassing Deepmind’s own use by 2029....”

“vLLM and SGLang are the two major open-source production inference serving stacks, and TorchTPU will be the go-to backend for them going for...”

“Red Hat is investing heavily in collaboration with Google to make the TorchTPU backend a first-class experience....”

“SGLang's current public TPU stack uses SGLang-JAX, and TorchTPU will be the go-to backend for SGLang on TPUs going forward....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: SemiAnalysis•Published: Sep 7, 2026
Original Coverage Title: “TPU Inference Externalization Full Steam Ahead - InferenceX”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

InfrastructureNov 28, 2025

Google Commercializes TPUv7, Challenging Nvidia

SemiAnalysis describes Google’s shift from internal TPU use to commercializing TPUv7 (Ironwood) hardware and renting systems via GCP, highlighting Anthropic’s confirmed 1 million TPU order split between direct purchases and GCP rentals. The report argues TPUv7 closes much of the performance gap with recent Nvidia GPUs while offering materially lower total cost of ownership (TCO) in many training workloads, and details the TPUv7 hardware, 3D-torus ICI scale-up network, OCS-based optical routing, software stack changes (native PyTorch support, Pallas kernel support), and ecosystem partners (Broadcom, Fluidstack, TeraWulf, Cipher Mining). SemiAnalysis frames this as a meaningful merchant-silicon challenge to Nvidia, with implications for datacenter power, neocloud hosting, supply chains, and open-source compiler/runtime adoption required to broaden TPU external adoption.

Read assessment
Large Language Models (LLM) & AIApr 22, 2026

Google launches separate TPUs for training and inference

Google Cloud announced its eighth-generation custom Tensor Processing Units (TPUs), splitting the family into two purpose-built chips: the TPU 8t for model training and the TPU 8i for inference. Google claims up to ~2.8–3x faster training versus prior generation Ironwood at comparable price, about 80% better performance per dollar on inference workloads, and the ability to cluster more than one million TPUs. Google said the TPUs will supplement — not immediately replace — Nvidia GPU offerings in its cloud, and that Nvidia’s Vera Rubin GPU will be available in Google Cloud later this year. Google also disclosed a collaboration with Nvidia to improve software-based networking (Falcon) for more efficient Nvidia system performance; Falcon was open sourced in 2023 under the Open Compute Project. The move positions Google’s cloud hardware as an alternative compute path for large AI workloads while maintaining interoperability with Nvidia-based stacks.

Read assessment
Large Language Models (LLM) & AIJun 27, 2026

Google sharpens TPU advantage in AI compute race

Alphabet’s homegrown tensor processing units (TPUs) are gaining prominence as a cost- and energy-efficient alternative to Nvidia GPUs, powering Google’s Gemini models and fueling Google Cloud’s enterprise growth. Google announced eighth-generation TPUs with distinct variants for training (TPU 8t) and inference (TPU 8i), claiming up to 3x faster training and 80% better performance-per-dollar, and has expanded commercialization—renting TPUs via cloud, selling hardware to customers, and launching a TPU cloud joint venture with Blackstone. Major AI labs and enterprises, including Anthropic and Meta, are adopting TPU capacity. Analysts and executives say TPU monetization and efficiency advantages could materially accelerate Google Cloud revenue and shift compute economics in the AI era.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.