Observed Signal · Jul 23, 2026 · Partnership · Source: a16z · Impact: 3/5 · Sentiment: Positive
Etched Builds Specialized Hardware for AI Inference
This a16z opinion piece argues that AI inference is becoming the largest, most stable computing workload and therefore will favor specialized hardware over general-purpose GPUs. It highlights Etched, a startup founded by Gavin Uberti, Robert Wachen, and Chris Zhu, which is building full-stack inference systems (chips, boards, interconnects, racks) optimized for tokens-per-watt via "low-voltage inference" and "cluster-scale memory." Etched reportedly taped out first-pass silicon on TSMC's N4P process, has hired 400+ engineers from major hardware companies, and plans to ship its first racks to customers this summer. a16z says it is partnering with Etched.
Describes a startup building full-stack inference hardware that claims first-pass silicon and near-term shipping — relevant to inference cost, tokens-per-watt, and enabling non-hyperscalers to access specialized inference infrastructure.
Track Google Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- In May 2026 Google said it was processing 3.2 quadrillion tokens per month across its products.
- Etched was founded by Gavin Uberti, Robert Wachen, and Chris Zhu who left Harvard to build an inference system from scratch.
- Etched has hired more than 400 engineers from companies including Nvidia, Google’s TPU group, Broadcom, Apple, SK Hynix, and TSMC.
- Etched’s first production chip taped out on TSMC’s N4P process and worked on the first attempt (first-pass silicon).
- Etched plans to ship its first racks to customers in summer 2026 and is standing up a 10 megawatt site; a16z states it is partnering with Etched.
Connected Companies & Entities
8 Entities mapped“In May 2026, Google announced that it was processing 3.2 quadrillion tokens per month across its products, roughly 300 times what it had bee...”
“We see the same story with OpenAI, which now has roughly a billion monthly active users and a rapidly growing coding agent platform in Codex...”
“and with Anthropic, which has seen enormous success with its Claude Code product ......”
“the internet’s trillions of daily packets long ago stopped moving through CPUs and started moving through custom switch chips from companies...”
“Other hyperscalers, like Amazon, Meta, and Microsoft, have followed;...”
“Other hyperscalers, like Amazon, Meta, and Microsoft, have followed;...”
“Etched’s first production chip taped out on TSMC’s N4P process and worked on the first attempt—a “first-pass silicon.”...”
“Etched has hired more than 400 engineers from Nvidia, Google’s TPU group, Broadcom, Apple, SK Hynix, and TSMC....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Etched Hits $5B Valuation, $1B in AI Chip Orders
Etched, an AI chip startup founded in 2022 and positioned as a competitor to Nvidia, said TSMC successfully manufactured its chip earlier in 2026 and the company has booked $1 billion in contract orders for full systems powered by that chip. Etched calls the systems "frontier inference clusters" and is currently testing the first product with customers. The startup has raised $800 million to date, including an unannounced $500 million tranche closed in December at a $5 billion post-money valuation. Notable investors include VentureTech Alliance, Jane Street, Hudson River Trading, Two Sigma and Ribbit Capital; angel backers named include Andrej Karpathy, Geoffrey Hinton and Fei-Fei Li. The announcement comes amid heightened investor interest in inference-accelerating chip technology from other players such as Cerebras, Groq, hyperscalers and OpenAI/Broadcom.
AI Hardware Stack Rebuilt from the Wafer Up
The article explains that modern AI accelerators rely on a constrained hardware stack beginning at wafer fabrication and advanced EUV lithography. TSMC (72% share) and ASML (EUV machines) are central bottlenecks, but the immediate chokepoint is CoWoS packaging for stacking HBM, capacity for which is sold out through 2026. TSMC plans $52–56 billion capex in 2026, yet wafer demand for AI accelerators is projected to rise 11x from 2022–2026. The piece argues GPUs (e.g., NVIDIA H100/B200) are optimized for training and often over-provisioned for latency-sensitive inference. It highlights Cerebras’ wafer-scale WSE-3 (trillions of transistors, massive on-die bandwidth) and cited benchmarks showing material inference throughput and cost advantages versus NVIDIA B200. The article notes OpenAI signed a $20B+ agreement with Cerebras for large-scale inference capacity and recommends builders benchmark their own workloads on emerging inference hardware.
Multi‑Silicon Era: Disaggregated AI Inference Emerges
At GTC 2026 NVIDIA unveiled a broad set of inference infrastructure updates and new systems, showcasing a multi-silicon, disaggregated inference strategy. Announcements include three new systems (Groq LPX, Vera ETL256, and STX), updates to the Kyber rack family, and multi-rack world-size SKUs such as Rubin Ultra NVL576 and the planned Feynman NVL1152. NVIDIA paid Groq $20B to license Groq IP and hire most of its team, enabling rapid integration of Groq LPUs (Groq LPU 3 / LP30 and refresh LP35) into NVIDIA’s Vera/Rubin stacks; NVIDIA plans an LP40 on TSMC N3P with CoWoS-R and NVLink support. The company also promoted Attention and FFN Disaggregation (AFD) using LPUs for low-latency decode, described LPX rack architecture (LPUs + Fabric Expansion Logic FPGAs), outlined Vera ETL256 (256-CPU rack) and STX/CMX storage rack designs, and issued a CPO/optics roadmap for large world-size scale-up.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
