Observed Signal · Apr 20, 2026 · Interview · Source: Chipstrat · Impact: 4/5 · Sentiment: Positive
Meta VP Matt Steiner on Ads Infrastructure
Meta VP Matt Steiner explains how Meta’s ads stack drives hardware and software design across recommender systems and generative AI. He describes a two-stage ad-serving pipeline: retrieval (powered by Andromeda, running on a co‑designed NVIDIA Grace Hopper SKU) and ranking (consolidated into a single model called Lattice). Meta trained a large foundation model called GEM and distilled it into a servable adaptive ranking model at roughly one trillion parameters that operates at sub‑second latency. Recommender workloads are memory‑bound with a different compute‑to‑memory profile than standard LLM GPUs, motivating Meta’s MTIA custom silicon work. Steiner also highlights LLM‑written kernels (e.g., KernelEvolve / Alpha Evolve) to automate hardware‑specific optimizations, and predicts rising demand for many more optimized kernels per chip as Meta’s heterogeneous fleet grows. He emphasizes end‑to‑end co‑optimization of hardware, networking, models and software over the next two years.
Provides substantive, technical insight from a Meta VP about large‑scale ad infrastructure, model consolidation (Lattice), a foundation model (GEM) distilled into a 1T‑parameter servable ranking model, custom silicon (MTIA) and hardware–software co‑design — information relevant to industry infrastructure, hardware vendors and AdTech strategy.
Track Meta Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Matt Steiner is VP of Monetization Infrastructure, Ranking & AI Foundations at Meta.
- Meta separates ad serving into retrieval (Andromeda) and ranking (Lattice) stages.
- Meta trained GEM (Generative Ads Recommendation foundation model) and distilled it into an adaptive ranking model served at roughly one trillion parameters with sub‑second latency.
- Recommender retrieval workloads at Meta are memory‑bound; Meta co‑designed a custom NVIDIA Grace Hopper SKU to meet retrieval memory/bandwidth needs.
- Meta is developing Meta Training and Inference Accelerators (MTIA) and uses LLM‑written kernels (Alpha Evolve / KernelEvolve) to optimize software across heterogeneous hardware.
Connected Companies & Entities
3 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Meta's ROIC Strategy: GEM Now, LLMs Later
ChipStrat analyzes Meta’s capital allocation approach: prioritize GEM (Generative Ads Recommendation Model) now to generate immediate, measurable ROI in ad ranking and monetization, while investing in frontier LLMs later for layered upside. Meta treats GEM as a large-scale teacher model that transfers knowledge to smaller, latency-sensitive serving models, keeping inference costs low. The company reports concrete ad-performance gains tied to recent model and infrastructure investments (+5% Instagram conversions; +3% Facebook Feed conversions; +3.5% Facebook ad clicks in Q4). Meta is also diversifying compute (NVIDIA, AMD, and custom MTIA silicon) and unifying ranking across paid and organic content. Concurrently, Meta funds Superintelligence Labs (co-led by Alexandr Wang and Nat Friedman) to build frontier LLMs that could further enhance recommendations, creative generation, and content localization.
Meta Unveils Custom AI Chips Amid Nvidia, AMD Partnerships
Meta revealed four custom in-house AI chips in its MTIA (Meta Training and Inference Accelerator) family as part of a rapid data-center expansion. Meta has deployed MTIA 300 (for training smaller models used in ranking, recommendations and ad delivery) and completed testing of MTIA 400, which is optimized for generative-AI inference and is slated for near-term deployment; MTIA 450 and MTIA 500 are planned to be operational in 2027. Meta said the chips are manufactured by Taiwan Semiconductor and that one data-center rack will hold 72 MTIA 400 chips. Meta framed the custom silicon as a way to improve price/performance, diversify silicon supply and hedge against vendor price changes, while noting concerns about securing high-bandwidth memory (HBM). The company has also signed large multi-year deals for Nvidia and AMD GPUs to preserve options.
Meta's Kunlun Recommender Shows Predictable Scaling
This Import AI issue surveys recent AI research and benchmarks, with a focus on Meta/Facebook’s new Kunlun recommender architecture and its discovered scaling laws. Meta describes Kunlun’s modular design (Transformer and Interaction blocks) that raises Model FLOPs Utilization (MFU) from 17% to 37% on NVIDIA B200 GPUs and demonstrates predictable power-law scaling in recommender performance measured by normalized entropy (NE). Meta reports deployment of Kunlun across major Meta Ads models with a reported 1.2% improvement in topline metrics. The newsletter also summarizes AIRS-BENCH (a 20-task benchmark from Meta and UK universities) showing current agents trail best-in-class humans, First Proof (a sealed 10-question frontier-math benchmark from multiple universities) which state-of-the-art models currently fail in one-shot settings, and a Nick Bostrom paper arguing trade-offs in timing development of superintelligence.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
