Observed Signal · Mar 31, 2026 · Technical Release · Source: SemiAnalysis · Impact: 2/5 · Sentiment: Neutral

SemiAnalysis Microbenchmarks Nvidia Blackwell Tensor Cores

Executive Signal Summary

SemiAnalysis published an in-depth microbenchmarking study of Nvidia’s datacenter Blackwell (SM100) GPU, analyzing PTX and SASS instruction performance for AI workloads. The report measures low-level behavior and practical upper bounds for tensor-related primitives (UMMA, TMA, cp.async/LDGSTS), TMEM, DSMEM, multicast, cluster scheduling, and new 2SM MMA (.cta_group::2) instructions. Key empirical findings include memory throughput and latency curves for async copy and TMA, a measured die-to-die L2 penalty (~300 cycles), and shape- and layout-dependent MMA throughput and latency (including cases where M=64 is SMEM‑bound at ~50% peak). The authors open-sourced their benchmarking repo, acknowledge cloud node providers and reviewers, and plan further kernel benchmarking across Blackwell variants and other AI accelerators (TPU Pallas, Trainium, AMD CDNA4).

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Provides actionable, low-level performance measurements for Nvidia Blackwell GPUs that matter to ML systems and kernel developers; useful for infrastructure teams but only indirectly relevant to the broader AdTech industry.

SIGNAL RADAR

Track NVIDIA Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • SemiAnalysis open-sourced a Blackwell micro-architecture benchmarking repository covering PTX and SASS instructions (UMMA, TMA, cp.async).
  • Async copy (LDGSTS / cp.async) throughput saturates at about 6.6 TB/s at ~32 KiB bytes-in-flight and has a baseline latency ~600 ns that roughly doubles after 8 KiB in flight.
  • TMA (cp.async.bulk.tensor / UTMALDG) scales to much larger bytes-in-flight (up to at least 128 KiB) and catches up to or exceeds async copy throughput beyond ~32 bytes-in-flight; TMA latency grows sharply after ~12 KiB.
  • Measured die-to-die L2 crossing penalty on the tested B200 hardware is roughly 300 cycles, and logical GPC/cluster mappings can vary between chips causing potential non-determinism.
  • Blackwell 2SM MMA (.cta_group::2) shows near-perfect weak scaling (up to ~2x) versus 1SM MMA; 1SM MMA with M=64 hits ~50% of theoretical peak while M=128+ approaches near 100% for many shapes.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: SemiAnalysis•Published: Mar 31, 2026
Original Coverage Title: “Dissecting Nvidia Blackwell - Tensor Cores, PTX Instructions, SASS, Floorsweep, Yield”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIFeb 16, 2026

InferenceX v2 Benchmarks Blackwell vs AMD & Hopper

SemiAnalysis released InferenceX v2 (formerly InferenceMAX), an open-source Apache 2.0 continuous inference benchmark that expands coverage across ~1,000 frontier GPUs and new distributed inference modes. InferenceXv2 adds large-scale disaggregated prefill (disagg) with wide expert parallelism (wideEP) testing for six recent NVIDIA GPU SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and all recent AMD western SKUs including MI355X. The release includes the first third‑party Pareto-frontier benchmarks for Blackwell Ultra GB300 NVL72 and multi-node MI355X disagg+wideEP FP4/FP8. Key findings: NVIDIA Blackwell rack-scale systems lead for MoE/disaggregated inference and energy efficiency; AMD MI355X is competitive on some FP8 and single-node perf/TCO but suffers composability and FP4 multi-node software gaps. The report also highlights MTP (multi-token/speculative decoding) as a major cost reducer and documents software stacks such as SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI and Mooncake.

Read assessment
Large Language Models (LLM) & AIJul 14, 2026

2026 GPU Comparison: NVIDIA, AMD, Intel for AI

This article evaluates workstation and prosumer GPUs for local LLM inference and AI workloads in mid-2026, comparing NVIDIA's Blackwell (RTX 50-series), AMD's Radeon AI Pro R9700, and Intel's Arc Pro B70. It argues that VRAM capacity, memory bandwidth, and software ecosystem maturity matter more than peak theoretical compute (AI TOPS) for real-world transformer inference. The piece provides recommended VRAM ranges for common model sizes, a complete spec and price table for relevant consumer and professional cards, and practical guidance on power, thermal behavior, form factor, PCIe bandwidth, and multi-GPU considerations. Conclusions highlight NVIDIA's Blackwell family as the inference benchmark due to bandwidth and CUDA/TensorRT maturity, AMD's R9700 as a value workstation option with ROCm support, and Intel's B70 as an affordable 32 GB workstation GPU with a maturing oneAPI ecosystem.

Read assessment
Large Language Models (LLM) & AIJul 23, 2026

NVIDIA Vera Rubin NVL72 Inference TCO Analysis

SemiAnalysis analyzes early engineering-sample metrics and architectural changes for NVIDIA's Vera Rubin NVL72 (Oberon SM_107) and compares its inference performance and total cost of ownership (TCO) against GB200 and GB300 NVL72 baselines. CoreWeave-reported results using DeepSeek R1 show large per-MW and per-dollar gains (CoreWeave claims ~5.4x per-MW and ~5x per-dollar vs a 2025 GB200 baseline), but SemiAnalysis highlights benchmarking nuances: different baselines (2025 vs 2026), single-turn workloads, pre-production rack hardware, and unverified metrics. Rubin's architectural changes include larger SMEM/TMEM configurations, inline TMA descriptor overrides, doubled FP8/FP4 tensor throughput, 2:4 activation sparsity, a 3-bit LUT B tensor-core decompression mode, and 3D-stacked HBM4 giving ~2.8x global memory bandwidth versus Blackwell. SemiAnalysis reports Rubin TCO per GPU of $3.57 (operator ownership) versus $1.84 for GB200 and $2.36 for GB300, and notes NVIDIA plans to submit verifiable InferenceX numbers by Q3 CY2026.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.