Observed Signal · Mar 31, 2026 · Technical Release · Source: SemiAnalysis · Impact: 2/5 · Sentiment: Neutral
SemiAnalysis Microbenchmarks Nvidia Blackwell Tensor Cores
SemiAnalysis published an in-depth microbenchmarking study of Nvidia’s datacenter Blackwell (SM100) GPU, analyzing PTX and SASS instruction performance for AI workloads. The report measures low-level behavior and practical upper bounds for tensor-related primitives (UMMA, TMA, cp.async/LDGSTS), TMEM, DSMEM, multicast, cluster scheduling, and new 2SM MMA (.cta_group::2) instructions. Key empirical findings include memory throughput and latency curves for async copy and TMA, a measured die-to-die L2 penalty (~300 cycles), and shape- and layout-dependent MMA throughput and latency (including cases where M=64 is SMEM‑bound at ~50% peak). The authors open-sourced their benchmarking repo, acknowledge cloud node providers and reviewers, and plan further kernel benchmarking across Blackwell variants and other AI accelerators (TPU Pallas, Trainium, AMD CDNA4).
Provides actionable, low-level performance measurements for Nvidia Blackwell GPUs that matter to ML systems and kernel developers; useful for infrastructure teams but only indirectly relevant to the broader AdTech industry.
Track NVIDIA Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- SemiAnalysis open-sourced a Blackwell micro-architecture benchmarking repository covering PTX and SASS instructions (UMMA, TMA, cp.async).
- Async copy (LDGSTS / cp.async) throughput saturates at about 6.6 TB/s at ~32 KiB bytes-in-flight and has a baseline latency ~600 ns that roughly doubles after 8 KiB in flight.
- TMA (cp.async.bulk.tensor / UTMALDG) scales to much larger bytes-in-flight (up to at least 128 KiB) and catches up to or exceeds async copy throughput beyond ~32 bytes-in-flight; TMA latency grows sharply after ~12 KiB.
- Measured die-to-die L2 crossing penalty on the tested B200 hardware is roughly 300 cycles, and logical GPC/cluster mappings can vary between chips causing potential non-determinism.
- Blackwell 2SM MMA (.cta_group::2) shows near-perfect weak scaling (up to ~2x) versus 1SM MMA; 1SM MMA with M=64 hits ~50% of theoretical peak while M=128+ approaches near 100% for many shapes.
Connected Companies & Entities
3 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
InferenceX v2 Benchmarks Blackwell vs AMD & Hopper
SemiAnalysis released InferenceX v2 (formerly InferenceMAX), an open-source Apache 2.0 continuous inference benchmark that expands coverage across ~1,000 frontier GPUs and new distributed inference modes. InferenceXv2 adds large-scale disaggregated prefill (disagg) with wide expert parallelism (wideEP) testing for six recent NVIDIA GPU SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and all recent AMD western SKUs including MI355X. The release includes the first third‑party Pareto-frontier benchmarks for Blackwell Ultra GB300 NVL72 and multi-node MI355X disagg+wideEP FP4/FP8. Key findings: NVIDIA Blackwell rack-scale systems lead for MoE/disaggregated inference and energy efficiency; AMD MI355X is competitive on some FP8 and single-node perf/TCO but suffers composability and FP4 multi-node software gaps. The report also highlights MTP (multi-token/speculative decoding) as a major cost reducer and documents software stacks such as SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI and Mooncake.
2026 GPU Comparison: NVIDIA, AMD, Intel for AI
This article evaluates workstation and prosumer GPUs for local LLM inference and AI workloads in mid-2026, comparing NVIDIA's Blackwell (RTX 50-series), AMD's Radeon AI Pro R9700, and Intel's Arc Pro B70. It argues that VRAM capacity, memory bandwidth, and software ecosystem maturity matter more than peak theoretical compute (AI TOPS) for real-world transformer inference. The piece provides recommended VRAM ranges for common model sizes, a complete spec and price table for relevant consumer and professional cards, and practical guidance on power, thermal behavior, form factor, PCIe bandwidth, and multi-GPU considerations. Conclusions highlight NVIDIA's Blackwell family as the inference benchmark due to bandwidth and CUDA/TensorRT maturity, AMD's R9700 as a value workstation option with ROCm support, and Intel's B70 as an affordable 32 GB workstation GPU with a maturing oneAPI ecosystem.
NVIDIA Vera Rubin NVL72 Inference TCO Analysis
SemiAnalysis analyzes early engineering-sample metrics and architectural changes for NVIDIA's Vera Rubin NVL72 (Oberon SM_107) and compares its inference performance and total cost of ownership (TCO) against GB200 and GB300 NVL72 baselines. CoreWeave-reported results using DeepSeek R1 show large per-MW and per-dollar gains (CoreWeave claims ~5.4x per-MW and ~5x per-dollar vs a 2025 GB200 baseline), but SemiAnalysis highlights benchmarking nuances: different baselines (2025 vs 2026), single-turn workloads, pre-production rack hardware, and unverified metrics. Rubin's architectural changes include larger SMEM/TMEM configurations, inline TMA descriptor overrides, doubled FP8/FP4 tensor throughput, 2:4 activation sparsity, a 3-bit LUT B tensor-core decompression mode, and 3D-stacked HBM4 giving ~2.8x global memory bandwidth versus Blackwell. SemiAnalysis reports Rubin TCO per GPU of $3.57 (operator ownership) versus $1.84 for GB200 and $2.36 for GB300, and notes NVIDIA plans to submit verifiable InferenceX numbers by Q3 CY2026.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
