Observed Signal · Feb 16, 2026 · Technical Release · Source: SemiAnalysis · Impact: 3/5 · Sentiment: Positive

InferenceX v2 Benchmarks Blackwell vs AMD & Hopper

Executive Signal Summary

SemiAnalysis released InferenceX v2 (formerly InferenceMAX), an open-source Apache 2.0 continuous inference benchmark that expands coverage across ~1,000 frontier GPUs and new distributed inference modes. InferenceXv2 adds large-scale disaggregated prefill (disagg) with wide expert parallelism (wideEP) testing for six recent NVIDIA GPU SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and all recent AMD western SKUs including MI355X. The release includes the first third‑party Pareto-frontier benchmarks for Blackwell Ultra GB300 NVL72 and multi-node MI355X disagg+wideEP FP4/FP8. Key findings: NVIDIA Blackwell rack-scale systems lead for MoE/disaggregated inference and energy efficiency; AMD MI355X is competitive on some FP8 and single-node perf/TCO but suffers composability and FP4 multi-node software gaps. The report also highlights MTP (multi-token/speculative decoding) as a major cost reducer and documents software stacks such as SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI and Mooncake.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Open-source, large-scale inference benchmark expands cross-vendor, multi-node MoE and disaggregated testing (including first third-party Blackwell Ultra and MI355X disagg+wideEP results). Results influence LLM inference TCO, software priorities (composability), and vendor selection for production inference infrastructure.

SIGNAL RADAR

Track NVIDIA Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • SemiAnalysis published InferenceX v2, an open-source (Apache 2.0) continuous inference benchmark.
  • InferenceXv2 benchmarks nearly 1,000 GPUs across NVIDIA SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and recent AMD SKUs (including MI355X).
  • InferenceXv2 is the first third-party suite to benchmark Blackwell Ultra GB300 NVL72 across the Pareto frontier and to test MI355X disagg+wideEP multi-node FP4 and FP8 performance.
  • Findings: NVIDIA Blackwell systems lead in large-scale MoE disaggregated inference and energy efficiency; AMD MI355X is competitive on FP8 single-node and some FP8 disagg configs but faces composability and FP4 multi-node software limitations.
  • Speculative decoding / MTP materially reduces cost per token across many configurations; the report documents production software stacks (SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI, Mooncake).
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: SemiAnalysis•Published: Feb 16, 2026
Original Coverage Title: “InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper - Formerly InferenceMAX”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

InfrastructureMar 31, 2026

SemiAnalysis Microbenchmarks Nvidia Blackwell Tensor Cores

SemiAnalysis published an in-depth microbenchmarking study of Nvidia’s datacenter Blackwell (SM100) GPU, analyzing PTX and SASS instruction performance for AI workloads. The report measures low-level behavior and practical upper bounds for tensor-related primitives (UMMA, TMA, cp.async/LDGSTS), TMEM, DSMEM, multicast, cluster scheduling, and new 2SM MMA (.cta_group::2) instructions. Key empirical findings include memory throughput and latency curves for async copy and TMA, a measured die-to-die L2 penalty (~300 cycles), and shape- and layout-dependent MMA throughput and latency (including cases where M=64 is SMEM‑bound at ~50% peak). The authors open-sourced their benchmarking repo, acknowledge cloud node providers and reviewers, and plan further kernel benchmarking across Blackwell variants and other AI accelerators (TPU Pallas, Trainium, AMD CDNA4).

Read assessment
Large Language Models (LLM) & AIAug 25, 2026

Qwen3-8B inference benchmark and FP8 on Blackwell

Independent benchmarks compare Qwen3-8B inference on an RTX PRO 6000 Blackwell (96 GB) across three serving stacks (vLLM 0.27.1, SGLang 0.5.9, and llama.cpp CUDA). At concurrency 32 using BF16, vLLM achieved 1,725 aggregate tokens/s (TTFT p50 39 ms), SGLang 1,327 tok/s (TTFT p50 42 ms), and llama.cpp 428 tok/s (TTFT p50 316 ms). Applying an FP8 checkpoint to vLLM increased throughput by ~1.5x (aggregate 1,725 -> 2,597 tok/s; single-stream 86 -> 130 tok/s) with lower latency and no detected regressions on a fixed factual check. The author documents methodology, reproductions, and an sm_120-specific kernel workaround required to run FP8 on workstation Blackwell hardware.

Read assessment
InfrastructureAug 24, 2026

AgentX 1.0: Open-source Agentic Inference Benchmark

SemiAnalysis announced AgentX 1.0 and InferenceXv3, an open-source agentic inference benchmark and benchmark implementation designed for long-context, multi-turn agentic coding workloads up to 1M context (Apache 2.0). The project open-sourced its dataset, tooling, and dashboard after spending more than $3M building the traces and running a matrix on ~2MW across 1,000+ chips. AgentX has already driven 50+ upstream PRs and cross-project optimizations across vLLM, SGLang, TensorRT-LLM, ATOM, LMCache, Mooncake and Dynamo. Results show mixed vendor outcomes: NVIDIA leads on many frontier models and configurations, AMD (with ATOM) is competitive in parts, and recent optimizations have shifted some perf-per-dollar comparisons. The dataset (393-session subset) is available on HuggingFace and the results/dashboards are public on InferenceX.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.