Observed Signal · Feb 16, 2026 · Technical Release · Source: SemiAnalysis · Impact: 3/5 · Sentiment: Positive
InferenceX v2 Benchmarks Blackwell vs AMD & Hopper
SemiAnalysis released InferenceX v2 (formerly InferenceMAX), an open-source Apache 2.0 continuous inference benchmark that expands coverage across ~1,000 frontier GPUs and new distributed inference modes. InferenceXv2 adds large-scale disaggregated prefill (disagg) with wide expert parallelism (wideEP) testing for six recent NVIDIA GPU SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and all recent AMD western SKUs including MI355X. The release includes the first third‑party Pareto-frontier benchmarks for Blackwell Ultra GB300 NVL72 and multi-node MI355X disagg+wideEP FP4/FP8. Key findings: NVIDIA Blackwell rack-scale systems lead for MoE/disaggregated inference and energy efficiency; AMD MI355X is competitive on some FP8 and single-node perf/TCO but suffers composability and FP4 multi-node software gaps. The report also highlights MTP (multi-token/speculative decoding) as a major cost reducer and documents software stacks such as SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI and Mooncake.
Open-source, large-scale inference benchmark expands cross-vendor, multi-node MoE and disaggregated testing (including first third-party Blackwell Ultra and MI355X disagg+wideEP results). Results influence LLM inference TCO, software priorities (composability), and vendor selection for production inference infrastructure.
Track NVIDIA Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- SemiAnalysis published InferenceX v2, an open-source (Apache 2.0) continuous inference benchmark.
- InferenceXv2 benchmarks nearly 1,000 GPUs across NVIDIA SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and recent AMD SKUs (including MI355X).
- InferenceXv2 is the first third-party suite to benchmark Blackwell Ultra GB300 NVL72 across the Pareto frontier and to test MI355X disagg+wideEP multi-node FP4 and FP8 performance.
- Findings: NVIDIA Blackwell systems lead in large-scale MoE disaggregated inference and energy efficiency; AMD MI355X is competitive on FP8 single-node and some FP8 disagg configs but faces composability and FP4 multi-node software limitations.
- Speculative decoding / MTP materially reduces cost per token across many configurations; the report documents production software stacks (SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI, Mooncake).
Connected Companies & Entities
7 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
SemiAnalysis Microbenchmarks Nvidia Blackwell Tensor Cores
SemiAnalysis published an in-depth microbenchmarking study of Nvidia’s datacenter Blackwell (SM100) GPU, analyzing PTX and SASS instruction performance for AI workloads. The report measures low-level behavior and practical upper bounds for tensor-related primitives (UMMA, TMA, cp.async/LDGSTS), TMEM, DSMEM, multicast, cluster scheduling, and new 2SM MMA (.cta_group::2) instructions. Key empirical findings include memory throughput and latency curves for async copy and TMA, a measured die-to-die L2 penalty (~300 cycles), and shape- and layout-dependent MMA throughput and latency (including cases where M=64 is SMEM‑bound at ~50% peak). The authors open-sourced their benchmarking repo, acknowledge cloud node providers and reviewers, and plan further kernel benchmarking across Blackwell variants and other AI accelerators (TPU Pallas, Trainium, AMD CDNA4).
Qwen3-8B inference benchmark and FP8 on Blackwell
Independent benchmarks compare Qwen3-8B inference on an RTX PRO 6000 Blackwell (96 GB) across three serving stacks (vLLM 0.27.1, SGLang 0.5.9, and llama.cpp CUDA). At concurrency 32 using BF16, vLLM achieved 1,725 aggregate tokens/s (TTFT p50 39 ms), SGLang 1,327 tok/s (TTFT p50 42 ms), and llama.cpp 428 tok/s (TTFT p50 316 ms). Applying an FP8 checkpoint to vLLM increased throughput by ~1.5x (aggregate 1,725 -> 2,597 tok/s; single-stream 86 -> 130 tok/s) with lower latency and no detected regressions on a fixed factual check. The author documents methodology, reproductions, and an sm_120-specific kernel workaround required to run FP8 on workstation Blackwell hardware.
AgentX 1.0: Open-source Agentic Inference Benchmark
SemiAnalysis announced AgentX 1.0 and InferenceXv3, an open-source agentic inference benchmark and benchmark implementation designed for long-context, multi-turn agentic coding workloads up to 1M context (Apache 2.0). The project open-sourced its dataset, tooling, and dashboard after spending more than $3M building the traces and running a matrix on ~2MW across 1,000+ chips. AgentX has already driven 50+ upstream PRs and cross-project optimizations across vLLM, SGLang, TensorRT-LLM, ATOM, LMCache, Mooncake and Dynamo. Results show mixed vendor outcomes: NVIDIA leads on many frontier models and configurations, AMD (with ATOM) is competitive in parts, and recent optimizations have shifted some perf-per-dollar comparisons. The dataset (393-session subset) is available on HuggingFace and the results/dashboards are public on InferenceX.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
