Observed Signal · Jun 9, 2026 · Technical Release · Source: SemiAnalysis · Impact: 3/5 · Sentiment: Positive
DeepSeek v4 Day‑0 to Day‑43 Inference Performance Report
SemiAnalysis’ InferenceX published an engineering analysis of DeepSeek v4 Pro’s Day‑0 through Day‑43 inference performance across multiple accelerator SKUs (GB300 NVL72, Huawei Ascend 950DT, MI355X, B200/B300, H200). The report documents Day‑0 support on CUDA and Huawei CANN, early ROCm/AMD regressions and a subsequent >100x AMD throughput improvement by Day 26 led by HaiShaw’s team. It details kernel and runtime issues (notably a TensorRT‑LLM fused HC hidden‑size guard), PRs submitted to TensorRT‑LLM, and many incremental software optimizations (MTP, MegaMoE, FP4 paths, AITER kernels, Triton/TileLang/FlyDSL integration). The article highlights rack‑scale GB300 NVL72 SGLang results (with CoreWeave providing GB300 racks), discusses DeepSeek v4 architectural features (HCA/CSA, MegaMoE) and reports throughput and tokens‑per‑MW efficiency gains tied to software improvements.
Documents cross‑stack Day‑0 support and multi‑week software optimizations for a major open LLM, showing material throughput and efficiency gains (including >100x AMD improvement) that affect inference TCO and deployability across key accelerator vendors.
Track Huawei Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- InferenceX (SemiAnalysis) measured DeepSeek v4 Pro inference performance from Day 0 to Day 43 across GPUs including GB300 NVL72, Huawei Ascend 950DT, MI355X, B200, B300 and H200.
- vLLM and SGLang provided native Day‑0 support for DeepSeek v4 on CUDA; Huawei CANN also documented Day‑0 Ascend support.
- AMD ROCm performance for MI355X was poor at Day 0 but improved by more than 100x by Day 26 under technical leadership credited to HaiShaw.
- SemiAnalysis authored a PR fixing a TensorRT‑LLM fused HC kernel issue (TensorRT‑LLM PR #13710) and referenced subsequent upstream merges/patches (e.g., PR #13771) to address a hidden‑size guard problem.
- CoreWeave contributed two GB300 NVL72 racks enabling the reported GB300 SGLang performance results.
Connected Companies & Entities
8 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
DeepSeek V4 (MODEL1) Expected with Engram, mHC
DeepSeek, a Chinese open-source AI startup, is expected to release DeepSeek V4 (rumored codename MODEL1) around the Lunar New Year (week of Feb 17, 2026). Reporting and code commits indicate V4 will be a major architectural overhaul focused on extreme long-context coding and software-engineering tasks. Key innovations described include Engram (a conditional memory lookup to separate factual recall from reasoning and enable multi-million-token knowledge stores), Manifold-Constrained Hyper-Connections (mHC) to stabilize rich cross-layer connectivity, and DeepSeek Sparse Attention (DSA) for 1M+ token contexts. DeepSeek reportedly delayed its R2 training after hardware instability with Huawei Ascend chips and reverted to Nvidia GPUs for final training. The article places DeepSeek within a broader surge of Chinese open-weight model activity (names cited include Qwen/Alibaba Cloud, Zhipu AI, Moonshot AI, and Minimax).
InferenceX v2 Benchmarks Blackwell vs AMD & Hopper
SemiAnalysis released InferenceX v2 (formerly InferenceMAX), an open-source Apache 2.0 continuous inference benchmark that expands coverage across ~1,000 frontier GPUs and new distributed inference modes. InferenceXv2 adds large-scale disaggregated prefill (disagg) with wide expert parallelism (wideEP) testing for six recent NVIDIA GPU SKUs (including GB200/GB300 NVL72, B200, B300, Blackwell Ultra) and all recent AMD western SKUs including MI355X. The release includes the first third‑party Pareto-frontier benchmarks for Blackwell Ultra GB300 NVL72 and multi-node MI355X disagg+wideEP FP4/FP8. Key findings: NVIDIA Blackwell rack-scale systems lead for MoE/disaggregated inference and energy efficiency; AMD MI355X is competitive on some FP8 and single-node perf/TCO but suffers composability and FP4 multi-node software gaps. The report also highlights MTP (multi-token/speculative decoding) as a major cost reducer and documents software stacks such as SGLang, vLLM, TensorRT‑LLM, Dynamo, MoRI and Mooncake.
DeepSeek Founder: Compute Is the Primary Constraint
Hello China Tech published a July 2026 selection from a nearly four-hour investor transcript in which DeepSeek founder Liang Wenfeng argued that compute availability is the principal gap between DeepSeek and leading US AI labs. Liang framed differences in talent, model capability, and applications as downstream consequences of smaller compute budgets and limited chip supply. He placed DeepSeek’s first external round at over Rmb 50bn (~$7.4bn) (not officially confirmed by the company), described work to reduce dependence on Nvidia via a high-level compiler called TileLang, and forecast that within a year domestic Chinese chips could be verified as usable for training. Liang also discussed pricing, open-weight releases, team retention via option grants, and V4 multimodality commitments.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
