Observed Signal · Sep 28, 2026 · Technical Release · Source: SemiAnalysis · Impact: 4/5 · Sentiment: Positive
GLM-5.3 Sparse Attention Impact on DRAM Memory TAM
This article analyzes the impact of sparse attention mechanisms, specifically DeepSeek Sparse Attention (DSA) used in Z.ai's GLM-5.3 model, on the total addressable market (TAM) for DRAM memory, including HBM and NAND. It explains that while sparse attention reduces KV cache memory and bandwidth during the attention operation, it does not reduce overall memory capacity requirements because the top-k selection still requires full context in HBM. The article discusses system optimizations like HiSparse, which offloads KV cache to host DRAM to overcome capacity bottlenecks. It also provides detailed performance and cost comparisons for serving GLM-5.3 on different hardware (GB200, GB300, MI355X) using inference engines like Dynamo-SGLang, Dynamo-TRT-LLM, and ATOM, highlighting cost-efficiency and interactivity trade-offs. The analysis includes a deep dive into GLM-5's architecture, including the lightning indexer, MLA configuration, and post-training pipeline.
News on large language model architecture and inference optimization directly impacting the AI infrastructure and memory market, with detailed cost-performance analysis relevant to AI/AdTech vendors.
Track NVIDIA Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- GLM-5.3 uses DeepSeek Sparse Attention (DSA) with a lightning indexer for top-k token selection.
- Sparse attention does not reduce overall memory capacity usage due to the need for full context in HBM.
- HiSparse, a hierarchical memory system, offloads KV cache to host DRAM to improve throughput at high concurrency.
- GB200 with Dynamo-SGLang costs $0.044 per million total tokens at 150 tokens per second, 12% less than MI355X with ATOM.
- GLM-5.3 on TileRT with MI355X achieves 2x P90 interactivity compared to FP4 MI335X config.
Connected Companies & Entities
5 Entities mapped“GB200, GB300...”
“MI355X...”
“DeepSeek Sparse Attention (DSA), introduced in DeepSeek V3.2...”
“Z.ai's GLM-5 model series...”
“SGLang team designed HiSparse...”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Samsung to invest $1B in AI infrastructure firm Helix
Samsung Electronics and five affiliates will invest a combined $1 billion in Helix Digital Infrastructure, an AI infrastructure company launched by KKR and backed by Nvidia. Samsung Electronics contributes $500 million, with the rest from Samsung C&T, Samsung SDS, Samsung SDI, Samsung Life Insurance, and Samsung Fire & Marine Insurance. Helix, led by former AWS CEO Adam Selipsky, focuses on hyperscale data centers, power generation, transmission, and fiber-optic networks. The investment adds to over $10 billion already committed by other investors including KKR, Kuwait Investment Authority, Nvidia, and Vistra. The move allows Samsung to leverage its semiconductor, cooling, data center construction, and battery capabilities to expand in the AI infrastructure market.
Alibaba's T-Head AI Chips: Cloud Customers or Qwen Training?
Alibaba announced at its Apsara conference that its new Zhenwu V900 AI chip will enter mass production and go on sale in Q1 2027, two quarters earlier than planned. This follows Huawei's announcement of its Ascend 960DT chip being ready in Q1 2027. Both companies face high demand and limited supply for their chips. IDC data shows Nvidia holds 55% of China's server AI accelerator shipments, Huawei 20%, and T-Head 7%. Alibaba plans to train Qwen models with 5-10 trillion parameters, but has not disclosed which chips will be used, raising concerns about competition between internal model training and paying cloud customers for scarce chip capacity.
Nscale Files for US IPO Targeting $30B Valuation
British AI infrastructure company Nscale has filed its IPO prospectus with the US SEC, aiming to list on the New York Stock Exchange under the ticker 'NSCL'. The company seeks a valuation of around $30 billion, as reported by CNBC via Reuters, a significant jump from its last private round valuation of $14.6 billion. Nscale, which transitioned from Bitcoin mining to AI compute, reported first-half 2026 revenue of $140.6 million, a 1,252% increase, but also a net loss of $1.02 billion. Its order book stands at $103.4 billion, with Microsoft and Anthropic contracts totaling $88.4 billion. The company recently secured $3.1 billion in convertible notes, with Nvidia as a key investor. Nscale faces high customer concentration, a going-concern note, and dependence on a few suppliers. Goldman Sachs, J.P. Morgan, and Morgan Stanley are lead underwriters.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
