GLM-5.3 Sparse Attention Impact on DRAM Memory TAM
This article analyzes the impact of sparse attention mechanisms, specifically DeepSeek Sparse Attention (DSA) used in Z.ai's GLM-5.3 model, on the total addressable market (TAM) for DRAM memory, including HBM and NAND. It explains that while sparse attention reduces KV cache memory and bandwidth during the attention operation, it does not reduce overall memory capacity requirements because the top-k selection still requires full context in HBM. The article discusses system optimizations like HiSparse, which offloads KV cache to host DRAM to overcome capacity bottlenecks. It also provides detailed performance and cost comparisons for serving GLM-5.3 on different hardware (GB200, GB300, MI355X) using inference engines like Dynamo-SGLang, Dynamo-TRT-LLM, and ATOM, highlighting cost-efficiency and interactivity trade-offs. The analysis includes a deep dive into GLM-5's architecture, including the lightning indexer, MLA configuration, and post-training pipeline.
- •GLM-5.3 uses DeepSeek Sparse Attention (DSA) with a lightning indexer for top-k token selection.
- •Sparse attention does not reduce overall memory capacity usage due to the need for full context in HBM.
- •HiSparse, a hierarchical memory system, offloads KV cache to host DRAM to improve throughput at high concurrency.
