Observed Signal · Sep 18, 2026 · Technical Release · Source: SemiAnalysis · Impact: 4/5 · Sentiment: Positive

AI Infrastructure Market: SemiAnalysis Tests Engram Offloading to DRAM and SSD

Executive Signal Summary

SemiAnalysis analyzes the Engram architecture, a model design that extends standard token embeddings with learned multi-token lookups, allowing for efficient parameter offloading to DRAM or SSD. This reduces HBM capacity requirements for models like DeepSeek-V4.1-Flash. Their experiments show that offloading Engram tables to DRAM can improve performance per dollar, while SSD offloading is currently not economically viable due to overhead. The article also benchmarks inference performance across NVIDIA and AMD GPUs, highlighting NVIDIA's CUDA moat and AMD's slower software support. The analysis includes findings on model behavior, such as gate scan results and ablation studies, and discusses the implications for HBM demand and model architecture innovation.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Technical release of a novel model architecture optimization that could reduce HBM demand and impact inference costs, relevant to AI/AdTech infrastructure.

Key Takeaways & Evidence Grounding

  • SemiAnalysis tested offloading Engram tables to DRAM and SSD for DeepSeek-V4.1-Flash.
  • Offloading to DRAM improved performance per dollar, reducing needed HBM capacity.
  • SSD offloading was not economically viable, with DRAM delivering 121 million tokens per dollar versus 52 million for SSD.
  • NVIDIA vLLM worked out of the box for DeepSeekV4.1 Flash, while AMD's vLLM support was delayed.
  • AMD MI355X performance per dollar remains 2-4x worse compared to NVIDIA B200.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: SemiAnalysisPublished: Sep 18, 2026
Original Coverage Title: Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.