Observed Signal · Sep 18, 2026 · Technical Release · Source: SemiAnalysis · Impact: 4/5 · Sentiment: Positive
AI Infrastructure Market: SemiAnalysis Tests Engram Offloading to DRAM and SSD
SemiAnalysis analyzes the Engram architecture, a model design that extends standard token embeddings with learned multi-token lookups, allowing for efficient parameter offloading to DRAM or SSD. This reduces HBM capacity requirements for models like DeepSeek-V4.1-Flash. Their experiments show that offloading Engram tables to DRAM can improve performance per dollar, while SSD offloading is currently not economically viable due to overhead. The article also benchmarks inference performance across NVIDIA and AMD GPUs, highlighting NVIDIA's CUDA moat and AMD's slower software support. The analysis includes findings on model behavior, such as gate scan results and ablation studies, and discusses the implications for HBM demand and model architecture innovation.
Technical release of a novel model architecture optimization that could reduce HBM demand and impact inference costs, relevant to AI/AdTech infrastructure.
Wichtigste Kernpunkte & Evidenz
- SemiAnalysis tested offloading Engram tables to DRAM and SSD for DeepSeek-V4.1-Flash.
- Offloading to DRAM improved performance per dollar, reducing needed HBM capacity.
- SSD offloading was not economically viable, with DRAM delivering 121 million tokens per dollar versus 52 million for SSD.
- NVIDIA vLLM worked out of the box for DeepSeekV4.1 Flash, while AMD's vLLM support was delayed.
- AMD MI355X performance per dollar remains 2-4x worse compared to NVIDIA B200.
Connected Companies & Entities
9 Entities mappedSemiAnalysis
AI infrastructure and semiconductor research, data models, tools and consulting.
“SemiAnalysis analyzes the Engram architecture and provides benchmarks....”
Meta
Consumer internet platforms monetised through advertising, apps, subscriptions and VR.
“SemiAnalysis benchmarks are supported by Meta....”
NVIDIA
Accelerated computing company spanning AI software, cloud and gaming.
“NVIDIA vLLM worked out of the box for DeepSeekV4.1 Flash across all 6 SKUs....”
AMD
US semiconductor company designing processors and computing chips.
“AMD's vLLM support for DeepSeekV4.1 Flash was delayed, and MI355X performance lags....”
DeepSeek
LLM developer offering AI chat and API access.
“DeepSeek did not release the original Engram models; SemiAnalysis replicated the setup....”
Microsoft
Diversified software, cloud, advertising and gaming platform company.
“SemiAnalysis benchmarks are validated by Microsoft Azure among others....”
OpenAI
Foundation model company selling AI software, APIs and subscriptions.
“SemiAnalysis benchmarks are supported by major labs like OpenAI....”
Hugging Face
Open AI model hub with hosted inference and collaboration.
“SemiAnalysis benchmarks are supported by the ML community including Hugging Face....”
Google Cloud
Enterprise cloud, data and AI services for businesses.
“SemiAnalysis benchmarks are validated by Google Cloud among others....”
Ontology Mapping & Concepts
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
