Observed Signal · Sep 12, 2026 · Technical Release · Source: AINews swyx · Impact: 5/5 · Sentiment: Positive
AI Model Launch Market: DeepSeek Launches V4.1 Flash with Novel Encoder-Decoder Architecture
DeepSeek released DeepSeek-V4.1-Flash, a 763B-parameter mixture-of-experts model employing a novel causal encoder-decoder architecture with 8B active parameters for prefill and 16B for decode. It features native vision understanding, 1M token context, an MIT license, and extreme inference efficiency, claiming up to 1/8 KV cache footprint versus V4 Flash. Independent evals (Artificial Analysis Index 40, Vals Index #1 open-weight) show it surpasses V4 Pro at lower cost. API pricing is $0.30/1M input and $1.20/1M output tokens. DeepSeek has soft-retired V4 Pro, routing traffic to V4.1 Flash. The model supports SSD offload and local deployment, with Ollama and Baseten offering day-0 support. Technical discussions highlight the architecture's novelty and potential impact on long-context agents.
Technical release from DeepSeek, a major AI platform, introducing a revolutionary architecture with industry-shifting implications for inference cost and efficiency. Sets new benchmarks for open-weight models and could force competitive responses.
Wichtigste Kernpunkte & Evidenz
- DeepSeek launched V4.1-Flash with a causal encoder-decoder architecture, 763B total params (8B prefill/16B decode active).
- Artificial Analysis Index scores V4.1-Flash at 40, above V4 Pro and below GLM-5.3-Flash.
- API pricing: $0.30 per 1M input tokens, $1.20 per 1M output tokens, cached input $0.006 per 1M.
- V4.1-Flash is #1 open-weight model on Vals Index, costing $0.30 per test.
- DeepSeek soft-retired V4 Pro, routing traffic to V4.1 Flash at Flash pricing.
- Independent eval: AutomationBench-AA 69%, GDPval-AA v2 Elo 1632 (surpassing Kimi K3), AA-LCR 84%.
- Local deployment reports: 200+ TPS on 4 Max-Qs with NVMe offload; 300+ TPS on 4 RTX Pros.
- Baseten and Ollama provide day-0 support for V4.1-Flash.
Connected Companies & Entities
4 Entities mappedBaseten
B2B platform for AI model serving, inference and deployment.
“Baseten shipped day-0 support for DeepSeek V4.1-Flash....”
DeepSeek
LLM developer offering AI chat and API access.
“DeepSeek launched V4.1-Flash, its new open-weight flagship....”
Ollama
Local and cloud infrastructure for open-model AI development.
“Ollama began rolling out V4.1-Flash to Max and Team accounts....”
Artificial Analysis
Independent AI model benchmarking and selection platform.
“Independent benchmark account Artificial Analysis reported scores and pricing for V4.1-Flash....”
Ontology Mapping & Concepts
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
