AI Model Economics Shift: Cheaper Inference, New Options
This weekly newsletter covers developments in AI model deployment and economics from October 4-10, 2026. Anthropic released Claude Haiku 5.5, a cost-efficient small model with adjustable effort settings, reducing prices for high-volume tasks. Mistral launched the public API preview of Mistral Large 4 ("Le Chonk"), a 1.05-trillion-parameter multimodal model with a one-million-token context, with weights promised later. Liquid AI introduced open-weight decision models (d1-3B and d1-omni-600M) that output probabilities in a single forward pass without generating text. Google open-sourced ML Drift, the GPU acceleration engine for LiteRT, supporting multiple graphics APIs and improving edge inference. Additionally, JFrog disclosed a critical remote-code-execution vulnerability (CVE-2026-105192) in LMCache, highlighting security risks in shared inference infrastructure.
- •Anthropic released Claude Haiku 5.5 on October 7, 2026, with pricing at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens.
- •Mistral launched the public API preview of Mistral Large 4 on October 6, 2026, featuring 1.05 trillion total parameters and a one-million-token context window.
- •Liquid AI released open-weight decision models d1-3B and d1-omni-600M on October 7, 2026.
