Observed Signal · Jul 22, 2026 · Technical Release · Source: TheSequence · Impact: 2/5 · Sentiment: Neutral
Inkling: Trillion-Parameter Model Wakes 41B at a Time
The article describes Inkling, a large-language model with a headline size of roughly 975 billion parameters but which activates only about 41 billion parameters per token (≈4.2%). Rather than behaving as a single monolithic network, Inkling is framed as a vast repository of specialist capacity: a routing mechanism selects a small working set of specialists for each token (the piece describes selecting six specialist departments plus two always-attending general-purpose departments). This design produces sparse arithmetic during inference while requiring large storage, networking, and deployment resources. The piece was published on 2026-07-22.
Describes a large-scale LLM architecture (sparse activation / router-based experts) that may influence AI capabilities used in marketing and creative tools, but is not a major platform policy change or a direct AdTech product launch.
Track The Sequence Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Inkling's headline parameter count is reported as 975 billion parameters.
- Approximately 41 billion parameters are active per token in Inkling (about 4.2% of total).
- A router selects a small working set for each token; the article describes selecting six specialist departments plus two general-purpose departments per token.
- The design yields sparse arithmetic while storage, networking, and deployment remain large.
- Article publication date: 2026-07-22.
Connected Companies & Entities
1 Entity mapped“Title: The Sequence AI of the Week #899: Inside Inkling: A Trillion-Parameter Model That Only Wakes Up 41 Billion at a Time...”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Thinking Machines releases open-weight model Inkling
Thinking Machines released Inkling, an open-weights, general-purpose multimodal Mixture-of-Experts foundation model: a 66-layer decoder-only transformer with a sparse MoE backbone (≈975B total parameters, ≈41B active per task). It natively accepts text, images and audio and produces UTF-8 text outputs. Trained from scratch on roughly 45 trillion tokens assembled from public, third-party and synthetic sources on NVIDIA GB300 NVL72 systems under a strategic NVIDIA partnership, Inkling ships with checkpointed support for up to a 1,000,000-token context (Tinker API exposes 256K) and includes a smaller preview (Inkling‑Small, ≈276B/12B). Published under Apache‑2.0 with day‑0 ecosystem and inference/hosting support (Hugging Face, vLLM, SGLang, TokenSpeed, Modal, Databricks, Baseten, NVIDIA optimizations and community quantization), it is positioned for enterprise fine‑tuning via Tinker, evaluated against benchmarks (2026-07-14) and accompanied by safety guidance.
ByteDance Bets on 5–10 Trillion Parameter Model
Reports in August 2026 indicate ByteDance is concentrating talent, data, and compute on a single long-duration language-model project—nicknamed in coverage as a Seed/Seedance effort—reported by LatePost as potentially exceeding 5 trillion parameters and by the Financial Times as up to 10 trillion parameters in early-stage training. Company leadership, including founder Zhang Yiming and CEO Liang Rubo, has signaled AI as a long-term strategic priority and directed the Seed team to prioritize frontier capability over benchmarking or distillation of rivals. The article contrasts ByteDance’s historical “ship fast, measure quickly, concentrate resources” operating system (which succeeded in short-form consumer apps) with the long feedback loops and high costs of frontier model training, and notes related company moves such as the March 2026 sale of gaming studio Moonton.
Tasteful Tokenmaxxing: AI Leaders Favor Depth Over Breadth
This AINews roundup (Apr 23, 2026) synthesizes industry conversations and product announcements focused on efficient AI usage and model/platform progress. The newsletter highlights a growing practice labeled “Tokenmaxxing” — using more model tokens while avoiding waste — and reports that many engineering leaders prefer deeper, serial autoresearch loops over massively parallel LLM runs. Major technical announcements covered include Google’s TPU v8 family (TPU 8t for training, TPU 8i for inference) and the Gemini Enterprise Agent Platform and Workspace Intelligence, Alibaba’s open-source Qwen3.6-27B, OpenAI’s Apache‑2.0 Privacy Filter for PII detection/redaction, and Xiaomi’s MiMo-V2.5 models. The piece also surveys trends: hardening agent harness abstractions, bring-your-own-model support in developer tooling, traces/agent data as a core primitive, post-training RL improvements, and ongoing inference-efficiency innovations.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
