Observed Signal · Aug 14, 2026 · Technical Release · Source: techcrunch · Impact: 3/5 · Sentiment: Positive

Kog Optimizes GPUs to Speed LLM Inference

Executive Signal Summary

French startup Kog is developing low-level software optimizations to accelerate AI inference on standard datacenter GPUs (e.g., AMD MI300X, Nvidia H200). Founder and CEO Gaël Delalleau says an early tech preview demonstrated 3,000 per-request tokens-per-second using a small, purpose-built 2B-parameter model (Laneformer 2B), which has since been open-sourced. Kog has gathered customer interest, design partners, and support from Scaleway, Bpifrance and the French Tech 2030 program, and says Varsity VC co-led its seed round. The company (11 people) aims to apply its approach to larger LLMs and expects to implement a first major model at 10x speed before pursuing a Series A.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Software-driven GPU inference improvements from startups can materially lower LLM inference cost/latency and enable broader real-time applications, but this is an early-stage company demonstration rather than a major platform policy or release.

SIGNAL RADAR

Track AMD Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Kog is a French startup building software optimizations to speed AI inference on standard datacenter GPUs.
  • Kog's tech preview demo achieved about 3,000 per-request tokens-per-second using a small ~2B-parameter model (Laneformer 2B), which is now open-sourced.
  • The demo used standard datacenter GPUs including AMD MI300X and Nvidia H200.
  • Kog reported around 200 tangible business leads following its May tech preview; the company has a team of 11 people.
  • Varsity VC co-led Kog's seed round; the startup is supported by Scaleway and backed by France’s Bpifrance and the French Tech 2030 program.

Connected Companies & Entities

5 Entities mapped

“such as the AMD MI300X and Nvidia H200 GPUs it used for its demo....”

“The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May....”

“such as the AMD MI300X and Nvidia H200 GPUs it used for its demo....”

“Anthropic itself understands that speed is worth money, and charges a price multiple for Claude’s Fast Mode....”

“this could add sovereignty tailwinds for the startup, which is already supported by Scaleway and backed by France’s Bpifrance and French Tec...”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: techcrunch•Published: Aug 14, 2026
Original Coverage Title: “Kog is going deeper to squeeze more inference out of GPUs”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIMar 26, 2026

Optimizing LLM Costs: TurboQuant and Production Strategies

A practitioner post from 498Advance describes a three-layer approach to reduce production LLM costs (fallback policies, task-aware routing, and selective local model hosting) and highlights a new Google Research paper, TurboQuant (ICLR 2026). TurboQuant, authored by Amir Zandieh and Vahab Mirrokni, introduces a compression pipeline combining PolarQuant and a Quantized Johnson‑Lindenstrauss (QJL) correction to dramatically reduce KV cache size and attention cost without retraining. Reported headline results include up to 6x KV cache memory reduction, 8x attention speedup with 4‑bit quantization on H100 GPUs, and effective 3‑bit KV cache quantization with no measured accuracy loss. The article also cites industry examples (LinkedIn, Roblox, Red Hat) using model optimization, quantization, sparsity, distillation, Ray and vLLM for scalable inference.

Read assessment
Video / AI InfrastructureMay 11, 2026

AKOOL Launches Real-Time AI Video Inference Engine

AKOOL announced a production-grade AI video inference engine that it says delivers 10–20× faster performance than conventional approaches and enables real-time AI video at global scale. The company claims single-clip generation can drop from tens of seconds to 1–3 seconds and that the system supports sub-30 millisecond latency per frame for live streaming. AKOOL attributes the gains to a full-stack redesign spanning algorithm, GPU parallelism, runtime overhead reduction, and next-generation GPU architectures. The engine runs across cloud, streaming, and on-device deployments and includes production reliability features such as real-time monitoring, automated quality controls, staged deployments and per-model cost tracking. AKOOL is already using the inference engine in its products (including Akool Live Camera) for live digital avatars, live translation, and interactive video experiences. The announcement was published May 11, 2026.

Read assessment
Large Language Models (LLM) & AIMar 26, 2026

Google achieves 6x KV-cache compression without training

This edition of The Tokenizer curates recent AI/ML research and tools centered on speed and efficiency. Highlights include Google Research's TurboQuant, a training-free KV-cache compression method using polar-coordinate transforms and random projections that enables 3-bit KV quantization, a reported 6x KV memory reduction and up to 8x performance gains on H100 GPUs. Other items cover a diffusion-based OCR approach up to 3.2x faster throughput, SkillNet (an npm-like package manager for agent skills) showing reward and step-count improvements, Stripe’s internal AI coding agents shipping ~1,300 PRs per week, ByteDance’s OpenViking filesystem-based context DB for agents, DeepSeek’s Engram memory module adding O(1) lookup to transformers, and a widely circulated Claude Code cheat sheet. The newsletter summarizes papers, implementations, datasets, and practical walkthroughs that emphasize inference, memory, and agent efficiency.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.