Observed Signal · Jul 21, 2026 · Partnership · Source: The Pragmatic Engineer · Impact: 2/5 · Sentiment: Positive

Napkin Math Powers Turbopuffer's Efficient Search

Executive Signal Summary

This Pragmatic Engineer newsletter summarizes an interview with Simon Eskildsen, co-founder and CEO of turbopuffer, on using “napkin math” — quick first-principles calculations of latency and cost — to find theoretical limits and identify large performance or cost gaps in systems. Eskildsen described how long tenure at Shopify and tools like the open-sourced toxiproxy influenced his approach to infrastructure. He built turbopuffer, an S3-backed vector search architecture motivated by excessive costs of existing solutions; Cursor became the first customer and turbopuffer claimed large unit-cost reductions. The piece covers early fundraising (roughly $700K initially), product design choices (S3 + clustering + simple caching via NGINX) and six reasons founders raise venture capital.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Engineering-focused profile about a startup vector-search architecture and cost/performance techniques; relevant to cloud infrastructure and vector databases but not directly to core AdTech/MarTech product changes.

SIGNAL RADAR

Track Shopify Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Simon Eskildsen is co-founder and CEO of turbopuffer.
  • Simon worked nearly a decade at Shopify and created the open-source toxiproxy (open-sourced in 2014).
  • Turbopuffer used an S3-backed clustering approach and an NGINX-level cache as a simple initial architecture to reduce search latency and cost.
  • Cursor became turbopuffer’s first customer and, according to the article, turbopuffer reduced Cursor’s indexing/search cost from about $80K/month to $4K/month (≈95% reduction).
  • Turbopuffer raised approximately $700K initially to hire two engineers, and the article states the company later crossed $100M in annual run rate with less than $1M initial funding.

Connected Companies & Entities

7 Entities mapped

“While at Shopify, Simon became interested in figuring out the theoretical limits of certain computer operations....”

“Cursor migrated their local code-base search over to this brand new product....”

“Cursor migrated their local code-base search over to this brand new product....”

“For the first version, I didn’t even implement a dedicated caching layer. I just put the reverse proxy (NGINX) in front of S3, and that was ...”

“Growing up in Denmark, he dabbled in HTML with tools like Microsoft FrontPage and Adobe’s Dreamweaver....”

“In 2024, the company was valued at $400M, and then its valuation surged to $29.3B. It was sold to SpaceX for $60B this year....”

“Growing up in Denmark, he dabbled in HTML with tools like Microsoft FrontPage and Adobe’s Dreamweaver....”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: The Pragmatic Engineer•Published: Jul 21, 2026
Original Coverage Title: “Pushing software engineering limits with “napkin math””

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Infrastructure / Retrieval & SearchMar 12, 2026

Turbopuffer Builds Search Engine for AI Retrieval

Turbopuffer — founded by Simon Hørup Eskildsen from work that began at Readwise — is positioning itself as a search engine for unstructured data by combining object storage (S3/GCS) with NVMe and memory tiering. The company’s architecture intentionally avoids a traditional consensus layer and relies on modern cloud primitives (object-store consistency, compare-and-swap on object storage, NVMe SSDs) to reduce cost and operational complexity. Early customers (Cursor, Notion) used Turbopuffer to cut costs and improve semantic/code search; the company reports heavy vector and full‑text workloads and is optimizing for agentic retrieval patterns that produce high concurrency. The interview covers origin stories, architectural tradeoffs, tiered storage strategy, pricing evolution, hiring philosophy (‘P99 engineer’), and roadmaps for ANN/ANNV versions and full-text search feature expansion.

Read assessment
Large Language Models (LLM) & AIMar 26, 2026

Google achieves 6x KV-cache compression without training

This edition of The Tokenizer curates recent AI/ML research and tools centered on speed and efficiency. Highlights include Google Research's TurboQuant, a training-free KV-cache compression method using polar-coordinate transforms and random projections that enables 3-bit KV quantization, a reported 6x KV memory reduction and up to 8x performance gains on H100 GPUs. Other items cover a diffusion-based OCR approach up to 3.2x faster throughput, SkillNet (an npm-like package manager for agent skills) showing reward and step-count improvements, Stripe’s internal AI coding agents shipping ~1,300 PRs per week, ByteDance’s OpenViking filesystem-based context DB for agents, DeepSeek’s Engram memory module adding O(1) lookup to transformers, and a widely circulated Claude Code cheat sheet. The newsletter summarizes papers, implementations, datasets, and practical walkthroughs that emphasize inference, memory, and agent efficiency.

Read assessment
Application Performance Monitoring (APM)Aug 6, 2026

Profiling-First Cuts E‑Commerce Latency 30%

A technical case study describes how a team reduced latency by roughly 30% on a high-traffic e-commerce platform over about two and a half years while preserving 99.9% uptime during peak trading periods. The improvements came from a repeatable, data-driven process: profile in production-like traffic before changing code, fix true bottlenecks (not assumed ones), and apply targeted fixes across three areas — query optimization (batching and selective denormalization), scoped refactors of hot code paths, and infrastructure tuning (right-sizing cloud instances and function configurations). The team also automated internal processes to cut manual data work by 80% and emphasizes incremental rollouts with p50/p95/p99 monitoring and rollback capability.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.