Observed Signal · Jun 30, 2026 · Product Review · Source: Lennys Newsletter · Impact: 3/5 · Sentiment: Neutral
Sonnet 5 review: 64-run How I AI benchmark
Claire Vo built the "How I AI Bench" live using Claude Code and ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT‑5.5, Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The benchmark comprised ~64 generations and combined human "vibe" scoring (70%) with LLM-as-judge scoring (30%). Results surprised the author: Gemini 3 Pro, Sonnet 5, and GPT‑5.5 ranked highly on the automated leaderboard, but Claire's personal taste favored different models (Sonnet 4.6 / Opus 4.8). The piece also notes Sonnet 5's introductory pricing and Anthropic's positioning of Sonnet 5 as a lower-cost, more agentic model for running tool-using workflows.
Provides comparative benchmarking and cost/performance data for frontier LLMs (including Anthropic's Sonnet 5), informing model selection and cost trade-offs for teams building agentic workflows and creative tooling.
Track Anthropic Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Claire Vo built the How I AI Bench live using Claude Code and ran roughly 64 model generations across multiple tasks.
- Five frontier models were tested blind: Sonnet 5, Sonnet 4.6, Opus 4.8, GPT‑5.5, and Gemini 3 Pro.
- The benchmark combined human "vibe" scoring (70%) with LLM judge scoring (30%).
- Anthropic positioned Sonnet 5 as a lower-cost, agentic model; introductory pricing cited at $2 per million input tokens and $10 per million output tokens through the end of summer.
- Automated leaderboard placed Gemini 3 Pro, Sonnet 5, and GPT‑5.5 near the top, while Claire's personal preference ranked Sonnet 5 lower.
Connected Companies & Entities
12 Entities mapped“We've got a new model, people, and it's from Anthropic....”
“This episode is brought to you by Runway a new kind of creative platform that has everything you need to generate any image video or piece o...”
“Runway brings together the world's most advanced AI models, which is why enterprises like Microsoft, Robinhood, Amazon, and Adobe, along wit...”
“Runway brings together the world's most advanced AI models, which is why enterprises like Microsoft, Robinhood, Amazon, and Adobe, along wit...”
“Runway brings together the world's most advanced AI models, which is why enterprises like Microsoft, Robinhood, Amazon, and Adobe, along wit...”
“Runway brings together the world's most advanced AI models, which is why enterprises like Microsoft, Robinhood, Amazon, and Adobe, along wit...”
“Runway brings together the world's most advanced AI models, which is why enterprises like Microsoft, Robinhood, Amazon, and Adobe, along wit...”
“hyperagent was built by the team behind Airtable and how I AI listeners get $1,000 in free inference to start building...”
“Tools referenced: Cursor: https://www.cursor.com/...”
“Tools referenced: Cursor: https://www.cursor.com/...”
“Tools referenced: Gemini 3 Pro (Google DeepMind): https://deepmind.google/models/gemini/pro/...”
“Tools referenced: GPT-5.5 (OpenAI): https://openai.com/index/introducing-gpt-5-5/...”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
How I AI: Sonnet 5 Benchmark and Agent Workflows
A two-part How I AI newsletter episode: Claire benchmarks Anthropic’s new Sonnet 5 using a repeatable “How I AI Bench” built with Claude Code, blind-testing Sonnet 5 against Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro across PRDs, prototypes, agentic tasks, and personality. Sonnet 5’s introductory pricing ($2 per million input tokens, $10 per million output tokens through the end of summer) positions it between prior Sonnet releases and Opus; Claire found human judgement diverged sharply from LLM-as-judge scores. Separately, Alessio Fanelli (founder of Kernel Labs) demonstrates managing autonomous coding agents from mobile using OpenAI Symphony, Linear, and cloud VPS, highlights token-cost tracking, skills-file hygiene, and perception tooling (Kernel Labs’ Glimpse) to extend autonomous runs. The issue emphasizes building repeatable benchmarks and operational practices for agentic workflows.
Anthropic's Sonnet 5.5 Overtakes OpenAI's GPT-6 in AI Ranking
Anthropic has released Claude Sonnet 5.5, a midrange AI model that scores 56 points on the Artificial Analysis Intelligence Index, ranking second overall behind its own flagship Opus 5.5 (58 points) and ahead of OpenAI's GPT-6 Astra (53) and GPT-6 Sol (48). Sonnet 5.5 shows an 18-point improvement over its predecessor Sonnet 5, excelling in agentic tasks and office work, nearly matching Opus 5.5, though it lags in factual knowledge. However, the model consumes significantly more tokens per task (about 193,000 in its highest reasoning mode), making it more expensive per task despite the same list price of $2 per million input tokens and $10 per million output tokens. Anthropic claims up to 30% cost reduction for most work due to efficiency. The model is available on major cloud platforms, and Anthropic has introduced distillation safeguards for the first time on a Sonnet model.
GPT‑5.6 Sol Outperforms Claude Fable in Benchmark
Claire Vo published a review (How I AI, Jul 9, 2026) comparing OpenAI’s GPT‑5.6 family (Sol, Terra, Luna) against Anthropic’s Claude Fable 5 and other models using her five‑category “How I AI” benchmark. Vo reports GPT‑5.6 Sol ranked highest on her Claire Weighted Index (70% human taste, 30% Terminal Bench 2.1), winning on prototypes, PRDs, browser automation, and one‑shot product prototypes. She notes Sol’s lower reported API pricing versus Fable, praises Sol’s practical, collaborator‑friendly outputs, and highlights use cases including building a gamified homework app via Codex, automated video clipping, and Chrome/browser automation. Sonnet 5 remains preferred for agentic voice. The piece is a hands‑on model evaluation and feature/use‑case walkthrough rather than a primary product announcement.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
