Observed Signal · Jul 6, 2026 · Technical Release · Source: Lennys Newsletter · Impact: 3/5 · Sentiment: Neutral
How I AI: Sonnet 5 Benchmark and Agent Workflows
A two-part How I AI newsletter episode: Claire benchmarks Anthropic’s new Sonnet 5 using a repeatable “How I AI Bench” built with Claude Code, blind-testing Sonnet 5 against Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro across PRDs, prototypes, agentic tasks, and personality. Sonnet 5’s introductory pricing ($2 per million input tokens, $10 per million output tokens through the end of summer) positions it between prior Sonnet releases and Opus; Claire found human judgement diverged sharply from LLM-as-judge scores. Separately, Alessio Fanelli (founder of Kernel Labs) demonstrates managing autonomous coding agents from mobile using OpenAI Symphony, Linear, and cloud VPS, highlights token-cost tracking, skills-file hygiene, and perception tooling (Kernel Labs’ Glimpse) to extend autonomous runs. The issue emphasizes building repeatable benchmarks and operational practices for agentic workflows.
Provides benchmark results, pricing, and operational lessons for a newly released major model (Sonnet 5) and practical guidance on agent management — relevant to builders choosing LLMs and designing agentic workflows but not a platform-wide policy or industry-shifting regulatory change.
Track Anthropic Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Anthropic released Sonnet 5 and Claire benchmarked it using a repeatable How I AI Bench.
- Sonnet 5 introductory price: $2 per million input tokens and $10 per million output tokens through the end of summer.
- Claire blind-tested Sonnet 5 against Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro across PRDs, prototypes, agentic tasks, and personality.
- Claude Code was used to build the benchmark (it can read old session history to generate tailored eval tasks) and to create an HTML scoring page that exports JSON.
- Alessio Fanelli (founder of Kernel Labs) demoed managing autonomous coding agents using OpenAI Symphony, Linear, a cloud VPS, and Kernel Labs’ Glimpse (a Playwright extension).
Connected Companies & Entities
8 Entities mapped“Claire puts Anthropic’s new Sonnet 5 through a real benchmark....”
“Alessio Fanelli ... shows Claire how he manages autonomous coding agents from his phone using OpenAI Symphony, Linear, and a cloud VPS....”
“She blind-tests Sonnet 5 against Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro across PRDs, prototypes, agentic tasks, and agent personali...”
“Brought to you by: Runway—The creative AI platform for images, video, and more...”
“Brought to you by: Jira Product Discovery—Prioritize with insights, build with confidence...”
“Alessio Fanelli ... manages autonomous coding agents from his phone using OpenAI Symphony, Linear, and a cloud VPS....”
“Alessio showed tasks ranging from 15 million to 221 million tokens, and the 221-million-token job (making an app deployable on Vercel) made ...”
“He demos a very different use case: using Codex with browser access to hunt for underpriced Pokémon cards on eBay....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Sonnet 5 review: 64-run How I AI benchmark
Claire Vo built the "How I AI Bench" live using Claude Code and ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT‑5.5, Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The benchmark comprised ~64 generations and combined human "vibe" scoring (70%) with LLM-as-judge scoring (30%). Results surprised the author: Gemini 3 Pro, Sonnet 5, and GPT‑5.5 ranked highly on the automated leaderboard, but Claire's personal taste favored different models (Sonnet 4.6 / Opus 4.8). The piece also notes Sonnet 5's introductory pricing and Anthropic's positioning of Sonnet 5 as a lower-cost, more agentic model for running tool-using workflows.
How I AI: GPT-5.6, Local AI Fleets, Agent Harnesses
This newsletter episode reviews GPT-5.6 (Sol) against other models, explains the concept and engineering value of "agent harnesses," and describes a 24/7 local AI fleet built by a solo operator. Claire demonstrates a Claude Agent SDK harness used to automate Sentry bug triage with structured outputs and permission encoding. Alex Finn details a multi-machine local setup (Mac Studio, DGX Spark, RTX 5090) that routes workloads across models like GLM, Qwen, and Ornith to make always-on inference economically viable. Claire's benchmark finds GPT-5.6 Sol practically most effective for product work, while also noting model-specific strengths for Terra, Fable, and Sonnet.
Anthropic launches Claude Sonnet 5, cheaper agentic LLM
Anthropic announced Claude Sonnet 5, a midsize, more agentic foundation model intended to run autonomous agent workflows at lower cost. Claude Sonnet 5 can plan, use tools (browsers, terminals) and run end-to-end tasks; it will be the default model for Anthropic’s free and Pro plans. Introductory pricing runs at $2 per million input tokens and $10 per million output tokens through August 31, rising to $3 per million input afterward. Anthropic says Sonnet 5 approaches Opus 4.8’s performance on many tasks while costing less; on an agentic coding benchmark Sonnet 5 scored 63.2% versus Opus 4.8’s 69.2% and Sonnet 4.6’s 58.1%. The company reports safety improvements over Sonnet 4.6, though Opus models and Claude Mythos Preview are still stronger on alignment and dangerous-task resistance. Testers including Zapier highlighted Sonnet 5’s ability to finish complex automation end-to-end.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
