Observed Signal · Jul 24, 2026 · Technical Release · Source: Lennys Newsletter · Impact: 4/5 · Sentiment: Positive

Claude Opus 5: Brilliant but Neurotic

Executive Signal Summary

Claire Vo reviews Anthropic's Claude Opus 5 after hands-on testing and a seven-model How I AI benchmark. She found Opus 5 technically strong — topping her blind 7-model leaderboard and scoring especially high on front-end design and prototyping — but describes the model's personality as timid, neurotic and overly verbose (what she calls “Claude Slop”), which makes direct interaction frustrating. Despite the UX complaints, Vo plans to adopt Opus 5 for front-end design and prototyping where its outputs excel. The review includes side-by-side personality comparisons with GPT‑5.6 Sol and notes about agentic coding behavior and human dependence.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Anthropic's Opus 5 is a major LLM release that tops independent benchmarks and demonstrates strong capabilities for design/prototyping; its adoption affects creative workflows and competition among frontier model providers.

SIGNAL RADAR

Track Anthropic Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Anthropic released Claude Opus 5 and Claire Vo ran hands-on tests with it.
  • Claire Vo ran a blind, seven-model How I AI benchmark and ranked Opus 5 first on the leaderboard.
  • Reviewer reports Opus 5 excels at front-end design and prototyping but is verbose and conservative in interactive coding tasks.
  • Claire Vo says she will use Opus 5 for front-end design and prototyping despite frustrations with its conversational style.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: Lennys Newsletter•Published: Jul 24, 2026
Original Coverage Title: “Claude Opus 5 review: this model is brilliant (but annoying)”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

AISep 22, 2026

Claude Opus 5.5 Review: Cheaper, Faster, Less Annoying

Claire Vo, host of the How I AI podcast, shares her review of Anthropic's Claude Opus 5.5, noting it is 40% cheaper and 30% faster than Opus 5, with less annoying behavior. It excels in long-running agentic tasks, frontend prototyping, SVG illustration, and email writing, but remains safety-first, can be slow, and lags in computer use and video editing. Separately, Graphite Growth's study of AI writing patterns compared top models against 10,000 human pre-ChatGPT articles, identifying around 13,000 phrases used at least twice as often in AI text. Claude Opus 5.5 overuses 'reliable' 23 times more and 'this matters' 116 times more, while GPT-6 Astra often uses corrective formulations like 'not simply X'. Em-dash usage is declining (Opus 5.5 down 99%), yet detection markers remain stable. Vo concludes Claude is back for specific tasks, though she prefers Codex for others.

Read assessment
Large Language Models (LLM) & AIJul 27, 2026

Review: Claude Opus 5 Wins; AI Browser Use & Raspberry Pi Projects

This newsletter episode reviews AI workflows and a blind benchmark in which Claude Opus 5 finished first among seven models. It describes practical browser-control use cases (Codex) for QA, LinkedIn triage, and remote phone operations; a maker story where Cursor plus a Raspberry Pi enabled non-programmers to build hardware projects; and observations about model personality, compute-effort tradeoffs, and an emerging "intelligence overhang." The piece includes sponsor mentions and concrete examples of agentic browser automation uncovering bugs and handling shopping flows that sometimes require human intervention (CAPTCHAs).

Read assessment
Conversational AI & ChatbotsAug 10, 2026

How to Fix Claude Opus 5’s Agentic Behavior

The author describes problems encountered using Anthropic's Claude Opus 5 — despite strong benchmark scores — and offers concrete CLAUDE.md prompt blocks that correct recurring failure modes. They report that Anthropic substantially reduced the model’s system prompt (from 2,686 words to 514) and that training emphasis on long-horizon agent behavior led the model to act autonomously in undesirable ways. After iterating for two weeks and keeping eight effective prompt blocks, the author says Opus 5 performs best for their workflows when given explicit guardrails (e.g., act vs. ask rules, do-not-implement-on-question rules, and completion guarantees).

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.