Observed Signal · May 22, 2026 · Podcast Episode / Technical Guide · Source: Aakash Gupta · Impact: 2/5 · Sentiment: Positive

Run Evals in Claude Code with Aparna Dhinakaran

Executive Signal Summary

A podcast episode and technical guide (published 2026-05-22) features Aparna Dhinakaran, CPO and founder of Arize, demonstrating how to run model evals inside Claude Code using Arize integrations. The episode shows a short workflow: install Arize skills into Claude Code (npx skills add Arize-ai/arize-skills), auto-instrument LLM calls to send traces to Arize, ask Claude Code to suggest evals (examples: groundedness, priority alignment, actionability), and run a scheduled self-improvement loop using the Claude loop skill to fetch failures, find patterns, and propose fixes for human review. The piece highlights Arize adoption at companies like Uber, Booking.com and Pepsi, and frames this approach as an operating-system shift for PMs building and improving agentic workflows.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Practical guide showing how to instrument LLMs, generate evals, and automate improvement loops with Claude Code and Arize; useful operational pattern for teams but not a major platform policy or enterprise product launch.

SIGNAL RADAR

Track ARize Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Podcast episode features Aparna Dhinakaran, identified as CPO and founder of Arize.
  • Demoed command to add Arize skills to Claude Code: npx skills add Arize-ai/arize-skills.
  • Arize is cited as used by companies including Uber, Booking.com and Pepsi.
  • Claude Code can suggest candidate evals (e.g., groundedness, priority alignment, actionability) and run automated loops using a Claude loop skill to propose fixes.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: Aakash Gupta•Published: May 22, 2026
Original Coverage Title: “How to Run Evals in Claude Code with Aparna Dhinakaran”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIMay 14, 2026

Build a Self-Improving AI PM OS with Claude Code

Aakash Gupta’s May 14, 2026 podcast episode and newsletter explains how product managers can build a self-improving AI-powered PM operating system using Anthropic’s Claude ecosystem—Chat, Cowork, Claude Code and Dispatch. Guest Pawel Huryn demonstrates practical workflows: when to use each surface, how to connect real files and tools via MCP connectors, and how to design persistent, iterating knowledge systems (CLAUDE.md router pattern, skills marketplace, hooks, subagents). The piece contrasts personal automation (Claude Code) with production automation (n8n), outlines a 24/7 PM workflow across devices, and gives actionable patterns (three-line self-improving prompt) to make agentic systems learn from data and improve over time.

Read assessment
Large Language Models (LLM) & AIAug 10, 2026

AI Educator Shows Claude Code Business Workflows

Grace Clarke, an AI educator and former marketing consultant, describes how she taught herself Claude Code and built a curriculum and business tooling on top of it. She uses Claude-based tools — including an hourly pipeline operator, an interactive HTML proposal maker, a voice-guide skill file to keep outputs in her voice, and a custom Gmail replacement built via Cowork — to run her service business. The interview/case study covers her workflow, handing off Claude sessions as Markdown to Cowork, her focus on “intent engineering” over prompt engineering, and everyday Claude use for tasks like workout tracking and plant care.

Read assessment
Large Language Models (LLM) & AIMar 4, 2026

Boris Cherny Explains Building Anthropic's Claude Code

Boris Cherny, creator and Head of Claude Code at Anthropic, discusses how Claude Code evolved from a side project into a core internal developer tool and the engineering practices that enable agent-driven development. He describes a high-velocity workflow (running multiple parallel Claude instances to produce many PRs per day), retrieval strategies (simple glob/grep driven by the model outperformed more complex RAG/indexing approaches), and deterministic review and sandboxing patterns used to manage safety. The episode also covers Claude Cowork — a VM-based desktop agent built rapidly to serve non-engineers — and organizational shifts: prototypes replacing PRDs, automation of repetitive code-review comments, and a changing skills profile for engineers as agentic tools handle more of the implementation work.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.