Observed Signal · Jun 15, 2026 · Interview / Podcast Episode · Source: Lennys Newsletter · Impact: 2/5 · Sentiment: Positive
Braintrust Uses AI Agents, Evals, and CI to Ship Software
Ankur Goyal, founder and CEO of Braintrust, discusses how his company uses AI agents, automated evals, and improved CI to accelerate engineering velocity and product quality. He describes running exhaustive, week‑long benchmark experiments with coding agents (using Codex) across database indexes, column-store formats and execution engines; introduces the “agent line” framework to decide what to delegate to agents; and explains how evals function as modern PRDs and scoring functions to encode “what good looks like.” The conversation covers operational patterns (foreground vs background agents, concurrent agents), capturing designer taste through repeatable evals, prompt iteration inside safe playgrounds, and why investing in CI/CD is high leverage for AI‑accelerated engineering teams.
Practical discussion of AI agents, automated evals, and CI provides concrete operational patterns for engineering teams to accelerate development and quality. The content is relevant to AI/LLM practitioners but is not an industry‑shifting platform announcement.
Track Vercel Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Ankur Goyal is founder and CEO of Braintrust, an AI evals and observability platform.
- Braintrust is used by teams including Notion, Stripe, Vercel, and Zapier (as named in the episode).
- Goyal described using Codex to run week‑long benchmarking experiments across database indexes, column store formats, and execution engines.
- He presented the “agent line” framework for deciding which decisions and interactions to delegate to AI agents.
- The episode argues that automated evals can serve as modern PRDs and that fixing CI/CD is critical to increase engineering velocity for AI workflows.
Connected Companies & Entities
5 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Evals Become the Modern PRD for AI Product Teams
A podcast episode and live demo with Ankur Goyal (Founder & CEO of Braintrust) argues that structured "evals"—datasets, task definitions and scoring functions—should replace traditional PRDs for AI product development. Braintrust, which announced a Series B at an $800 million valuation, powers eval workflows for customers including Replit, Vercel, Airtable, Ramp, Zapier and Notion. In the demo the hosts connected to Linear’s MCP server, auto-generated test data with Opus, iterated prompts and scoring functions, and improved a model score from 0 to 0.75. The piece explains offline vs. online evals, the data-task-scores framework, and recommends PMs own evals and scoring functions to create durable, model-agnostic product specifications.
Anthropic's AI Tools Reshape Software Engineering
The Pragmatic Engineer visited Anthropic’s San Francisco lab and interviewed four engineers to describe how improved AI tooling is changing software development. Key examples: the Claude Platform team built and launched Claude Managed Agents after a six-month project and re-architected its platform layer (migrating from Python to Rust); Bun creator Jarred Sumner completed a Zig→Rust rewrite in 11 days using 64 parallel AI agents and about $165,000 in tokens, with substantial verification and testing work after the initial implementation. The article documents shifts in team practices — more agent-driven prototyping, heavy use of automated code review and security scanners, increased fluidity between teams, and continued reliance on planning and PRDs for complex projects.
Coinbase Scaled AI Across 1,000+ Engineers
Chintan Turakhia, Senior Director of Engineering at Coinbase, describes how his team used AI tools and custom agents to transform a 1,000+ engineer organization and accelerate product development. Tasked with rewriting Coinbase’s self-custody wallet into a consumer social app in six to nine months, the team used AI as a force multiplier to cut PR review times from 150 hours to 15 hours, compress feedback-to-release cycles, and run a “PR speed run” where 100 engineers pushed 70 PRs in 15 minutes. The discussion covers leadership demonstration, hands-on adoption, metrics for engineering velocity, integrating tools like Cursor, Linear, Slack, GitHub Copilot, ChatGPT and Claude, building custom Slack bots and agents, and demos for real-time feedback capture and feature delivery.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
