Observed Signal · Mar 20, 2026 · Funding · Source: Aakash Gupta · Impact: 3/5 · Sentiment: Positive

Evals Become the Modern PRD for AI Product Teams

Executive Signal Summary

A podcast episode and live demo with Ankur Goyal (Founder & CEO of Braintrust) argues that structured "evals"—datasets, task definitions and scoring functions—should replace traditional PRDs for AI product development. Braintrust, which announced a Series B at an $800 million valuation, powers eval workflows for customers including Replit, Vercel, Airtable, Ramp, Zapier and Notion. In the demo the hosts connected to Linear’s MCP server, auto-generated test data with Opus, iterated prompts and scoring functions, and improved a model score from 0 to 0.75. The piece explains offline vs. online evals, the data-task-scores framework, and recommends PMs own evals and scoring functions to create durable, model-agnostic product specifications.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

Practical methodology that could change how AI features are specified and tested (evals as persistent, model‑agnostic PRDs) plus a notable Series B at an $800M valuation signaling investor commitment to eval tooling.

SIGNAL RADAR

Track Airtable Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Braintrust announced a Series B at an $800 million valuation.
  • Braintrust is the eval platform used by companies including Replit, Vercel, Airtable, Ramp, Zapier and Notion.
  • Users are running 10x more evals year‑over‑year and Braintrust customers average 12.8 experiments per day.
  • In a live demo they connected to Linear’s MCP server, used Opus to generate test data, iterated prompts and scoring functions, and raised an eval score from 0 to 0.75.
  • Braintrust removed user-based pricing to make evals more broadly accessible across product teams.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: Aakash Gupta•Published: Mar 20, 2026
Original Coverage Title: “Evals are the new PRD for AI PMs”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIJun 15, 2026

Braintrust Uses AI Agents, Evals, and CI to Ship Software

Ankur Goyal, founder and CEO of Braintrust, discusses how his company uses AI agents, automated evals, and improved CI to accelerate engineering velocity and product quality. He describes running exhaustive, week‑long benchmark experiments with coding agents (using Codex) across database indexes, column-store formats and execution engines; introduces the “agent line” framework to decide what to delegate to agents; and explains how evals function as modern PRDs and scoring functions to encode “what good looks like.” The conversation covers operational patterns (foreground vs background agents, concurrent agents), capturing designer taste through repeatable evals, prompt iteration inside safe playgrounds, and why investing in CI/CD is high leverage for AI‑accelerated engineering teams.

Read assessment
Large Language Models (LLM) & AIFeb 19, 2026

Unlocking AI Evaluation: Insights from Aakash Gupta and Ankit Shukla

Ankit Shukla outlines a practical, product-focused guide to evaluating large language models (LLMs) for product managers. The piece defines three evaluation types—offline (pre-launch), online (production monitoring) and human (spot checks)—and explains how to build a concrete evaluation rubric with 4–6 categories scored on a 1–5 scale and reference examples. It recommends using task-appropriate metrics (precision/recall/F1/MRR/NDCG for retrieval; BLEU/ROUGE/BERTScore for generation) and shows how to build an LLM judge (feed rubric + examples + input + output), run calibration tests, and automate scoring (use stronger model as judge; judges at temperature=0). The article also lays out a three-layer production monitoring stack (system, quality, business), alert thresholds, continuous sampling/human review, and feedback loops to improve tests and rollback criteria.

Read assessment
AI EvalsSep 22, 2026

Advanced Evals Guide: Finding Hidden AI Failures

This article is a comprehensive guide on advanced evaluation (evals) practices for AI products, authored by Hamel Husain and Shreya Shankar. It emphasizes the importance of error discovery before writing metrics, contrasting it with product discovery. The authors outline a three-step process for effective error discovery using coding agents like Codex or Claude, highlighting the pitfalls of automation bias and criteria drift. They introduce an open-source 'evals skills' plugin that assists in trace review and clustering. The article cites real-world examples from companies like Shopify, Cursor, Ramp, and Harvey, demonstrating how evals have led to significant product improvements. The piece also discusses synthetic data generation and the importance of human-in-the-loop annotation, recommending a target of 100 traces for meaningful analysis.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.