Observed Signal · May 14, 2026 · Funding · Source: techcrunch · Impact: 3/5 · Sentiment: Positive
Former Meta News Chief Launches Forum AI to Judge LLMs
Campbell Brown, formerly Meta’s news chief, founded Forum AI to evaluate foundation models on "high‑stakes" topics such as geopolitics, mental health, finance and hiring. Forum AI recruits domain experts to design benchmarks and trains "AI judges" to evaluate models at scale; Brown says the group has achieved roughly 90% consensus between model evaluations and human experts. Founded about 17 months ago in New York, Forum AI raised a $3 million round led by Lerer Hippeau. Brown argues current model evaluation and compliance practices are inadequate and hopes enterprise demand for reliable AI (for credit, hiring, insurance, etc.) will drive improvements in accuracy and auditability.
Forum AI is building domain-expert benchmarks and scalable evaluation (AI judges) for foundation models on high-stakes topics; that work could influence trust, compliance and enterprise adoption of LLMs, addressing model accuracy and auditability gaps relevant across AI, media and advertising ecosystems.
Track Meta Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Campbell Brown founded Forum AI approximately 17 months ago in New York.
- Forum AI evaluates foundation models on "high-stakes topics" including geopolitics, mental health, finance and hiring.
- Forum AI recruited experts — including Niall Ferguson, Fareed Zakaria, Tony Blinken, Kevin McCarthy and Anne Neuberger — to architect benchmarks for geopolitics work.
- Forum AI trains "AI judges" and reports achieving roughly 90% consensus between those judges and human experts.
- Forum AI raised $3 million in a funding round led by Lerer Hippeau (announced last fall).
Connected Companies & Entities
3 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
In‑House Model Fever: Companies Build Their Own LLMs
The AI Secret newsletter reports a growing trend of companies building proprietary, narrowly specialized models instead of renting frontier models — exemplified by Base44's Base One and MyClaw.ai's upcoming MyClaw Pro. Meta released Brain2Qwerty v2, a non‑invasive brain‑activity decoder and open‑source research release that reached substantially higher word‑accuracy than prior non‑invasive systems. The newsletter also highlights AI-driven disruption in animation production reducing costs and jobs, and a revived Health and Location Data Protection Act introduced by Elizabeth Warren and Mary Gay Scanlon that would ban selling health and location data to brokers and explicitly cover user inputs to AI systems. A short TL;DR lists other industry moves (Anthropic, Google/Gemini, Adobe & Disney, OpenAI partnerships) underscoring rapid commercial, technical and regulatory shifts affecting AI adoption.
AEF-1 Standard for Third-Party AI Evaluators Emerges; xAI, OpenAI, Anthropic Co-sign
The AI Evaluator Forum (AEF) published AEF-1, a proposed baseline standard for independent third-party evaluations of frontier AI systems, covering access, conflict of interest, funding, recusal, and transparency. Major AI labs including xAI, OpenAI, and Anthropic co-signed the standard. The article also covers a public safety debate around 'Pacing the Frontier', where Anthropic's Dario Amodei proposes embedded evaluators with unprecedented access, amid criticism of potential conflicts of interest within the Anthropic-linked safety ecosystem. Other topics include agent harness engineering, new model releases (DeepSeek-V4.1-Flash, Cohere Parse 5), and robotics foundation models.
Arena: The Unbiased Leaderboard Shaping AI's Future
Arena (formerly LM Arena) has emerged as the public leaderboard for frontier large language models, influencing funding, product launches and PR cycles. Originating as a UC Berkeley PhD research project, the startup scaled rapidly and reached a reported $1.7 billion valuation within seven months. TechCrunch’s Equity host Rebecca Bellan interviewed Arena co-founders Anastasios Angelopoulos and Wei-Lin Chiang about how the platform operates, its approach to “structural neutrality,” and why Arena’s dynamic evaluations are harder to game than static benchmarks. The piece notes that companies including OpenAI, Google and Anthropic back the project, that Anthropic’s Claude currently tops expert leaderboards in legal and medical tasks, and that Arena is expanding beyond chat to benchmark agents, coding and real-world tasks with a new enterprise product.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
