Observed Signal · May 4, 2026 · Analysis/Forecast · Source: Import AI · Impact: 4/5 · Sentiment: Neutral

AI Systems May Begin Automating AI R&D

Executive Signal Summary

This Import AI essay argues there is a >60% chance that AI systems capable of conducting end-to-end AI research without human involvement — effectively building their own successors — could emerge by the end of 2028. The author surveys public benchmarks, agentic coding tools, reproducibility and model‑training evaluations (SWE‑Bench, METR time horizons, CORE‑Bench, MLE‑Bench, PostTrainBench) and engineering work (kernel design, training optimization) to show aggregate capability trends. Examples cited include rapid improvements in coding and experiment automation, large speedups in model training optimizations by Anthropic models, and proof‑of‑concept automated alignment research. The piece is an evidence-driven forecast about the plausibility and timing of automated AI R&D and discusses implications for research, governance, and preparedness.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

The essay forecasts a plausible, near-term shift to automated AI R&D, which would materially affect how models are developed, compute demand, engineering workflows, and governance — an industry-level technological inflection with major operational and regulatory implications.

SIGNAL RADAR

Track Anthropic Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • The author estimates a >60% chance that fully no-human-involved AI R&D (an AI system able to plausibly autonomously build its successor) occurs by the end of 2028.
  • SWE-Bench coding benchmark: Claude 2 scored ~2% at late-2023 launch; Claude Mythos Preview later reached ~93.9%, effectively saturating that benchmark.
  • METR 'time horizon' progression reported: GPT-3.5 (2022) ~30 seconds; GPT-4 (2023) ~4 minutes; o1 (2024) ~40 minutes; GPT 5.2 High (2025) ~6 hours; Opus 4.6 (2026) ~12 hours.
  • CORE-Bench reproducibility benchmark: GPT-4o CORE-Agent ~21.5% (Sept 2024); an Opus 4.5 system was later reported to reach 95.5% and the benchmark was declared 'solved' (Dec 2025).
  • Anthropic model training optimization results: Claude Opus 4 achieved 2.9× speedup (May 2025), Opus 4.5 16.5× (Nov 2025), Opus 4.6 30× (Feb 2026), and Claude Mythos Preview 52× (Apr 2026) on a CPU-only training optimization task.
Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: Import AI•Published: May 4, 2026
Original Coverage Title: “Import AI 455: AI systems are about to start building themselves.”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIApr 6, 2026

Studies: AI Raises Cyberattack Capabilities and Automates Work

Import AI summarizes multiple new research reports showing accelerating AI capabilities and mixed economic expectations. Lyptus Research finds a clear scaling trend in AI offensive-cyber tasks (model capability doubling times: 9.8 months since 2019; 5.7 months since 2024), with recent frontier models achieving ~50% success on tasks that take human experts ~3.1–3.2 hours. A field experiment run by INSEAD and Harvard Business School across 515 startups reports that firms taught to integrate AI discovered 44% more AI use cases, completed 12% more tasks, were 18% more likely to acquire paying customers, and generated 1.9x higher revenue. MIT research finds AI automation is a 'rising tide'—projecting most text-based labor tasks could reach 80–95% success by 2029. A Forecasting Research Institute survey shows experts expect relatively rapid AI progress but only modest GDP gains (~+1 percentage point by 2030).

Read assessment
Large Language Models (LLM) & AIMay 10, 2026

Are AI Labs Preparing for an Intelligence Explosion?

Exponential View issue #573 (May 10, 2026) by Azeem Azhar discusses the possibility of frontier AI models training their successors (a scenario Jack Clark estimates as 60% likely by 2028). The piece examines observable behaviours labs would exhibit if they expected rapid automated R&D, arguing that constraints on land, power, chips and skilled electricians make scaling an industrial problem as well as an algorithmic one. Azhar suggests labs would shift hiring toward engineers who can build automated research workflows, and would pre‑commit to large compute, memory and data‑centre capacity — tolerating near‑term cash burn to secure future acceleration.

Read assessment
Large Language Models & AIMay 26, 2026

Anthropic Co‑founder Warns of Rapid AI Singularity

An Import AI newsletter issue presents a long-form essay and Oxford talk reflecting on accelerated AI progress, personal and organizational impacts, and speculative timelines for recursive self-improvement. The author uses the Epoch Capabilities Index to frame recent benchmark successes (e.g., legal and math achievements) and argues continued investment in compute and data makes further rapid advances likely. The piece describes how Anthropic (and its Claude family of models) is increasingly automating coding and analysis—citing an internal release, Opus 4.6, that raised automation and observability needs—and anticipates organizational shifts toward verification, observability, and a new “trust economy.” The author makes dated predictions (e.g., biology relevance by Nov 2026, Nobel-linked discovery by Apr 2027, autonomous revenue-generating companies by Nov 2027, autonomous successor design by Dec 2028) and includes a short fictional story exploring uplift and life-extension themes. The talk was given at Oxford on 2026-05-20 and the piece was published 2026-05-26.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.