Observed Signal · Jul 29, 2026 · Technical Release · Source: OpenAI Blog · Impact: 4/5 · Sentiment: Positive

Two API Settings Tripled ARC‑AGI‑3 Scores

Executive Signal Summary

OpenAI reports that enabling two Responses API settings—retained reasoning and compaction—substantially improved agent performance on the ARC‑AGI‑3 2D puzzle benchmark. Using their Responses API harness (which retains private reasoning messages across turns and compacts long histories instead of rolling truncation), GPT‑5.6 Sol's score on the ARC‑AGI‑3 public set rose from 13.3% with the official harness to 38.3%, roughly a 3x improvement, while output tokens fell by about 6x. The post explains that the official ARC harness discarded private reasoning and used rolling truncation, which prevented models from carrying forward internal thoughts and older actions. OpenAI recommends using retained reasoning and compaction in evaluations to better match production deployments like ChatGPT and Codex.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

A major AI provider (OpenAI) demonstrates that evaluation harness choices (retained reasoning and compaction) materially change benchmark results and token costs; this affects reproducibility, model comparisons, and best practices for deploying LLM agents in production.

SIGNAL RADAR

Track OpenAI Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Enabling retained reasoning and compaction increased GPT‑5.6 Sol's ARC‑AGI‑3 public-set score from 13.3% (official harness) to 38.3% with OpenAI's Responses API harness.
  • OpenAI reports the two settings produced roughly 3x the score while reducing output tokens by about 6x on the public task set.
  • GPT‑5.6 Sol previously scored 7.8% on ARC‑AGI‑3 (overall) and GPT‑5.5 scored 0.4% on the benchmark before harness changes.
  • The official ARC harness discarded private reasoning after each action and used rolling truncation (dropping older messages once context exceeded 175,000 characters), which harmed agent learning over time.
  • OpenAI implemented the ARC‑AGI‑3 harness with its Responses API to retain private reasoning across tool calls and to use compaction instead of rolling truncation.

Connected Companies & Entities

1 Entity mapped

“To better match our production setup, we implemented the ARC-AGI-3 harness with our Responses API. Our API makes it easy to manage context: ...”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: OpenAI Blog•Published: Jul 29, 2026
Original Coverage Title: “How enabling two settings tripled our scores on the ARC-AGI-3 benchmark”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Large Language Models (LLM) & AIJul 9, 2026

GPT-5.6 Scores 7.8% on ARC‑AGI‑3 Benchmark

An analysis of OpenAI’s GPT-5.6 (Sol variant) highlights a surprising benchmark result: the model scored 7.8% on the ARC-AGI-3 benchmark, a test of fluid intelligence that humans typically score above 90%. Although GPT-5.6 shows large gains versus GPT-5.5 (0.43%) and strong performance on other ARC-AGI versions, the low absolute ARC-AGI-3 score reveals remaining gaps in planning and long-chain reasoning despite improved scene comprehension and orientation in novel environments. The author notes the evaluation cost (at max reasoning effort) approached $20,000 and cites ARC Prize commentary that Sol succeeded by correctly orienting itself in new environments but fails more often in deeper planning/execution stages.

Read assessment
Large Language Models (LLM) & AIJul 29, 2026

OpenAI releases GPT-5.6 with major efficiency gains

OpenAI announced the GPT-5.6 model family—flagship GPT-5.6 Sol plus lower-cost Terra and Luna—designed to balance capability and serving cost by routing workloads to appropriate variants. Sol is available in ChatGPT, Codex, and the API with listed pricing of $5 per million input tokens and $30 per million output tokens. The release emphasizes deployment efficiency and capabilities such as Programmatic Tool Calling and multi-agent support, and describes system-level runtime optimizations (load balancing, KV-cache tuning, prompt caching, routing, kernel and implementation improvements) and an agentic harness used by Codex and ChatGPT Work. OpenAI reports that Sol outperforms Claude Fable 5 on a coding-agent index at under half the cost and attributes ~20% lower end-to-end serving costs and >15% higher token-generation efficiency to those optimizations, though some internal figures were not fully documented in first-party materials. Buyers are advised to evaluate end-to-end deployment economics rather than only published token prices.

Read assessment
Large Language Models (LLM) & AIMar 6, 2026

GPT-5.4 Arrives: ChatGPT Reveals Reasoning, Controls Apps

OpenAI is rolling out GPT-5.4 across ChatGPT, the API, and Codex, introducing GPT-5.4 Thinking and GPT-5.4 Pro for more complex tasks. The update presents a reasoning-first interface, showing users the planned solution path and enabling intervention before final answers. GPT-5.4 Thinking will replace GPT-5.2 Thinking for Plus, Team, and Pro users, with GPT-5.2 remaining as a Legacy option until June 5, 2026. OpenAI reports reliability gains, citing a 33% reduction in incorrect statements versus GPT-5.2 and an 18% decrease in errors in complete answers. In GDPval benchmarks across 44 professions, GPT-5.4 meets or surpasses industry experts in 83% of cases. The release also includes an Excel Add-in enabling natural-language creation, analysis, and updating of tables. Overall, OpenAI emphasizes embedding AI more deeply into real-world workflows and enabling agents to operate software across environments.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.