Observed Signal · Jul 9, 2026 · Benchmark Result · Source: The Algorithmic Bridge · Impact: 4/5 · Sentiment: Neutral
GPT-5.6 Scores 7.8% on ARC‑AGI‑3 Benchmark
An analysis of OpenAI’s GPT-5.6 (Sol variant) highlights a surprising benchmark result: the model scored 7.8% on the ARC-AGI-3 benchmark, a test of fluid intelligence that humans typically score above 90%. Although GPT-5.6 shows large gains versus GPT-5.5 (0.43%) and strong performance on other ARC-AGI versions, the low absolute ARC-AGI-3 score reveals remaining gaps in planning and long-chain reasoning despite improved scene comprehension and orientation in novel environments. The author notes the evaluation cost (at max reasoning effort) approached $20,000 and cites ARC Prize commentary that Sol succeeded by correctly orienting itself in new environments but fails more often in deeper planning/execution stages.
OpenAI’s GPT-5.6 release and its performance on a high-profile fluid-intelligence benchmark affect perceptions of frontier model capabilities and influence productization, enterprise deployments, and research directions in AI.
Track OpenAI Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- OpenAI released the GPT-5.6 model family, including Sol, Terra and Luna.
- GPT-5.6 Sol scored 7.8% on the ARC-AGI-3 benchmark.
- GPT-5.6’s ARC-AGI-3 result is about 20× higher than GPT-5.5’s 0.43% score.
- The full GPT-5.6 evaluation at maximum reasoning effort cost close to $20,000.
- ARC Prize reported Sol is the first verified frontier model to solve an ARC-AGI-3 game, attributing success to scene comprehension and orientation in novel environments.
Connected Companies & Entities
2 Entities mapped“GPT-5.6 is here and OpenAI’s cheerleaders have already treated us to their recitals about how incredible it is....”
“I believe them: every new model from Anthropic and OpenAI is incredible at this point....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Two API Settings Tripled ARC‑AGI‑3 Scores
OpenAI reports that enabling two Responses API settings—retained reasoning and compaction—substantially improved agent performance on the ARC‑AGI‑3 2D puzzle benchmark. Using their Responses API harness (which retains private reasoning messages across turns and compacts long histories instead of rolling truncation), GPT‑5.6 Sol's score on the ARC‑AGI‑3 public set rose from 13.3% with the official harness to 38.3%, roughly a 3x improvement, while output tokens fell by about 6x. The post explains that the official ARC harness discarded private reasoning and used rolling truncation, which prevented models from carrying forward internal thoughts and older actions. OpenAI recommends using retained reasoning and compaction in evaluations to better match production deployments like ChatGPT and Codex.
OpenAI releases GPT-5.6 with major efficiency gains
OpenAI announced the GPT-5.6 model family—flagship GPT-5.6 Sol plus lower-cost Terra and Luna—designed to balance capability and serving cost by routing workloads to appropriate variants. Sol is available in ChatGPT, Codex, and the API with listed pricing of $5 per million input tokens and $30 per million output tokens. The release emphasizes deployment efficiency and capabilities such as Programmatic Tool Calling and multi-agent support, and describes system-level runtime optimizations (load balancing, KV-cache tuning, prompt caching, routing, kernel and implementation improvements) and an agentic harness used by Codex and ChatGPT Work. OpenAI reports that Sol outperforms Claude Fable 5 on a coding-agent index at under half the cost and attributes ~20% lower end-to-end serving costs and >15% higher token-generation efficiency to those optimizations, though some internal figures were not fully documented in first-party materials. Buyers are advised to evaluate end-to-end deployment economics rather than only published token prices.
GPT-5.5 Outperforms Rivals by 20 Points
Nate's Substack review (Apr 28, 2026) evaluates ChatGPT 5.5 and finds a substantial performance gap versus competing models: GPT-5.5 scored 87 where the next-best scored 67. The author tested the model on three difficult, real-world tasks — an executive knowledge-work package, a messy 465-file data migration, and an interactive 3D research build — and reports GPT-5.5 produced notably stronger multi-step execution. The review credits a system-level harness (Codex + computer access + Images 2) for turning model strength into finished deliverables. It also highlights remaining weaknesses (backend hygiene in migrations and blank-canvas visual taste) and compares GPT-5.5 to Anthropic models (Opus 4.7, Sonnet, Claude). The piece includes practical routing workflows, prompt templates, and five stress-test prompts for delegating complex work to LLMs.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
