OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

The Decoder The Decoder

https://the-decoder.com/wp-content/uploads/2026/06/openai_gpt56_sol.png" style="height: auto; margin-bottom: 10px;" width="1376" />


OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only through its own API with retained reasoning and context compaction.

In the official test environment, the model managed just 7.8 percent.

Opus 5 hit its 30.2 percent without such aids.


The article https://the-decoder.com/openai-claims-gpt-5-6-sol-beats-opus-5-on-arc-agi-3-but-only-with-its-own-custom-test-harness/">OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness appeared first on https://the-decoder.com">The Decoder.

Read full article at The Decoder →