OpenAI has announced that its GPT-5.6 Sol model achieved a score of 38.3 percent on the ARC-AGI-3 benchmark, surpassing Anthropic's Opus 5, which scored 30.2 percent. However, this achievement was only possible using OpenAI's own custom test harness, which allowed for retained reasoning and context compaction. In a standard testing environment, GPT-5.6 Sol's performance dropped dramatically to just 7.8 percent, highlighting a significant discrepancy in the results depending on the testing conditions. This raises important questions regarding the reliability and comparability of AI benchmarks in the rapidly evolving landscape of artificial intelligence.

Source: The Decoder