OpenAI: Two API Settings Took GPT-5.6 Sol From 13.3% to 38.3% on ARC-AGI-3 With 6x Fewer Output Tokens
OpenAI Blog·high signal
OpenAI published that enabling retained reasoning and compaction on the Responses API tripled GPT-5.6 Sol's ARC-AGI-3 public-set score from 13.3% to 38.3% while cutting output tokens 6x; under the official evaluation harness the model scored as low as 7.8% because its private chain of thought was discarded after every move, forcing it to re-derive the game each turn. Compaction summarizes old context instead of truncating it. The builder takeaway is blunt: published benchmark numbers are measuring harness and API configuration at least as much as model capability.