The viral 'binaries are now editable code' claim rests on a saturation report the benchmark host's own leaderboard does not show
A 667-upvote r/singularity gallery post amplified a claim, sourced to X posts from Chris and OpenAI's Boris Power, that GPT 'saturated' Vals AI's SRE-Bench for binary reverse engineering less than a month after launch. Fetching vals.ai/benchmarks/srebench directly (updated September 5) shows GPT-5.6 Sol leading with sharply uneven per-domain results across 19 in-house programs averaging 16,900+ lines: RevFirmware 5.00/6 with 83% solved, RevProtocol 4.42/6 at 66%, RevGame 3.52/6 with zero instances fully solved, RevMalware 2.06/6 at 8%. The site's own summary says 'even the strongest cybersecurity-oriented models exhibit limited RE capability,' and success rates roughly halve against protected binaries. The top comment on the Reddit thread makes the same correction, citing 5.5% on the companion ProgramBench.
↳ Follow the thread