Ofir Press: Anthropic's Opus 5.5 ProgramBench number covers 166 of 200 tasks and counts tests passed, not tasks solved
Latent Space (AINews) / swyx gist·medium signal
ProgramBench co-author Ofir Press pointed out that Anthropic ran 166 of the benchmark's 200 instances and got near-100% test pass rates on that subset, which leaves open whether the hardest programs were dropped. Anthropic also reports average tests passed, while the official ProgramBench leaderboard counts only fully completed tasks, and a partial solve often passes 60-70% of tests. Read the 91.2% system-card figure as a different metric from the leaderboard before comparing models on it.