Google's Stellar Colosseum posts a 4263 Codeforces rating, above the best human score on record, using a many-agent proof harness
arXiv 2609.15983 (submitted 2026-09-14, revised 09-15) from Honghao Lin, David P. Woodruff, Vahab Mirrokni and colleagues describes a many-agent harness that explores alternative proof strategies in parallel, uses a readiness gate before decomposing a plan into interdependent subproblems, and routes verifier feedback back to the specific part of the argument that failed. It solves 218 of 222 Codeforces problems for a 4263 rating against a 4039 best human score, and hits 71.0% on TCS-Bench (research-level theorem proving drawn from FOCS, STOC and SODA papers) using Gemini 3.1 Pro and Gemini 3.7 Flash. The authors report new results on open problems from FOCS and JMLR papers, and the harness has been folded into Antigravity's Teamwork framework.
↳ Follow the thread