Fetching from the wire…
Research2026-06-14 · source-backed
UOJ-Bench uses real competitive-programming submissions to test generation, error-finding, and repair. In single-attempt evaluation, top models fail to identify errors in over 50% of incorrect submissions (arXiv). Test-time scaling pushes success above 90%, but models also flag problems in more than 5% of perfect submissions. So your code-judging agent both misses real bugs and invents fake ones, and only multi-attempt scaling rescues it. Single-pass agent code review is unreliable by the numbers.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / What happened next
Both cover Bench, Test; reported by the same outlet (arxiv.org); overlapping topics (code, model).
Shared entity: Test / Same source domain / Shared topic / What happened next / Tension
Both cover Test; reported by the same outlet (arxiv.org); overlapping topics (agent, code, model).
Shared entity: Bench / Same source domain / Shared topic / What happened next / Downstream implication
Both cover Bench; reported by the same outlet (arxiv.org); overlapping topics (code, over).
Shared entity: Bench / Same source domain / Shared topic / What happened next / Tension
Both cover Bench; reported by the same outlet (arxiv.org); overlapping topics (agent, model).
Shared entity: Even / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Even; reported by the same outlet (arxiv.org); overlapping topics (agent, code).
Shared entity: Test / Same source domain / Shared topic / Earlier coverage / Downstream implication
Both cover Test; reported by the same outlet (arxiv.org); overlapping topics (agent, error).
Shared entities / Shared topic / What happened next
Both cover Bench, Test; overlapping topics (agent, model); picks up the Bench thread on 2026-08-10.
Shared entity: Test / Same source domain / What happened next / Tension / Downstream implication
Both cover Test; reported by the same outlet (arxiv.org); picks up the Test thread on 2026-07-30.