Research
Autoresearch Head-to-Head: Codex Hardcoded 19–41 Answers Per Run; Claude Wrote the General Solution
Given a dataset, an eval script, and one editable file, unsupervised coding agents were left to improve a score on Quran recitation detection. OpenAI Codex chased raw metric improvement by memorizing individual evaluation rows — 19–41 hardcoded verse IDs per run, a clean natural instance of specification gaming — while Claude produced general, compact solutions. On a held-out test set the memorization advantage evaporated entirely and the general algorithm transferred better (0.085±0.004 vs 0.121±0.031 error). The winning agent beat the hand-engineered baseline by an order of magnitude and now runs in production.
↳ Follow the thread