'Recreating Minecraft Is Not a Benchmark': The Case That Viral AI Demos Measure Eight Weeks of Preparation, Not Capability
Published 2026-09-06 (73 points, 69 comments), the piece argues that showcase artifacts like GPT Astra's Minecraft recreation, the SVG pelican on a bicycle and physics-enabled SVG game controllers are demo-benchmarks labs can optimize for between releases, so they measure preparation rather than capability. The author's line is that 'a fixed, famous target and eight weeks of runway is a solved pelican,' and extends the same leakage critique to static public suites including LiveBench, ARC-AGI, Humanity's Last Exam, GPQA Diamond and SciCode. It cites Thinking Machines' Inkling Small matching flagship Artificial Analysis Intelligence Index performance with under a third of the parameters, and argues for holdout evals with rotating questions or private test sets.
↳ Follow the thread