Fetching from the wire…
Public story · 2026-08-23 · high
aaabench hands agents a real game engine and professional-grade conditions, then publishes the test setup but no scores.
Why now: Covered in the August 23 briefing alongside GamePhanes, a second long-horizon game-building eval.
aaabench gives coding agents a real game engine, Unreal, and professional conditions and time to build an open-world game. The repository, from developer ukanwat, has picked up 375 stars and 70 forks since July 31, per GitHub. What's missing is the point: no results, no leaderboard, no scored runs. Just the apparatus.
That's an unusual call for a benchmark. Most eval releases lead with a table of numbers because that's what gets cited. Withholding results while shipping the harness first is defensible here specifically because a benchmark this expensive to run needs the test validated before anyone trusts the scores it produces. Building an open-world game in a professional engine isn't a five-minute eval loop.
aaabench isn't alone. GamePhanes, a Godot-based agent environment and benchmark, hit 106 stars in two days, per its GitHub repository. Two long-horizon game-building evals surfacing side by side says something about where benchmark designers think agents are falling short: not on toy tasks, but on sustained, multi-system work.
Games make sense as the test bed because they split a question that most coding benchmarks collapse into one. Does the code run is one thing. Is the game any good, meaning is it playable, is it fun, does it hold together as a system, is a separate and harder thing. An agent can pass the first test and fail the second, and most benchmarks never find out because they stop checking after the code compiles.
What aaabench doesn't say yet is how it's grading "any good" once it does publish scores, or whether that judgment comes from a person, another model, or some fixed rubric. That's the detail that will decide whether this benchmark tells builders anything they didn't already assume.
Each link below shares sources, entities, or timing with this story.
GamePhanes uses Godot / Shared entities / Same source domain / Shared topic / What happened next / Tension
Linked by a graph relationship (GamePhanes uses Godot); both cover GamePhanes, Godot; reported by the same outlet (github.com).
Shared entity: July / Shared topic / Earlier coverage / Tension
Both cover July; overlapping topics (agent, benchmark, coding, eval, harness); earlier July coverage from 2026-07-23.
Shared entity: July / Same source domain / Shared topic / Earlier coverage
Both cover July; reported by the same outlet (github.com); overlapping topics (agent, alongside, days, star).
Both cover July; reported by the same outlet (github.com); overlapping topics (agent, days, harness, star).
GamePhanes uses Godot / Shared entity: Godot / Same source domain / Earlier coverage
Linked by a graph relationship (GamePhanes uses Godot); both cover Godot; reported by the same outlet (github.com).
Linked by a graph relationship (GamePhanes uses Godot); both cover Godot; reported by the same outlet (github.com).
Shared entity: July / Same source domain / Shared topic / What happened next
Both cover July; reported by the same outlet (github.com); overlapping topics (agent, days, harness).
Shared entity: July / Same source domain / Shared topic / Earlier coverage / Tension
Both cover July; reported by the same outlet (github.com); overlapping topics (agent, days).