GameASG-Bench: 93.2 percent of individual checks pass, 55.3 percent of tasks actually succeed
arXiv 2609.21293 (18 Sep) builds 47 browser-native game generation tasks across 12 genres in 2D and 3D, each with an evaluation interface declared before generation that fixes legal starting scenarios, player-level actions, stable snapshots, rejection behavior and invariants while leaving implementations open. Checks split into static L1 source compliance and browser-executed L2 runtime evidence. Across nine agent stacks the highest mean L2 pass rate is 93.2% but the highest strict task success rate, requiring all applicable L1 and L2 checks, is 55.3% (26 of 47) — the headline argument that averaged check rates hide task-level compliance gaps. For DeepSeek-V4-Flash, full tool access and larger turn budgets helped while reasoning effort did not behave monotonically, and the two tested harnesses each got 18 strict successes but overlapped on only ten tasks.
Source
↳ Follow the thread