PlayCoder Benchmark Reveals Near-Zero Logical Correctness in LLM-Generated GUI Applications Despite High Compilation Rates
arXiv·high signal
PlayCoder introduces PlayEval, a benchmark spanning six GUI application categories, and Play@k, a metric that measures whether generated code can be played end-to-end without logic errors. Testing 10 state-of-the-art code LLMs shows that despite high compilation rates, they achieve near-zero Play@3 scores — meaning generated games and GUI apps compile but don't actually work correctly. PlayTester, an LLM-based agent, performs task-oriented GUI playthroughs to detect logic violations automatically.