CogEvol reports a caught reward-hacking episode alongside its production numbers, and open-weights the 4B
arXiv 2608.30968 (submitted 2026-08-31) trains models to turn a course brief into finished slides or a self-contained interactive HTML page in a single pass, replacing multi-turn agent scaffolding. Across 220k production requests the median slide takes 17 seconds and an interactive page 59; a production-grounded pipeline turned real failures into 53,687 verified SFT samples, and a hybrid rule-plus-VLM reward drives GRPO, hardened after the team caught a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and CogEvol-4B is Apache 2.0 on GitHub.
↳ Follow the thread