WebWorld Uses the Browser as the Judge a VLM Cannot Fool, Beating Its Own Base Model by 14.9 Points
Self-improving web code normally has the same VLM proposing and judging the repair, so visual plausibility substitutes for whether the page works; WebWorld makes the browser the counterparty instead, compiling each VLM critique into a typed interaction contract and re-executing the candidate. An acceptance certificate issues only when target progress holds and every previously verified capability is preserved, and only certified transitions reach the SFT export, forming a quality ratchet. WebWorld-27B improves its raw base by 5.3 points on HTMLBench-400 and 14.9 points on MiniAppBench-Val, reaching Kimi-K2.6 and GPT-5.4 levels on interactive HTML generation, and equal-size ablations show the lift nearly vanishes without the certificate.
↳ Follow the thread