Research
VibeLifeBench: Seven Frontier Models All Score Low on 200 Multi-Week Life-Assistant Tasks in a Silently Changing World
Existing agent evaluations use short self-contained requests in static environments; VibeLifeBench instead scripts 200 long-horizon tasks across ten everyday-life domains, each a multi-week timeline in a simulated world of 22 mock services whose clock advances on its own and whose changes are largely unannounced. Grading is done by fine-grained weighted checks that read only artifacts the agent actually left behind, covering end state, timeliness, and whether unstated constraints were upheld. All seven frontier models evaluated score low, and the authors will open-source the tasks, environments, and evaluation framework.
↳ Follow the thread