GDPevo measures agent self-evolution on enterprise workflows: +16.44 points from experience, still far below a 91.6% oracle
GDPevo grades whether an agent's persistent state actually transfers, using "rule hybridization" — decomposing enterprise workflows into atomic business rules, splitting them across training tasks, and recombining them in held-out tests so gains are attributable rather than contaminated. V1 covers CRM, ERP, finance, healthcare, legal, and data-centric workflows with 120 tasks in 12 groups; the fully automated pipeline regenerated a 240-task V2 in two days as a contamination defense. Across four agents (harness plus model) under four supervision types, self-evolution improved held-out accuracy by up to 16.44 percentage points while the best evolved agents stayed far under the fully informed oracle ceiling of 91.6%.
Source
↳ Follow the thread