Skills
A 32-probe test suite catches 99% of the faults an exhaustive suite catches, at 1.2-2.0% of the executions
arXiv 2608.26746 (2026-08-27) builds compact behavioral test suites for LLM-produced programs by pairing a fault-driven greedy selector with a mutation-independent diversity term covering probe families, cases, templates and time. Across twenty policies and 4.1M+ program-probe executions, 32 probes covered 99.0% of dynamically killable faults using 1.2-2.0% of exhaustive testing; the diversity term lifted scenario-family coverage from 84.6% to 94.9%, and deploying the suite took severe regressions from 15 of 20 program-environment groups to 0. The authors are explicit that this is prioritized evidence, not proof, and publish the budget, evidence source, split, and misses.
↳ Follow the thread