TLA+-Bench Grades by Model Checker, Not Resemblance — and Finds the Best Model Writes Correct Specs Only 16% of the Time
TLA+-Bench (arXiv 2607.23425) replaces reference-similarity and parse-success grading with execution: every one of its 403 model-checked gold specifications ships a configuration the TLA+ model checker runs over the full reachable state space, alongside 897 parse-only silver specs drawn from 13 public repositories. Its headline result is about measurement itself — varying only the grading choices earlier benchmarks left unstated moves the correct rate sixfold on one fixed set of outputs (10.0% to 1.7%), and adding the interface-supply choice widens that 'correctness envelope' elevenfold (18.7% to 1.7%). Inside the envelope the findings are stable: every model writes valid TLA+ far more often than correct TLA+, the strongest model is correct 16% of the time by default and 26% when given configuration names, and open models top out at 1%.
↳ Follow the thread