Fetching from the wire…
Public story · 2026-07-27 · high
Grading choices alone swing the score up to elevenfold on the same outputs, more than changing which model you test.
Why now: The paper, arXiv 2607.23425, posted in July 2026, puts a number on how much grading choices alone can move a formal-methods benchmark score.
TLA+-Bench finds the best AI model writes a correct formal specification only 16% of the time, per the benchmark paper posted to arXiv as 2607.23425.
The benchmark checks each spec by running it through TLA+'s model checker over the full reachable state space. A spec doesn't just need to read correctly, it has to verify.
Grading the same fixed set of outputs with different but equally reasonable rules moved the correct rate from 10.0% down to 1.7%. That's a sixfold swing on identical answers. Add whether graders hand over the interface and configuration names upfront, and the gap widens to elevenfold.
Inside that envelope, the model rankings hold steady. The strongest model scored 16% correct by default and 26% when given configuration names outright. Open-weight models never cleared 1%, and every model tested wrote valid TLA+ syntax far more often than TLA+ that actually verified.
The dataset pairs 403 gold specs, each shipped with a runnable model-checker config, against 897 parse-only silver specs pulled from 13 repositories.
No TLA+-generation number published before this paper is safely comparable to another, since none of them documented which grading choices produced it. That's the real story here, not that AI struggles with formal methods, which was already expected. Watch whether future formal-methods benchmarks start disclosing their grading harness now that the size of the gap is measured.
The paper, arXiv 2607.23425, posted in July 2026, puts a number on how much grading choices alone can move a formal-methods benchmark score.
Each link below shares sources, entities, or timing with this story.
Shared entity: Bench / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Bench; reported by the same outlet (arxiv.org); overlapping topics (best, model).
Shared entity: Bench / Shared topic / Earlier coverage / Tension
Both cover Bench; overlapping topics (benchmark, best, config, configuration); earlier Bench coverage from 2026-06-19.
Shared entity: Bench / Same source domain / Shared topic / Earlier coverage
Both cover Bench; reported by the same outlet (arxiv.org); overlapping topics (benchmark, configuration, model).
Both cover Bench; reported by the same outlet (arxiv.org); overlapping topics (benchmark, model).
Shared entity: Bench / Same source domain / Earlier coverage / Tension
Both cover Bench; reported by the same outlet (arxiv.org); earlier Bench coverage from 2026-07-22.
Shared entity: Bench / Shared topic / Earlier coverage / Tension
Both cover Bench; overlapping topics (benchmark, correct); earlier Bench coverage from 2026-07-15.
Both cover Bench; overlapping topics (benchmark, model); earlier Bench coverage from 2026-07-14.
Shared entity: Bench / Same source domain / Earlier coverage / Tension
Both cover Bench; reported by the same outlet (arxiv.org); earlier Bench coverage from 2026-06-18.