Criterion-revision benchmark CMB-0.1 finds zero passing model trials, and the authors call it an instrument failure
This paper studies whether an agent that fails can form and then persistently use a revised success criterion, requiring five non-compensatory conditions including new-episode transfer and intervention sensitivity on the claimed carrier. Across 84 deterministic scorer trials, 96 calls and 192 model-case-arm trials on four local quantized artifacts, no model trial satisfied all five, but the authors refuse to read that as capability absence: the harness performed the commits, several commitments disclosed the target distinction, and Qwen2.5-7B answered every transfer and preservation item with no revision state at all. They publish CMB-0.4, a trace-anchored successor protocol with concealed transfer and explicit WRITE/NO-WRITE/ESCALATE actions.
Source
↳ Follow the thread