Five reliability properties held while the agents' cognition was ablated, reset and swapped for a frontier LLM panel
Where Reliability Lives (arXiv 2609.03192, v1 2026-09-02) builds a persistent simulated settlement whose append-only ledger adjudicates every attempted act against world state, separating mind, institution and world before any experiment so that "where does this property live" becomes testable. Holding enforcement fixed, the authors ablated the native minds, killed and reset them mid-task, substituted a frozen frontier-LLM panel for the entire native cognition, and corrupted beliefs with trusted false testimony. Behavior changed dramatically, one falsehood cost each trusting run about 900 futile actions, yet five pre-declared properties never moved: accepted reality stayed singular, invalid attempts were refused with typed reasons, duties outlived their processes, no work was accepted twice, and across 2,581 substituted-panel claims no false completion was ever accepted. The authors scope the claim to this one designed world, but the design lesson is direct: put your invariants in an authoritative ledger, not in the prompt.
↳ Follow the thread