Preregistered Agent-System Experiments Refuted Their Own Central Prediction Twice, in Opposite Directions
arXiv 2609.03192 treats the location of a reliability property, in the model or in the machinery around it, as an experimental question, using a persistent simulated settlement whose append-only ledger adjudicates every attempted act. Holding cognition fixed and intervening on institutional mechanisms, preregistered predictions were refuted twice in opposite directions; holding enforcement fixed and intervening on cognition four ways (ablating the native minds, killing and resetting them mid-task, substituting a frozen frontier-LLM panel, and corrupting beliefs with trusted false testimony) changed behavior dramatically, with one falsehood costing each trusting run about 900 futile actions. Yet five pre-declared properties never moved, including no false completion ever accepted across 2,581 substituted-panel claims, which is direct evidence for putting guarantees in the harness rather than the model.
↳ Follow the thread