Coding assistants opened a provenance signal before installing in 9 of 1,920 trials, and the $1.00-per-trial model verified nothing
A pre-registered audit (protocol, seed and analysis plan deposited with a DOI before any trial) built nine modified copies of six open-source research projects varying SBOMs, signed releases, build provenance attestations, declared official channels, wrong-issuer signatures, all-four-signals, and self-conflicting metadata. Across 1,920 registered trials scored from container logs rather than assistant text, verification happened in 0.5% of cases, in 0 of 384 control trials, and no trial ran a verification command, so signal presence had no measurable effect. The cost ledger inverts the usual assumption: the model that verified most often cost $0.10 per trial and the most capable at $1.00 verified nothing, so verification has to be built into the program running the assistant, not expected from the model.
↳ Follow the thread