For Verifying Nondeterministic AI Pipelines, More Re-Executions Do Not Help: a Fixed Threshold Lets 27 of 29 Fabrications Through at k=5
arXiv 2609.10601 (submitted 7 Sep 2026) gives a protocol for trustless verification of compound AI pipelines that tolerates nondeterministic output, a dishonest executing node, and intermittent access to a shared record at once, deciding a challenge on the median of k re-executions with no quorum. The headline is where it breaks: on a synthetic HotpotQA pipeline a calibrated fixed threshold accepts 44 of 45 honest reproductions and rejects 104 of 105 divergent pairs, yet still passes same-input fabrication in 27 of 29 trials at k=5, because more sampling sharpens the estimate without moving it. A threshold derived per execution catches 19 of 29 where the best constant reaches 9, so the binding constraint is threshold choice, not sample size.
↳ Follow the thread