'Count-Scale Drift': Summing Evidence Weights Silently Moves Your Threshold as You Add Sources
This paper separates interpreting a source from aggregating interpretations, proposing a four-field evidence tuple (hypothesis, reliability bucket, rationale, provenance) as the interface, and names a failure mode that hits any additive scoring system: thresholding a sum of unnormalized weights is posterior thresholding at an operating point that slides with the number of sources consulted, and the slide grows with reader reliability. When source reliabilities differ, the vote rule and the posterior order instances differently and no threshold reconciles them; pooling calibrated log-likelihood ratios fixes both. The fix is arithmetic, not architectural, and the authors flag it as applying beyond LLMs to score-summing triage engines and additive multi-signal detectors — plus they state five falsifying predictions and three negative results.
↳ Follow the thread