Top Open-Source ASR Models Reproduce Benchmark Reference Text Verbatim Even When the Audio Is Masked or Contradicts It
arXiv 2608.19936, submitted August 20, presents a method for quantifying benchmark optimization in speech recognition by focusing on cases where the audio underdetermines the reference transcript. Three behavioral probe families (reference disagreement, masked-number recovery, and orthographic switching) show the highest-scoring open-source models emit verbatim reference spans even when the corresponding audio is contradictory, masked, or ambiguous. Mechanistic probing shows the models key off narrow acoustic cues to override faithful transcription. The general lesson beyond ASR: a public leaderboard score can partly measure memorization of the reference set, and probing underdetermined inputs is a cheap way to detect it in any benchmark you rely on.
↳ Follow the thread