Agents
Automated TTS judges collapse onto audio quality — bad news for anyone benchmarking a voice agent
Researchers decomposed 'naturalness' into ten linguistically grounded perceptual dimensions and built a benchmark of 860 utterances annotated by trained linguist raters, then tested both MOS predictors and Audio-LLM judges against it. MOS predictors collapse onto acoustic signal quality rather than the distinct speech aspects listeners actually perceive, and Audio-LLM judges show selective, prompt-dependent detection that does not generalize across dimensions — neither class reliably catches linguistically structured speech errors. The dataset, annotation schema, and evaluation code are public, which makes this immediately usable if you are picking an eval for a voice agent.
Source
↳ Follow the thread