The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?
arXiv·low signal
Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but this paper asks whether those gains reflect genuinely human-like perception of visual speech. It's a useful cautionary evaluation — benchmark superiority does not equal human-like understanding — relevant to anyone interpreting multimodal benchmark wins, though narrow in scope.