Humans Catch Partial Audio Deepfakes 9% of the Time and Call the Synthetic Sentence Real in 77% of Cases
Earlier studies reporting 70-80% human deepfake-audio detection relied on older synthesizers, so this study tested 82 IT professionals against tools released in 2019, 2022 and 2024 and benchmarked them against six pretrained detectors on the same material. For fully synthetic speech the human F1 score collapses from about 90% for RTVC and YourTTS to 48% for ElevenLabs, even though listeners were explicitly warned deepfakes were present. For partial spoofing, where a single sentence in an utterance is swapped, strict accuracy falls to 9% and listeners label the synthetic sentence as genuine 77% of the time; humans and detectors fail in complementary ways and neither reliably localizes short manipulations.
↳ Follow the thread