RA-Bench Uses Real Videos as Anchors Across 17,886 Clips — and No Detector Family Generalizes, Including Ten Zero-Shot MLLMs
RA-Bench (arXiv 2608.14391, submitted Aug 14; 72 upvotes and today's top HuggingFace Daily Paper, with Chunhua Shen, Yang You and Kaipeng Zhang among 30+ authors) anchors AI-generated-video detection to real footage: 1,830 authentic anchors and 16,056 synthetic clips from four open-source and five closed-source generators, spanning 10 social-risk categories tied to real crisis events. The result is uniformly negative — none of seven traditional detectors, ten zero-shot multimodal models, or two fine-tuned MLLMs generalize consistently. Two secondary findings sharpen it: the videos that fool humans are the same ones that defeat detectors, and simulated social dissemination further degrades detection reliability. For anyone planning to bolt a provenance check onto a media pipeline, this is the paper saying the detector layer is not yet load-bearing.
↳ Follow the thread