Research
Vision-Language Models Score Near Chance on Detecting Swapped Video Frames While Humans Are Near Ceiling
TimeCatch reframes temporal grounding as anomaly detection, creating temporal anomalies by swapping consecutive frames and frame-level anomalies by replacing a frame with Gaussian noise, then testing detection and localization across four synthetic and real-world datasets alongside a human study. VLMs consistently detect and often accurately localize frame-level anomalies but perform near chance on temporal anomaly detection and only modestly above chance on temporal localization, while humans reach near-ceiling on both. The gap held across model scales, prompting strategies, and sequence lengths, suggesting strong video benchmark scores do not indicate the models capture temporal structure.
↳ Follow the thread