StreamArena Makes Video Agents Watch 88 Minutes Straight and Answer Open-Ended Questions — Best Model Gets 44.5% on Real-Time Perception vs 91.8% for Humans
StreamArena (arXiv 2608.05703, Xiaohongshu with HKU, CUHK, HKUST) argues existing video benchmarks let weak models succeed via short clips and multiple choice, so it uses 243 videos averaging 88.8 minutes with 3,646 open-ended questions across real-time perception, historical retrospection, proactive interaction, and multimodal tool use (27% of candidate questions were cut in QC). Best-in-class StreamMind hits 44.5% real-time perception (baselines 8.0–28.1%), 34.6% retrospection, 56.1% tool use against ThinkStream's 1.8%, and just 9.5% on proactive interaction — while cutting query-to-answer latency 66.2%, from 81.4s to 27.5s, retaining 89.7% pooled accuracy. A quietly humbling control: human historical recall itself falls from 80.7% with rewatching to 63.4% under streaming conditions.
↳ Follow the thread