Fetching from the wire…
Public story · 2026-08-10 · high
Built from 243 videos averaging 88.8 minutes, the paper also found human recall drops 17 points without rewatching.
Why now: The paper posted to arXiv in August 2026, arguing existing video benchmarks let weak models pass by testing short clips and multiple-choice answers instead of sustained streams.
StreamArena forces video AI models to watch 88 minutes of footage and answer questions as it plays, per a paper from Xiaohongshu, HKU, CUHK, and HKUST.
The stakes are direct. Any product pitched at live video monitoring needs a model that keeps up in real time, and the best one tested here doesn't come close. StreamMind, the top performer, answers real-time questions correctly 44.5% of the time. Human viewers hit 91.8%.
The authors argue most video benchmarks let weak models pass by testing short clips with multiple-choice answers. StreamArena's 243 videos and 3,646 open-ended questions are built to close that loophole.
The numbers get worse by task. StreamMind manages 34.6% on retrospection, recalling what happened earlier in the stream, and 9.5% on proactive interaction, speaking up about something without being asked. Other baseline models score 8.0% to 28.1% on real-time perception, well below StreamMind's 44.5%.
One score breaks the pattern. StreamMind hits 56.1% on tool-use tasks, well past a rival system called ThinkStream's 1.8%. It also cuts the time between a query and an answer by 66.2%.
The paper's control group complicates the easy read. Historical recall for human viewers drops from 80.7%, when they're allowed to rewatch footage, to 63.4% once rewatching is off the table. Streaming comprehension is hard for people too, just less hard.
A 9.5% proactive-interaction score means these models can watch a live feed but can't be trusted to speak up on their own. That rules out real-time monitoring as a near-term use case, no matter how good the retrieval numbers look. Watch whether the next round of video-agent papers reports proactive interaction at all, or quietly drops it.
Each link below shares sources, entities, or timing with this story.
SkillDetonate built by HKUST / Shared topic
Linked by a graph relationship (SkillDetonate built by HKUST); overlapping topics (against, agent).
OpenCLI supports Xiaohongshu / Shared entity: Xiaohongshu / Earlier coverage
Linked by a graph relationship (OpenCLI supports Xiaohongshu); both cover Xiaohongshu; earlier Xiaohongshu coverage from 2026-08-09.
Same source domain / Shared topic
Reported by the same outlet (huggingface.co); overlapping topics (agent, benchmark, minut, model).
Reported by the same outlet (huggingface.co); overlapping topics (benchmark, best, human, model).
Shared topic / Tension / Downstream implication
Overlapping topics (against, agent, choice); pushes against this story (against); traces where this leads (downstream).
Same source domain / Shared topic / Tension
Reported by the same outlet (huggingface.co); overlapping topics (agent, control); pushes against this story (versus).
Shared topic / Tension
Overlapping topics (against, agent, choice, model); pushes against this story (against).
Overlapping topics (against, agent, best, human); pushes against this story (against).