OmniAssistBench Scores Gemini-3-Pro at 66.4 out of 100 on Real-Time Video Assistance, Qwen3-Omni-Instruct at 51.2
OmniAssistBench (arXiv 2608.21360, Aug 21) evaluates omni-modal models as real-time video assistants that must continuously perceive an environment and guide a user toward a goal, a setting static offline datasets cannot handle because the model's response changes the user's next action. The team constrains diverging interaction paths by giving models predefined priors from the source video and requiring them to guide users along identical routes, reverse-engineering internet videos into multi-turn clips over more than 1,000 expert person-hours. Gemini-3-Pro reached 66.4 of 100 versus 51.2 for Qwen3-Omni-Instruct, with both failing on hand-gesture visual prompts, multi-turn history retention, and delaying a response until the target event occurs.
↳ Follow the thread