Sources
MOSS-VL Ships an 11.3B Open-Weight VLM That Perceives While It Speaks — 66.0 vs 37.5 on OmniMMI Proactive Alerting and 5.1x Faster TTFT Than Qwen3-VL-8B
OpenMOSS (Xipeng Qiu's group, 32 authors) released MOSS-VL on 2026-08-15 — arXiv 2608.15045 — an 11.3B real-time vision-language model built on gated cross-attention so it can ingest incoming video frames during generation, with visual tokens kept outside the decoded sequence. It scores 66.0 on OmniMMI Proactive Alerting against a 37.5 baseline and hits time-to-first-token 5.1x faster than Qwen3-VL-8B. All five checkpoints, the staged training curriculum and inference code are released; at 430 upvotes it is by a wide margin the most-upvoted paper on HuggingFace Daily Papers today.
↳ Follow the thread