Sources
OpenMOSS Releases MOSS-VL-Realtime — an 11B Vision-Language Model for Continuous Video Streams With 256K Context
OpenMOSS shipped MOSS-VL-Realtime, an 11B vision-language model built to process continuous video streams with a 256K-token context window. Streaming-native VLMs at this size point toward real-time perception agents (monitoring, live assist, robotics) that don't need to chunk video into stills. It's an early but concrete signal that 'always-on' multimodal agents are becoming practical at self-hostable parameter counts.
↳ Follow the thread