Om AI Lab Releases VLX-Seek-1.5-10B, Which Converts Visual Regions Into Language-Addressable Tokens Instead of Emitting Bounding Boxes
Om AI Lab published VLX-Seek-1.5-10B on Hugging Face under Apache 2.0 — a 10B vision-language model for fine-grained perception and visual grounding in embodied settings like drones, robots, and surveillance. The architectural choice worth noting is "region-reference localization": rather than having the LM generate bounding-box coordinates directly, it converts visual regions into language-addressable tokens, on the theory that this aligns better with what language models are actually good at. It covers open-vocabulary detection, referring expression comprehension, multi-object grounding, and counting with region-level evidence, with published evaluations on drone-view perception, embodied spatial reasoning, and object hallucination. It sat at 349 downloads and zero Reddit comments despite 65 upvotes — near-invisible relative to the day's Meta news.
↳ Follow the thread