Research
LLaVA-OneVision-2: Codec-Stream Tokenization for Next-Generation Multimodal Intelligence
The most capable model in the LLaVA-OneVision series introduces codec-stream tokenization that treats compressed video as a continuous bit-cost stream, using motion-residual cues to concentrate token budgets on event-bearing content. Features a unified OneVision-Encoder for images and video with windowed attention for efficient local computation at native resolution. Achieves superior performance across broad multimodal benchmarks. Open framework on GitHub (EvolvingLMMs-Lab).
Source
↳ Follow the thread