Sources
ChronoVision Grounds Video Reasoning in Reconstructed Latent Visual States, Reporting 74.8% In-Domain on a New Image-Ordering Benchmark
Submitted Aug 6 (arXiv 2608.05631), ChronoVision adds a reconstructive visual head that predicts latent representations of transformed visual states plus an ROI attention module driven by semantic span queries, then post-trains with RL using implicit process grounding and a composite reward. On its own Vbvr-VQA — which reformulates video reasoning as a strict image-ordering task — it reports 74.8% in-domain and 71.6% out-of-domain, plus 55.0% on the cross-domain IntPhys2 physical-reasoning benchmark. Single-team results on a self-introduced benchmark, so treat the in-domain number cautiously; the out-of-domain and IntPhys2 transfer are the more credible signals.
↳ Follow the thread