Sources
GST-Bench: Top VLMs Score 42.68 vs 79.08 for Humans on Building a Consistent Spatial Map From Video
ByteDance Seed's benchmark (arXiv 2608.05747, Aug 6) covers 6,790 minutes of synthetic video with human-verified questions and isolates a specific failure: models handle local spatial relations competently but cannot consolidate long-horizon observations into a globally consistent scene representation. The ~36-point gap to human performance is the headline, and the authors also release GST-Train to attack it. This matters for anyone betting on computer-use or embodied agents that navigate over long episodes — the bottleneck being measured is persistent spatial memory, not perception.
↳ Follow the thread