Microsoft Research's MindTopo finds VLMs can see topology but lose it the moment they have to act
MindTopo is a new benchmark testing topological reasoning across five categories — continuity, separation, order, enclosure and knots — in both static scene analysis and interactive planning in simulated environments. Across a broad set of proprietary and open-weight models, performance was consistently stronger on static reasoning than on interactive planning, and both stayed well below human performance. The failure modes differ in a way that matters for agent builders: static errors are perception failures, while planning errors appear after the scene is correctly understood, as models lose track of relationships across multiple actions. Image and video generation tools helped little, often altering topology or violating task constraints.
↳ Follow the thread