RoboTok retrieves matching human manipulation demos from internet video by hand-trajectory similarity rather than appearance
arXiv 2609.03199 (2026-09-02) is the highest-upvoted paper on HuggingFace Daily Papers this week at 116. It attacks the robot-data bottleneck by learning a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames, so a manipulation behavior can be matched across changes in camera viewpoint, scene appearance and actor occlusion while staying compact enough to index internet-scale video continuously. Given a query human manipulation video it retrieves relevant web demonstrations for training dexterous policies, and the authors report better retrieval relevance and higher downstream task success than existing robot-data retrieval methods. The transferable idea for non-robotics builders is the embedding choice: index on the invariant (the motion), not the surface features.
↳ Follow the thread