HiFi-UMI Removes the Real-Robot Anchor: Policies Post-Trained Only on Handheld Capture Match Teleoperation Within 3 Points
HiFi-UMI (arXiv 2607.25895, 2026-07-28) is the top trending paper on HuggingFace today at 82 upvotes. Rather than shrinking the real-robot data fraction, it raises robot-free UMI capture fidelity — head-mounted offline stereo-inertial SLAM, native rather than reconstructed inter-gripper pose, a shared microsecond GPIO trigger, and two ~200-degree wide-angle cameras per hand — reaching 3 mm workspace-local end-effector accuracy with no external tracking. Policies post-trained solely on HiFi-UMI demonstrations deploy directly to a real robot within -2.5, +3.1, and -0.6 percentage points of in-domain teleoperation across StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA, hitting 85% on precision insertion. Pre-training on 4,000 hours lowers action error on ten unseen tasks by 41%; the team open-sources HiFi-UMI-2K, 2,000 hours of synchronized demonstrations.
↳ Follow the thread