Sources
Latent Interface Training blocks robot policies from cheating on task-irrelevant visual cues
arXiv 2609.12641, submitted 2026-09-11 to cs.RO, targets vision-action shortcuts: robot foundation models exploit visual cues that happen to correlate with demonstrated actions in training, then degrade under distribution shift. LIT is a framework-agnostic two-stage recipe. Stage 1 trains the action expert with no images at all, conditioning on language, robot state, and each chunk's terminal SE(3) end-effector pose. Stage 2 adds a latent interface supervised to reconstruct that terminal pose, and makes it the only path by which visual information reaches the action expert.
↳ Follow the thread