A Visuo-Tactile World Model Nearly Doubles Dexterous Manipulation Scores, and Ablating Tactile World Modeling Costs Most of the Gain
DexTacWAM (arXiv 2609.24976, 21 Sep 2026) extends World-Action Models, which are largely vision-centric and cannot model contact dynamics, by encoding each fingertip independently, compressing through a finger- and pose-aware tactile compressor, and injecting the tactile latent into a video diffusion world model. On six contact-rich tasks on a 22-DoF bimanual platform it tops every task, averaging 70.6 against 38.0 for the best baseline; removing tactile world modeling while keeping the same tactile features and action expert drops the four-task mean from 74.7 to 26.6. Four hours of tactile-encoder adaptation over a frozen pretrained vision VAE and roughly 100 demonstrations per task extend the pretrained video model to touch, with the compressor keeping 89.4% of pre-fusion contact recall at 2.26x faster training.
Source
↳ Follow the thread