Agents
ADEPT pre-trains one dexterity policy on object reposing, then post-trains downstream skills onto two robot hands
arXiv:2608.19182 (submitted 19 Aug 2026) applies the pre-train/post-train split to dexterous manipulation RL: a generic object-reposing policy is learned once, and downstream policies build on it instead of relearning the same skills per task. Transfer is stabilized with behavior-cloning distillation and conservative on-policy updates, plus a joint-space Geometric Fabric to keep the RL policy safe against real hardware. Distilled policies ran on a 23-DOF Kuka-Allegro with two RGB cameras and a 29-DOF Flexiv-Sharpa with five vision-based tactile sensors, reaching human-level speed on long-horizon tasks from hard initial states.
Source
↳ Follow the thread