Sources
EXIMO Finetunes Vision-Language-Action Robot Policies in Three Stages Using a VLM to Guide Exploration Instead of Teleoperation Data
arXiv 2608.19891, submitted August 20, targets the open problem of teaching a billion-parameter VLA policy a new task without collecting hundreds of hours of teleoperation. EXIMO runs explore, imitate, and optimize stages, using a vision-language model to direct exploration rather than relying on RL's sample inefficiency over long horizons or on new human demonstrations. The framing matters beyond robotics: it is the same pattern as using a stronger model to generate the trajectories that finetune a cheaper specialized policy, applied where data collection is physically expensive.
↳ Follow the thread