Figure's Helix 2.5 raises zero-shot household task success from 9% to 56% across 30 homes it never trained on
Figure announced Helix 2.5 on 2026-09-17: a humanoid policy pretrained on its Index dataset of human behavior and then evaluated in 30 Bay Area homes with no data collected in any of them, no fine-tuning and no adaptation to the objects manipulated. In a controlled comparison against an otherwise identical policy trained from scratch, Index pretraining lifted zero-shot success from 9% to 56%, and one foundation model produced three long-horizon whole-body behaviors: tidying living rooms, folding towels and making beds. Figure says Helix 2.5 needed 50% less task-specific data than Helix 02 while covering three times as many homes, and that Index currently ingests roughly 35 minutes of human behavior per second against a $3.5 billion compute commitment.
↳ Follow the thread