Fetching from the wire…
Public story · 2026-09-18 · high
Trained on human behavior data alone, the policy never touched any of the 30 homes it was tested in beforehand.
Why now: Figure published the results on September 17, the same announcement citing a $3.5 billion compute commitment behind the Index dataset.
Figure's Helix 2.5 reached 56% success on household tasks in 30 homes it never trained on, per Figure's September 17 announcement. An identical policy trained from scratch on the same tasks managed just 9%. The 47-point gap comes from pretraining alone. The policy never collected data in the 30 homes and received no fine-tuning once it arrived.
Helix 2.5 runs on Figure's Index dataset, built from human behavior footage ingested at about 35 minutes per second. The same policy tidied living rooms, folded towels, and made beds without any per-home adaptation to the objects it manipulated. Figure says Helix 2.5 also needed 50% less task-specific data than its predecessor, Helix 02, while covering three times as many homes.
Figure doesn't say how it selected the 30 homes or the full list of tasks it tested. The company is also the only party that has evaluated Helix 2.5 so far, and Index's build involves a $3.5 billion compute commitment.
Each link below shares sources, entities, or timing with this story.
Starting at $3.5 billion with plans to exceed $6 billion, initial deployment targeted for the second half of 2027 in Barstow, Texas, earmarked for training Helix, Figure's humanoid control model. Both sides said they'll explore using Figure's robots inside Nscale's supply chai...
TechCrunch: Greenoaks led on July 30, five months after Index led a $100M Series A. Founded by Stanford PhD Joon Sung Park, author of the "Smallville" generative-agents work. Revenue up 5x since a February 2026 launch, headcount past 50, tens of millions of simulations run for...
Company A did it. Then B. Then the survey data caught up. Something structural is breaking. Three independent moves in early July all point the same direction. OpenAI started token-based credit metering for ChatGPT Workspace agents on July 6. Microsoft made an E5 license a pre...
GitHub quietly updated its Copilot pricing multiplier table, and the numbers are jarring. Starting June 1, 2026, every Copilot interaction gets priced in "AI Credits" with per-model multipliers: Claude Opus at 27x the base rate, Claude Sonnet at 9x, and base completions at 1x....
For six years Shopify was the loudest big-company case for React Native. On September 10 it reversed that decision in public. In a Shopify Engineering post, engineer Mustafa Ali wrote that "LLMs changed one of the core assumptions behind our 2020 decision." The 2020 reasoning...
Intelligence raised a $7.9M seed led by Index with Conviction, A*, and YC. 5.3 million users across 190+ countries voting on pairs of AI-generated designs. The thesis frontier labs are now paying for: subjective quality can't be auto-graded, so human preference at scale become...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.