Meta's HumanCLAW Isolates VLM 'Action Intelligence' From Motor Control — Best of Nine Models Scores 16.8%
HumanCLAW (arXiv 2607.27180, 2026-07-29, Meta Research with Ziwei Liu, Ranjay Krishna, Manling Li and others) decouples decision-making from low-level execution: a harnessed off-the-shelf VLM issues atomic skill commands that a controller translates into sub-second chunks of physically simulated full-body motion, so balance and motor failures are factored out. On HumanCLAW-Bench — 1,218 long-horizon egocentric find-navigate-interact episodes across 41 indoor scenes — nine state-of-the-art VLMs all fail, with the best reaching 16.8% success. The diagnosis is specific: recognizing the target is not the bottleneck; models lose track of their own body and cannot tell whether they have arrived or collided.
↳ Follow the thread