Fetching from the wire…
Public story · 2026-07-20 · high
The model, called Xiaomi-Robotics-1, scored 74.5% on RoboCasa and a record 57.6% on RoboCasa365, beating GR00T N1.6, Pi-0.5, and three other rivals.
Why now: The paper's Hacker News post climbed to 342 points in nine hours, unusually fast for a robotics paper on a site that leans toward software.
Xiaomi open-sourced Xiaomi-Robotics-1, a robot model that beats five rival systems on a standard manipulation benchmark, per the arXiv paper.
The five it beat count among the strongest systems in robot learning. An open, free model that outperforms all of them resets the baseline every other lab has to clear.
The model scored a 74.5% average success rate on RoboCasa, ahead of RLDX-1, Cosmos Policy, GR00T N1.6, Pi-0.5, and Pi-0-FAST. On the harder RoboCasa365 suite, it hit 57.6%, a new state of the art over the prior best of 46.6%.
Xiaomi-Robotics-1 doesn't use a new architecture. Its edge comes from the training set: 100,000 hours of real manipulation trajectories, collected from physical robots instead of simulation.
That kind of real-world data is expensive and slow to collect at scale. That's exactly why it works as a barrier other labs can't easily clear.
Robot learning's next round of benchmarks will likely go to whoever has the most real-world manipulation hours, not whoever designs the cleverest model. Worth watching whether Xiaomi keeps releasing weights at this scale, or whether this is the high point before the real advantage gets held back.
The arXiv paper's Hacker News post drew 342 points in nine hours, fast for a robotics paper on a site that usually favors software.
Each link below shares sources, entities, or timing with this story.
The stealth model is Xiaomi's MiMo V2 Flash: 309B MoE with 15B active parameters, 256K context, hybrid-thinking toggle. 73.4% SWE-Bench Verified — top open-source globally — approaching GPT-5-High at roughly 3.5% of the cost. A successor model was teased in the same OpenClaw P...
arXiv 2607.25857 reformulates all content moderation as a single binary yes/no QA task, which lets heterogeneous safety datasets with incompatible taxonomies consolidate under one training framework instead of requiring a taxonomy merge. It sets a new SOTA on multimodal safety...
This one hit my inbox and I had to read it twice. Anthropic announced a partnership with SpaceXAI for the entire Colossus 1 data center in Memphis. 220,000 NVIDIA GPUs. 300+ megawatts. That's the largest single compute acquisition by any AI lab. Full stop. But the part that ma...
The top trending HuggingFace paper (274 upvotes) introduces dots.tts, a 2B continuous autoregressive text-to-speech model hitting best average Seed-TTS-Eval (WER 0.94%/1.30% zh/en) with strong cloning and emotional range. CFG-aware MeanFlow distillation gives 85ms first-packet...
Announced alongside FLUX 3, it bolts a lightweight action decoder onto intermediate FLUX 3 features to make a video-action model, beating prior VLA models even with the backbone frozen. Deployed at Audi for kitting, component insertion and flexible-material manipulation, with...
CNBC covered the July 16 reveal, timed to Jensen Huang's Japan visit, of a world model built to perceive and navigate physical environments in real time. Nvidia is forming a physical-AI coalition Fujitsu, Hitachi, and Kawasaki intend to join. This is the clearest signal yet th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.