Fetching from the wire…
Public story · 2026-09-06 · high
Both models flop on precision work, managing 2 of 20 tries each on a tighter task.
Why now: RoboCurve posted the results September 4.
OpenAI ran its Astra model on real bimanual robot arms and won big on one task, badly on another. In results published September 4, RoboCurve tested Astra and Fable 5.1 head to head, 20 trials each, on I2RT YAM arms through the Inspect Robots 0.58.0 harness.
On the coarse task, putting a block into a bowl, Astra hit 19 of 20 tries. Fable 5.1 managed 8, and an older Fable 5 managed 1. Astra also did it cheaper and faster: $0.94 a run against Fable 5.1's $2.12, 2.5 minutes against 6.8, and 2,100 output tokens against 12,900.
The second task tells a different story. Fitting a puzzle piece into a groove asks for tighter precision, and both models scored exactly 2 of 20 attempts. Whatever let Astra beat Fable 5.1 on block-into-bowl didn't carry over. The benchmark doesn't say why the gap closes on the harder task, only that it does.
That split matters more than the headline number. A model that's 19/20 on coarse placement and 2/20 on tight-tolerance work isn't ready to run a warehouse pick line or assemble anything with fitted parts. It's ready for tasks that look like dropping a block in a bowl, which is a narrower claim than "OpenAI's robot model wins."
Watch whether the next round of testing adds a task between those two extremes. Right now there's a wide gap between "easy and Astra dominates" and "hard and both models are equally bad," and nothing here tests what's in between.
Each link below shares sources, entities, or timing with this story.
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
The Register put the two side by side on August 8. Read together, frontier safety policy isn't converging on a posture, it's splitting by risk domain, with each lab tightening where its own evals scared it and loosening where false positives cost product usability. That's evid...
This one's been building for days and it crystallized this week. Per The Register, the incident behind the US export-control block on Anthropic's Fable 5 and Mythos 5 wasn't a jailbreak or a guardrail bypass. It was a plain three-word prompt, "fix this code," run against CVE-l...
Anthropic shipped Fable 5 on June 9. Willison spent ~5.5 hours stress-testing it: slow and expensive, but it handled everything he threw at it, including agentic coding. (Simon Willison) The tell that it's a real working model and not a benchmark queen: because it post-dated A...
Simon Willison surfaced Jarred Sumner's writeup of rewriting Bun's core from Zig to Rust this week, and the numbers stopped me cold. PR #30412, merged May 14, added roughly 1 million lines across 2,188 files, reached 99.8% test compatibility on Linux x64, and shrank the binary...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.