Fetching from the wire…
Top 5 · 2026-09-05 · source-backed
The Astra coverage went to price and context window. The number that changes how I'd deploy it went into a system card nobody read.
Artificial Analysis measured GPT-6 Astra's hallucination rate on AA-Omniscience at 51% at max effort, against 92% for its predecessor. Accuracy went up 4 points in the same measurement rather than being traded away. OpenAI's own system card reports an internal hallucination benchmark falling from 12.2% to 4.2% (Artificial Analysis). An r/OpenAI thread at 275 upvotes is a complaint that this got buried while benchmark rankings took the coverage.
The usual assumption is that cutting hallucination costs you coverage, because a model that refuses to guess answers fewer questions. That didn't happen here. Both moved the right way.
For anything running unattended, refusal-to-guess behavior determines whether a long chain survives. A model that fabricates one plausible intermediate fact at step 8 of a 30-step task produces a confident, coherent, wrong result at step 30, and you find out days later. Cutting the fabrication rate in half is worth more to an overnight agent loop than four points on any index.
The rest of the Astra picture is less flattering. OpenRouter lists it at $10/M input and $50/M output, cache read $1/M, 1,050,000-token context, up to 128,000 completion tokens, and web search at $10 per 1,000 calls. Peak throughput across providers is about 55 tokens per second with a P50 best latency of 2.89s (OpenRouter). Expensive and slow.
Simon Willison ran his pelican-on-a-bicycle SVG test across all five reasoning levels on September 4. Every Astra output from low upward beat the best GPT-5.6 result, and Astra at low cost 9.55 cents. He also notes Astra below max still can't place the pelican's legs on both sides of the frame (Simon Willison).
There's a counter-consideration on the cost math. A 536-upvote r/singularity thread argued Astra's real story is token efficiency, and the correction in the comments is that Astra reasons internally and emits no visible thinking tokens, so part of the apparent efficiency is a change in what gets billed and displayed. Check whether your token counter includes reasoning tokens before you conclude anything about spend this week.
Where Astra earns $1.50 a task: cross-file code review, where CodeRabbit measured a 14-point gap over Opus 5, and long unattended chains where fabrication compounds. Where it doesn't: anything latency-sensitive at 55 tok/s, and anything you're going to read yourself anyway.
One more thing from the safety hub, because it changes the browser-agent risk budget. Astra holds an 8.5% attack success rate across 1,810 curated indirect prompt injection attacks from Gray Swan's IPI Arena, down from 27.0% for GPT-5.6 Sol. Threefold reduction, still one attack in twelve getting through. Keep the sandbox.
Each link below shares sources, entities, or timing with this story.
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
GitHub published Project HydraFusion on September 4. Spotify published Portal on September 3. CodeRabbit published its Astra evaluation on September 4. None of them coordinated, and all three are the same argument. HydraFusion is a Copilot research preview that treats workflow...
Simon Willison pulled the numbers out of an FT report sourced to "people with knowledge of the matter": Anthropic's annualized revenue reached $65bn in July, up from $47bn in May. Six thousand customers spend $100,000 or more a year. The company told investors it expects a pro...
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.