Fetching from the wire…
Research2026-09-25 · source-backed
arXiv 2609.26550 (CMU, submitted September 22) compares Jev against sixteen generative and reward-model judges under blinded human adjudication. It costs $0.044 per 1,000 judgments at 152ms median and lands within 3 points of the strongest LLM judge on RewardBench-style preference and HaluEval factuality. It trails by 14.5 points on JudgeBench, where the judge has to check a derivation or resist a well-written wrong answer. The frozen cascade, accept confident verdicts and escalate the rest to GPT-6 Astra, is directly copyable for eval pipelines.
Each link below shares sources, entities, or timing with this story.
The Astra coverage went to price and context window. The number that changes how I'd deploy it went into a system card nobody read. Artificial Analysis measured GPT-6 Astra's hallucination rate on AA-Omniscience at 51% at max effort, against 92% for its predecessor. Accuracy w...
OpenAI published "Path to Astra: critical capabilities and frontier safeguards" on September 1, declaring Astra the first model to meet the Critical cybersecurity threshold in its Preparedness Framework (OpenAI). Critical, in their own definition, means the model can find and...
GitHub published Project HydraFusion on September 4. Spotify published Portal on September 3. CodeRabbit published its Astra evaluation on September 4. None of them coordinated, and all three are the same argument. HydraFusion is a Copilot research preview that treats workflow...
Ara Kharazian, who runs the Ramp AI Index, posted September 16 that OpenAI's growth is coming from shifts off GPT-5.6 Sol, off some Anthropic models, and from net-new usage, reading it as Anthropic's pace-the-frontier position costing it frontier adoption. That cuts against th...
Tom Gally funded a September 14 page of thirty SVG drawing prompts built in the style of Simon Willison's benchmark, on the premise the original is now too well known to measure anything (gally.net). Substitutes include an octopus operating a pipe organ and a giraffe assemblin...
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.