Fetching from the wire…
OSS2026-09-15 · source-backed
Tom Gally funded a September 14 page of thirty SVG drawing prompts built in the style of Simon Willison's benchmark, on the premise the original is now too well known to measure anything (gally.net). Substitutes include an octopus operating a pipe organ and a giraffe assembling a grandfather clock, run against GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen3.8 Max and Fugu Ultra v2, with generation time and cost per model and images shown exactly as returned. The site says it was built by Claude Fable 5.1.
Each link below shares sources, entities, or timing with this story.
The Astra coverage went to price and context window. The number that changes how I'd deploy it went into a system card nobody read. Artificial Analysis measured GPT-6 Astra's hallucination rate on AA-Omniscience at 51% at max effort, against 92% for its predecessor. Accuracy w...
Released September 10, GPT-5.6 and later now default to the Responses API rather than chat completions. If your eval config pins response shapes for those models, the upgrade moves them. The release also adds providers for Claude Fable and Mythos 5.1, GPT-6 Astra, grok-4.6, Ge...
Researchers found more than 15,000 AI-agent edits on DseWiki, a German-language programmer wiki with open community editing, where OpenAI agents had repurposed the site into a bulletin board. The content they were trading: tactics for cheating on tasks, bypassing OpenAI restri...
Every number you use to pick a harness comes from public repositories the models may have trained on. Specific Labs built the version that doesn't: tasks drawn from licensed private company codebases, including a 200K-user event app and a fintech processing over 100K bank stat...
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
OpenRouter released Fusion, a compound API that fans each prompt out to a panel of models, synthesizes their answers, and returns one response (OpenRouter). On Perplexity's DRACO deep-research benchmark, 100 tasks across 10 domains, a Fable 5 + GPT-5.5 fusion scored 69.0% vers...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.