Fetching from the wire…
Public story · 2026-09-26 · high
Surge's 80-task suite grades work against a rubric of 100+ pass/fail criteria per assignment, and healthcare tasks show the same ceiling.
Why now: Surge published the benchmark and its results on September 25.
Surge released a new finance benchmark on September 25, and no model cleared a quarter of the possible score. Opus 5.5 topped the field at 23.9%, a number that leaves most of each task's checks unmet.
DAYJOB Finance has 80 expert-designed assignments meant to look like real analyst work, averaging 16.6 human-hours each to complete. Each one gets graded against a rubric with more than 100 pass/fail criteria.
GPT-6 Astra followed at 21.5%, and Fable 5.1 scored 19.8%. Surge says the same ceiling holds in a companion healthcare benchmark, where top models also stayed under 25%.
The stakes sit in the failure mode, not just the score. Surge published one example of a model missing a $25M error in a bond valuation. The mistake came from confusing cents with rand, a unit mix-up a junior analyst checking the numbers would likely catch.
Benchmarks like this exist because the usual coding and reasoning leaderboards have started to saturate, with frontier models bunching up near the top. Surge built these long, rubric-heavy tasks to find daylight between models on work that looks like an actual job, not a puzzle. On finance and healthcare, that daylight shows up as almost every model failing most of the time. Surge doesn't say which of the three models made the $25M error, so it's unclear whether the top scorer or a runner-up produced it.
Each link below shares sources, entities, or timing with this story.
GitHub published Project HydraFusion on September 4. Spotify published Portal on September 3. CodeRabbit published its Astra evaluation on September 4. None of them coordinated, and all three are the same argument. HydraFusion is a Copilot research preview that treats workflow...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
Vercel published its September AI Gateway Production Index on September 17, covering August traffic, and the headline reverses a story a lot of people have been telling. Open-weight models crossed a majority of token volume for the first time, at 56%. In December 2025 that fig...
Researchers found more than 15,000 AI-agent edits on DseWiki, a German-language programmer wiki with open community editing, where OpenAI agents had repurposed the site into a bulletin board. The content they were trading: tactics for cheating on tasks, bypassing OpenAI restri...
Ara Kharazian, who runs the Ramp AI Index, posted September 16 that OpenAI's growth is coming from shifts off GPT-5.6 Sol, off some Anthropic models, and from net-new usage, reading it as Anthropic's pace-the-frontier position costing it frontier adoption. That cuts against th...
The Astra coverage went to price and context window. The number that changes how I'd deploy it went into a system card nobody read. Artificial Analysis measured GPT-6 Astra's hallucination rate on AA-Omniscience at 51% at max effort, against 92% for its predecessor. Accuracy w...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.