Fetching from the wire…
Public story · 2026-07-25 · high
GPT-5.5's Security score swung from 77.91 to 64.39 depending only on which agent harness ran it, per the leaderboard.
Why now: The swing surfaced in the same batch of coverage that had Anthropic reporting Opus 5's Frontier-Bench and ARC-AGI 3 numbers without naming a harness.
Tencent's WorkBuddy Bench scored GPT-5.5 at 77.91 on its Security subset under one agent harness, and 64.39 under another, per the leaderboard. That's a 13.5-point swing on the same model weights, enough to flip which model looks best to anyone comparing scores.
That gap showed up the same day Anthropic reported Opus 5 more than doubling Opus 4.8's Frontier-Bench score. Anthropic also cited 30.2% on ARC-AGI 3, versus a prior field best of 7.8%, without naming which harness ran either test.
The underlying paper, from Tencent Youtu Lab, Keen Security Lab and Yunding Security Lab, builds 260 tasks from real commits, PRs and business scenarios. Each one gets rewritten as a colloquial role-played request, so the original issue thread can't be found by search.
The suite also refuses one cross-subset average, since each subset uses a different scoring instrument. Token cost varies just as much as accuracy. MiniMax-M3 averaged roughly 11.1 million input tokens per Security run, while GPT-5.5 ran the smallest output budget in the field.
A model that scores three points higher while burning four times the tokens isn't the better pick for a production loop. Rank alone won't tell you that. The smart move: benchmark your own harness on your own tasks before you pick a model.
A related paper found that strong scores on fixed benchmarks don't survive once a user's request shifts mid-conversation, per Tack, Laban and Neville.
Each link below shares sources, entities, or timing with this story.
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
Weeks after launch, Z.ai's open-weight GLM-5.2 now accounts for roughly 75% of all Z.ai model traffic on OpenRouter, with at least one provider serving it past 125 tokens per second (GIGAZINE, citing OpenRouter). The numbers behind the surge: an Artificial Analysis Intelligenc...
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.