Fetching from the wire…
Public story · 2026-07-25 · high
GPT-5.5's Security score swung from 77.91 to 64.39 depending only on which agent harness ran it, per the leaderboard.
Why now: The swing surfaced in the same batch of coverage that had Anthropic reporting Opus 5's Frontier-Bench and ARC-AGI 3 numbers without naming a harness.
Tencent's WorkBuddy Bench scored GPT-5.5 at 77.91 on its Security subset under one agent harness, and 64.39 under another, per the leaderboard. That's a 13.5-point swing on the same model weights, enough to flip which model looks best to anyone comparing scores.
That gap showed up the same day Anthropic reported Opus 5 more than doubling Opus 4.8's Frontier-Bench score. Anthropic also cited 30.2% on ARC-AGI 3, versus a prior field best of 7.8%, without naming which harness ran either test.
The underlying paper, from Tencent Youtu Lab, Keen Security Lab and Yunding Security Lab, builds 260 tasks from real commits, PRs and business scenarios. Each one gets rewritten as a colloquial role-played request, so the original issue thread can't be found by search.
The suite also refuses one cross-subset average, since each subset uses a different scoring instrument. Token cost varies just as much as accuracy. MiniMax-M3 averaged roughly 11.1 million input tokens per Security run, while GPT-5.5 ran the smallest output budget in the field.
A model that scores three points higher while burning four times the tokens isn't the better pick for a production loop. Rank alone won't tell you that. The smart move: benchmark your own harness on your own tasks before you pick a model.
A related paper found that strong scores on fixed benchmarks don't survive once a user's request shifts mid-conversation, per Tack, Laban and Neville.
Each link below shares sources, entities, or timing with this story.
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Opus); both cover Bench, Claude Code, Claude Opus, Fable; overlapping topics (benchmark, claude, code, model, point).
output uses Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (output uses Claude Code); both cover Bench, Benchmark, Claude Opus, Fable; overlapping topics (claude, leaderboard, model, number).
Opus built by Anthropic / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Opus built by Anthropic); both cover Bench, Claude Code, Claude Opus, Frontier; overlapping topics (claude, code, model).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Opus); both cover Bench, Claude Opus, Frontier, GPT; overlapping topics (model, task, token).
HuggingFace released Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (HuggingFace released Claude Code); both cover AGI, ARC, Claude Opus, Code; reported by the same outlet (arxiv.org).
Opus built by Anthropic / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Opus built by Anthropic); both cover Bench, Claude Opus, GPT, Opus; reported by the same outlet (officechai.com).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Opus); both cover Bench, Claude Opus, GPT, Opus; overlapping topics (benchmark, code, number, same).
Linked by a graph relationship (Claude Code uses Opus); both cover AGI, ARC, Benchmark, Claude Opus; overlapping topics (benchmark, claude, code, model).