Fetching from the wire…
Public story · 2026-08-31 · high
HunterBench scores models on real exploit chains against a fintech SaaS and a legacy PHP app, and five Chinese open models beat OpenAI's entry.
Why now: The site went up and started circulating in the August 31 roundup of new evals.
HunterBench scores language models on live infrastructure instead of static challenge sets. One target is a Node, Flask, and legacy PHP app seeded with SQL injection, IDOR, path traversal, and broken auth. The other is a fintech SaaS with real accounts, roles, and exploit chains that require moving between them. Each model runs three times per lab, out of 1000 points, and the scores get averaged.
Glm-5.3 tops the board at 415. Glm-5.3-flash and glm-5.2 tie at 328, deepseek-v4-flash sits at 302, minimax-m3 at 237, kimi-k3 at 229, hunyuan hy3 at 218, and qwen3.8-27b at 180. GPT-5.6-luna comes in last at 86, trailing every open Chinese model on the list.
The author built this after watching CyberGym drift toward generating exploits for known bugs already logged in OSS-Fuzz, a task that rewards pattern matching against a known answer more than it rewards finding a new hole in a running system. HunterBench instead asks a model to work a target the way a human tester would: enumerate, chain, escalate.
That design choice is also its limit. Two private labs built by one person is a small, fixed sample, and nobody outside the author can check whether the challenges are balanced across the model list or whether one lab happens to reward a particular attack style. Read the ranking as directional, not as a claim that Chinese open models are categorically better at offensive security than GPT-5.6-luna. If you're picking a model for red-team tooling, that gap is large enough to justify running your own targets before trusting a score built on someone else's two boxes.
Each link below shares sources, entities, or timing with this story.
Claude Code benchmarked against GPT / Shared entity: GPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; overlapping topics (against, model).
GPT competes with Claude / Shared entity: GPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; overlapping topics (against, model, target).
Flask built by Armin Ronacher / Shared entities / Earlier coverage
Linked by a graph relationship (Flask built by Armin Ronacher); both cover Flask, GPT; earlier Flask coverage from 2026-06-17.
Claude Code benchmarked against GPT / Shared entities / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover Chinese, GPT; earlier Chinese coverage from 2026-07-08.
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover Chinese, GPT; earlier Chinese coverage from 2026-07-01.
Hermes Agent uses GPT / Shared entity: Chinese / Shared topic / Earlier coverage
Linked by a graph relationship (Hermes Agent uses GPT); both cover Chinese; overlapping topics (account, target).
Claude Code benchmarked against GPT / Shared entities / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover SaaS, Self; earlier SaaS coverage from 2026-07-19.
GPT competes with Claude / Shared entity: GPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; overlapping topics (against, model).