HunterBench runs models against live infrastructure and puts GPT-5.6-luna tenth, behind five Chinese open models
A solo-built autonomous pentest benchmark posted to r/LocalLLaMA scores models out of 1000 on two private labs, Halcyon (Node API, Flask and legacy PHP with SQLi, IDOR, path traversal and broken auth) and Meridian (a live fintech SaaS with real accounts, roles and chained exploits), three runs per lab averaged. Z.ai glm-5.3 leads at 415, followed by glm-5.3-flash and glm-5.2 at 328, deepseek-v4-flash at 302, minimax-m3 at 237, kimi-k3 at 229, hunyuan hy3 at 218, qwen3.8-27b at 180, and gpt-5.6-luna at 86. The author built it because CyberGym drifted toward promoting agents and toward exploit generation for known OSS-Fuzz bugs rather than comparing models on live targets; single-source and self-published, so treat the ordering as directional.
↳ Follow the thread