Fetching from the wire…
Public story · 2026-08-04 · high
It's trained on IBM's Granite base and beats open-weight rivals 200 times its size.
Why now: The paper is part of the August 4, 2026 research briefing, arriving as teams weigh how many security scans a CI budget can actually afford per build.
Antares-3B outperforms open-weight models 200 times its size at finding software vulnerabilities, per a paper posted to arXiv.
The economics are the point. A full 500-task evaluation sweep finishes in about 15 minutes on a single H100, working out to under 2 seconds and under $0.002 per task. That puts continuous vulnerability scanning inside a CI budget instead of a frontier-API bill.
The model starts from IBM's Granite base. Supervised fine-tuning on cybersecurity reasoning and repository-exploration data comes first, then reinforcement learning from verifiable rewards, run against real vulnerable repositories. The paper says Antares approaches GPT-5.5's accuracy. It doesn't say Antares beats it, or how much of the gap is left.
If a 3B-parameter model can get this close to frontier accuracy at $0.002 a task, the constraint on security scanning stops being model capability and starts being how often you're willing to run it. Teams that gated vulnerability scans to nightly or pre-release runs because of API cost lose that excuse. Worth watching whether Antares' recipe, RL over verifiable rewards on real vulnerable code, shows up in other narrow, checkable domains where a small model can beat a much larger one on price.
Each link below shares sources, entities, or timing with this story.
GPT competes with Claude / Shared entity: GPT / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Copilot uses GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT; overlapping topics (data, model).
Critique uses GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Critique uses GPT); both cover GPT; overlapping topics (model, over).
GPT competes with Claude / Shared entity: GPT / Same source domain / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Copilot uses GPT / Shared entity: GPT / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT; earlier GPT coverage from 2026-07-20.