Fetching from the wire…
Research2026-09-15 · source-backed
Agents get only a CWE description and terminal access, and have to find the implementing files across 500 real vulnerabilities from 290 repositories, six package ecosystems and 147 CWE categories (arXiv 2609.15939). Across 27 language models and four static-analysis tools on a standardized interface, the strongest system reached File F1 of 0.229, and 38.4% of tasks received no correct localization from any evaluated system. The authors also found systems that identify vulnerabilities well still report unsupported locations on already-patched repos, which separates localization from detection as a distinct capability. Anyone selling agentic vulnerability discovery should be asked about this number.
Each link below shares sources, entities, or timing with this story.
Researchers loaded five systems with a revoked policy and its replacement, then measured retrieval and downstream action across nine policy scenarios, nine models and six defense conditions. Wherever the revocation label was visible to the retrieval layer, the revoked fact cam...
arXiv 2608.24358 switched models mid-run on long coding tasks using cheap/expensive pairs from the Claude and GPT families. Full-trajectory escalation from weak to strong recovers under half the gap while costing a substantial premium, which the authors call the handoff tax. D...
A June 26 paper (arXiv:2606.26294) describes a self-improving architecture where the agent and the evaluator that scores it evolve together, specifically to avoid the stagnation of optimizing against a fixed, gameable reward. (arXiv) Anyone building a self-improving harness ha...
An agent proposes changes to a training pipeline, runs it, and keeps edits improving a verifiable in-loop metric. Looks like reliable progress. The authors name algorithmic mode collapse: surface edit diversity stays stable while semantic and mechanism-level diversity collapse...
2,910 programmatically verified tasks built from an ontology of 97 canonical UI components. Holding the harness fixed and changing only observation and action space, GPT-5 mini scores 83.1% with accessibility-tree observations and 48.9% with coordinate-only pixel control. Acro...
Material Discovery Bench measures LLM progress on discovering thermally conductive dielectrics for 3D chip packaging, built on the AI Security Institute's open-source Inspect framework, with models given web search, Python/bash sandboxes, and property-computation ML tools (Dis...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.