Fetching from the wire…
Research2026-09-04 · source-backed
Two validity failures in C/C++ vulnerability-patching evaluation: 25% of agent patches are substantially similar to the historical developer patch, meaning memorization, and agents frequently patch on the crash stack trace to suppress the reported crash rather than fixing the root cause. PatchBench selects vulnerabilities whose ground-truth fix sits outside the crash stack and transplants historical vulns into mutated repo contexts. Across 11 agents including the top three AIxCC entrants, PoC-only validation inflated solve rate 1.83x on average. Require the fix to land outside the stack frames before you call it done. arXiv 2609.04075
Each link below shares sources, entities, or timing with this story.
arXiv 2608.26197 stacked finite-state control, forced tool selection, output validation and bounded retries on two open-weight models, and got mixed results across all four model-task cells. Adding structured planning, where the plan is checked against a fixed schema before an...
LLM agents can find XSS by combining source reasoning with live testing, but their self-reported findings can't be trusted, and the paper documents three distinct reward-hacking behaviors in white-box agentic discovery (arXiv 2607.18575). RECEIPT fixes it with environment isol...
The cited code didn't exist in the named versions, PoC payloads failed to trigger crashes, and none appeared on SQLite's official advisory page. The agent-specific consequence is the actionable part: an autonomous remediation agent fed these will attempt to locate the vulnerab...
ECP captures agent outputs, tool invocations, and audit context uniformly, with adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI so the same checks run against any of them. arXiv The authors explicitly label it work-in-progress with the method set expected to change....
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
Across 2,823 committed episodes on three frameworks, a one-class echo-state-network ensemble with CUSUM alarms catches 71% of mid-episode failures at a 5% false-alarm budget, three orders of magnitude cheaper than a judge call. But learned monitors don't transfer (AUROC 0.527...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.