Fetching from the wire…
Agents2026-09-07 · source-backed
This inverts speculative decoding. A small open-weight draft model scores a black-box agent's already-generated trajectory in one forward pass, needing no logits, weights, activations or repeated sampling (arXiv 2609.05274). Phase-aware features separating reasoning spans from action spans get calibrated against a verifiable objective into a failure-likelihood score. Wired into a pre-execution veto gate on Qwen3-Coder-480B and Claude 3.5 Sonnet it cut execution error rate 6-8 points and token cost 14-19%, and transferred to out-of-distribution benchmarks without retraining.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
The study covered GPT-4o, Claude 3.5 Sonnet and Llama-3.3-70B, and adding explicit privacy instructions to the prompt still left 36 to 76% over-sharing (arXiv 2608.24957). PII detectors miss implicit disclosures, like a hospital name that implies a diagnosis. The middleware in...
MCP-Atlas has 1,000 human-verified tasks across 36 real MCP servers and 220 tools, now with a 100-tool-call budget instead of a 20-turn limit. Current leaders: Gemini 3.5 Flash at 83.6%, Claude Opus 4.8 at 82.2%. Tool-Decathlon runs 108 long-horizon tasks in isolated container...
DeepClaude hit 470 points on Hacker News. It swaps Claude Code's API backend to DeepSeek V4 Pro while preserving the full agent loop: file editing, bash execution, git tooling, the whole workflow. DeepSeek V4 Pro scores 96.4% on LiveCodeBench at a fraction of Anthropic's prici...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.