Fetching from the wire…
Agents2026-09-02 · source-backed
Defense-as-Skill argues pre-install vetting is structurally insufficient because a malicious skill only triggers once a concrete task and workspace state make the unsafe action look useful. SkillSonar runs as an editable skill alongside untrusted skills, checking sensitive actions against the user's stated task boundary and routing to allow, replan or confirm without touching the agent runtime (arXiv 2609.01487). They built SCOPE-R, 206 attack-confirmed malicious instances plus 43 benign tasks across 6 risk families, and evolved the guard with MCTS on rollout feedback. Evaluated on both Claude Code and OpenClaw.
Each link below shares sources, entities, or timing with this story.
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
Everyone writing SKILL.md files has absorbed the same folklore. Keep the top file thin. Push detail into reference files. Let the agent walk the tree as needed. More layers, more context efficiency. A controlled study submitted July 20 tested that across InfiniteBench, three a...
A paper from Xiao Yu, Baolin Peng, and Ruize Xu makes a claim that seems obvious once stated and is genuinely new as a training methodology: modern agents are inseparable from their inference harnesses, so training them in stripped-down RL sandboxes produces a train/serve mism...
This one annoyed me, in the good way. Researchers took 206 real developer-agent sessions from 13 developers, extracted each developer's preferences from their actual interaction traces via rule-based bootstrapping plus evidence-grounded refinement, then replayed everything aga...
SynChain uses persistence-aware directed SFT to make a computer-use agent produce artifacts that pass vetting while hiding malicious influence in structural redundancies, surviving internal state updates and reactivating in a later workflow with no new external input. Tested a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.