Fetching from the wire…
Security2026-09-15 · source-backed
A red agent generates context-compatible injections across nine threat types while recording the exact modification, and a blue agent analyzes complete skill packages and proposes patches, so a verifier scores at injection level rather than skill level (arXiv 2609.14079). With its best backend it's the only scanner in the comparison reaching 100% injection detection. Applied to popular published skills it found latent vulnerabilities in over 17% of those examined, and running those skills triggered real incidents.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.24306 tests each agent invocation locally for faithfulness against its own inputs, classifying errors as hallucination, uncited input reliance, uncited output or insufficient citation. Applied to three top-ranked open-source deep research systems, nearly every agent...
An attack-surface survey of MCP, Skills, and tool calling (arXiv 2608.17275) reports the share of deployed MCP tools that modify external state has climbed from 27% to 65%. That reframes the threat model entirely: agents act now, they don't read. Applied to blockchain executio...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
Thinkingbox is an MCP-compatible sandbox with isolated sessions, full execution traces, and outcome evaluation against terminal backend state, carrying 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank internal IT and consulting support (arXi...
It's a benchmark of 56 contract-defined backend tasks, judged only through black-box HTTP tests against an OpenAPI contract, so there's nowhere to hide (arXiv). GPT-5.5, the best model, succeeds on 55.4% under the base oracle and drops to 28.6% under the final hardened oracle....
Skill self-evolution methods revise skill text from execution feedback, but each oracle evaluation needs a full agent rollout, which confines search to patching whatever just failed (arXiv 2609.15396). SkillLift treats ranking as a smoother supervision target than absolute sco...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.