Fetching from the wire…
Security2026-09-15 · source-backed
SkillAtlas converts private security report bundles into reviewed, redacted, searchable public cases: 3,014 cases, 6,589 traces, 151,131 steps, 233 affected skills, 8 risk categories (arXiv 2609.13353). The defender-relevant number is that 42.5% of successful cases first fail, which means a single sandbox run or a stable signature misses them entirely. Trajectory-grounded labels raise pre-execution guard accuracy to 0.770.
Each link below shares sources, entities, or timing with this story.
The attack needs no instruction, trigger, or retriever optimization, just plainly worded false assertions generated in one pass against a LongMemEval corpus. A four-stage screening pipeline that reaches 0.832 recall on indirect prompt injection rejected none of the poisoned me...
SWE-Touch mines task-critical regions from repair trajectories, builds plausible Counter-Edits that conflict with task completion, and injects them with contextual user messages when the agent reaches that code. Across nine models on SWE-bench Verified, average resolve rate dr...
Give an agent one H100, a target LLM, and two hours of wall clock to deploy and optimize an OpenAI-compatible inference server across four scenarios (arXiv 2607.20468). Across 15 frontier agent configurations, agents reach up to 8.08x and often match default vLLM at 4.05x. A s...
Automatically extracts actionable learnings from execution traces and stores them for retrieval on future similar tasks. Decomposes trajectory value into sub-goal components with contextual retrieval, improving success rates on repeated tasks without retraining. arXiv 2603.106...
13 public sources consolidated into 9,740 skills (7,505 malicious, 2,235 benign) across 11 harmonized attack categories. Learned text detectors score 0.882-0.932 Macro-F1 under random splits but collapse to 0.653-0.665 source-disjoint. arXiv Three off-the-shelf skill scanners...
Reflex-Guard combines jailbreak-aware preprocessing, compact sentence-transformer embeddings and seven binary classifiers trained on 30,568 samples, reporting 95.9% recall end-to-end against 255ms for Llama Guard 2 and 723ms for SafeDecoding, with 100% detection of GCG suffix...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.