Fetching from the wire…
Security2026-09-10 · source-backed
MMPIBench pushed a fixed attack set through six visual carriers across 720 runs on six frameworks, five models and four attacker objectives. Visual attacks were attempted in 12.8% of runs but completed in about 1%, with nearly the whole gap closing at the planning step. Extend to audio and it flips: only two of five models ingest audio and only three of six frameworks deliver it, but where the signal arrives, the attack completed in 49% of cells and 75% for one model. Vision got hardened by training. The other channels didn't. If you're adding voice input to an agent, you're adding the unhardened path.
Each link below shares sources, entities, or timing with this story.
A fleet evaluation across 46 endpoints from six vendors found a recognition-enforcement gap: source-format features are linearly decodable from activations and models verbally identify forged authority when asked, but some configurations still emit the conflicting tool call. A...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
arXiv 2607.24174 (July 27) generated adversarial log entries from real attack traces and got multiple state-of-the-art LLMs to classify traces containing clear indicators of compromise as benign. The defensive gift: the natural-language explanations emitted alongside the class...
LongPIBench covers four realistic scenarios (peer review, resume screening, code review, email summary) at context lengths from thousands to tens of thousands of tokens, and reports that even simple heuristic injection attacks reach high success rates and frequently bypass cur...
The study covered GPT-4o, Claude 3.5 Sonnet and Llama-3.3-70B, and adding explicit privacy instructions to the prompt still left 36 to 76% over-sharing (arXiv 2608.24957). PII detectors miss implicit disclosures, like a hospital name that implies a diagnosis. The middleware in...
This is the most directly usable paper of the day and it does something rare: it bolts onto an existing agent without touching it. Ledger is a deterministic runtime wrapper that distills an agent's completed interactions into explicit state. What has been observed, what has be...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.