Fetching from the wire…
Agents2026-09-03 · source-backed
EAL-Bench names a failure with no external attacker: when persistent memory summarizes an evolving authorization state, compression washes away provenance and leaves a stored permission the event history never granted. Across five LLMs as memory writers and two as executors in procurement, cybersecurity and finance, incremental updates created false authority for up to 50.2% of unauthorized requests, and executors then acted on it in 98.6% of trials. Two safeguards help, requiring stored permissions to trace to valid source events and bounded event sourcing for permission changes, but both reject more legitimate actions too.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.26598 targets the failure where an agent recovers from an error within an episode but hits the identical failure in later tasks, because post-episode feedback never revises the persistent harness. Guided by a domain-level Evolution-SOP, it writes episodic memory rec...
arXiv 2607.26998 flips the pentest agent's observation-action loop against it, replacing static honeytokens with a trajectory-adaptive policy that constructs new decoy artifacts conditioned on the agent's interaction history, folding validated ones into a factually consistent...
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
"Remember When It Matters" (arXiv:2607.08716) attacks behavioral state decay, where task-critical instructions get buried or evicted on long runs, using a second memory agent that proactively injects reminders into the action agent's context. It gained +8.3pp pass@1 on Termina...
MemSyco-Bench points out that memory benchmarks test whether memories are correctly stored, retrieved, and updated, never whether the retrieved memory should have influenced the decision at all. Its five tasks check whether agents can reject memory as factual evidence, respect...
Every skill marketplace runs on one assumption: certify each package, and the ecosystem is safe. CompoSkill breaks that assumption by showing composition risk is a path property, not a node property. The attack works black-box. The attacker knows only a role profile. They down...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.