Fetching from the wire…
Agents2026-09-23 · source-backed
arXiv 2609.25960 replays payment-exception episodes and treats dropped settlement messages as intervenable causes alongside agent actions. Across 545 planted episodes, agent-step attribution blamed the agent for every infrastructure fault, and fixing what those methods named recovered 0.0% of the loss. Fixing a minimal sufficient message set recovered 100%. 27.8% of episodes didn't decompose additively at all. Post-mortems that only replay agent decisions will blame the agent for your network.
Each link below shares sources, entities, or timing with this story.
Blaming the right agent in a failed multi-agent trajectory is currently done with prompting or fine-tuned long-context models. AFANet models step-level semantic signals and agent-level relationships as a graph, and with far fewer parameters and near-zero inference cost it matc...
The authors define goal-directed execution as four repeated behaviors: selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, verifying completion against the environment. Post-training Qwen3.5-122B-A10B on 363 long-horizon multi-to...
arXiv 2608.11392 studies what happens when a long-running agent compacts its context: a standing constraint frequently persists as textual residue that no longer governs behavior. Behavioral replay shows models perform the prohibited action far more often with a degraded resid...
arXiv 2607.26791 benchmarks post-compromise incident response and reports agents struggle to proactively investigate silent intrusions. They respond to what they're pointed at. Read alongside the July intrusion post-mortem, that argues against putting an agent on the detection...
A paper submitted September 18 names a failure class I've been half-aware of and never had a word for, then reproduces it in software you probably have installed. Loopjacking describes the case where a human approves operation A and the runtime executes a materially different...
535 participants solved a 40-item reasoning battery either unaided or while required to consult GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash or Kimi K3, with each model also answering every item alone 100 times under matched elicitation. In a reference comparison, roughly h...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.