Fetching from the wire…
Agents2026-09-07 · source-backed
Eight teams per setting formed independently from one base model, each agent keeping a private notebook across ten formation episodes, then role-matched agents were traded between teams (arXiv 2609.05279). Against a placebo reproducing roster-change disruption without changing the occupant, the swap raised communication per unit of progress by 16 to 63 percent. In Hanabi a swapped agent costs more than an inexperienced one. Most of the extra communication in Collab-Overcooked comes from the agent that stayed, not the newcomer. The production assumption that any agent filling a role substitutes for any other is measurably wrong.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.26935 argues the human-vs-bot label space can't represent agent traffic: an MLP binary classifier misroutes 39.1% of real agent sessions as human, a SAINT transformer 34.5%, while adding an explicit third class yields agent F1 = 1.000 across all 30 runs. Against a f...
Self-hosted agents read and write their own memory and config to function, which means an attacker can compromise one entirely through legitimate OS system calls with no exploit involved (arXiv 2607.17986). The paper builds a 23-cell attack matrix across Target, Mechanism, Gra...
Crase (arXiv 2608.24809) is deliberately non-agentic for scholarly search: one search-engine query for seeds, expansion along the 1.5-hop citation neighborhood, pruning of citation edges whose claims lack entailment support, then a recency-aware random walk for ranking. Candid...
arXiv 2608.04682 removes the assumption that every SWE benchmark makes, that a high-quality issue report exists. Six bug categories, eight languages, multi-bug fixing and potential-bug discovery under dual-track evaluation. Most state-of-the-art coding agents perform poorly at...
In a preregistered 18,000-mission evaluation scored by deterministic code with no LLM judge, two instances of one model in a two-agent handoff co-failed on 90.0% of missions where either failed (log OR 6.66, phi 0.916). Swapping in a different model reduced the association in...
arXiv 2608.12921 uses causal inference to find which edges in a communication graph actually carry signal, then removes the rest. Most multi-agent stacks default to broadcast or a fully-connected mesh and pay for every edge in tokens, so edge-level attribution is a direct cost...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.