Fetching from the wire…
Agents2026-08-02 · source-backed
Most systems treat topology as a fixed design choice or an offline optimization target. MANTA initializes task-conditioned from prior structural experience, then monitors traces during deployment and applies bounded structural updates to agent roles, communication links, execution order, information visibility, and validation pathways: while preserving task interface and agent budget. 5.8 points above the strongest baseline.
Each link below shares sources, entities, or timing with this story.
Microsoft Research dropped a paper that should change how every builder thinks about their agent configuration files. SkillOpt (arXiv 2605.23904) treats a Markdown document as an external parameter of a frozen LLM and applies learning rate, batch, and momentum concepts in text...
This one is strange enough that I want to be careful about how strongly I state it. "Workspace Topology as an Attack Vector in Agentic Coding Assistants" (arXiv 2608.14876) is the first empirical study I've seen that treats repo layout as an attack surface. The variables: dire...
MobileWorldSafety (arXiv 2608.17659) embeds injection attacks in Android apps and evaluates with a two-stage pipeline combining rule-based verification with LLM adjudication, specifically to separate safety failures from capability failures. That distinction is the contributio...
Most long-horizon agent work invests in plan refinement and pre-flight safety checks, which leaves nothing once an early error has already corrupted both the agent context and the environment state (arXiv 2608.14380). AgentRewind records aligned checkpoints of context and a co...
BFCL v4 results show PTC matching or beating JSON tool calling on 11 of 14 models, with the GPT-5.6 family up 10.6% and better stability under context degradation and parallel execution. Most agent frameworks hard-code structured output as the default. On current models that d...
arXiv 2608.04682 removes the assumption that every SWE benchmark makes, that a high-quality issue report exists. Six bug categories, eight languages, multi-bug fixing and potential-bug discovery under dual-track evaluation. Most state-of-the-art coding agents perform poorly at...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.