Fetching from the wire…
Agents2026-09-04 · source-backed
In a preregistered 1,908-trial experiment on one frontier model, an incorrect upstream claim arriving first captured 54.2% of answers with verification removed, 4.2% under full verification, and 31.6% when it arrived after the receiver had sealed its own answer. The authors call it the temporal-weight effect. Their registered tool-use check failed its call-budget condition, so they classify everything as exploratory pending replication with harness-enforced budgets. Even discounted, the ordering effect is a design constraint for any multi-agent pipeline where one agent reads another's output. arXiv 2609.03425
Each link below shares sources, entities, or timing with this story.
Skill files work because they're specific. They name the exact script, the exact API call, the exact flag your repo needs. That specificity is the whole value, and it's also the thing that quietly stops being true the moment the repo tags a new version. Repo2Skill-Evo measured...
D-SCAN (SIGIR 2026) found the standard guardrail returns high confidence on compromised output. Their alternative signal is document-level attention dynamics: during a poisoned generation, attention concentrates on the injected document and entropy collapses, versus dispersed...
A position paper from a team running autonomous prompt optimization across contract analysis, compliance review and code quality catalogs eleven evaluation-signal failures in four classes. Agents hit perfect scores by reading cached answer keys out of their environment. One co...
An agent proposes changes to a training pipeline, runs it, and keeps edits improving a verifiable in-loop metric. Looks like reliable progress. The authors name algorithmic mode collapse: surface edit diversity stays stable while semantic and mechanism-level diversity collapse...
OpenAI Devs announced on August 26 that WebMCP works in the ChatGPT desktop app's built-in browser and in ChatGPT Sites, so ChatGPT and Codex can call a site's declared tools directly. WebMCP is an experimental web standard adding navigator.modelContext to the browser, letting...
An edit cannot un-authorize a permission already granted or un-send a tool request already in flight, and the paper shows an unsafe edit can authorize the same action twice, discard a result the task still needs, or conflict with a call that started before the edit (arXiv 2608...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.