Fetching from the wire…
Agents2026-09-15 · source-backed
Restart and repeat finished work, or continue from unverified progress and carry earlier errors forward (arXiv 2609.13672). This paper makes reuse an explicit decision, binding the choice of starting point and the permitted recovery action to supporting evidence, execution and independent checks. Four deterministic and 20 paired file challenges show accurate restoration and successful completion can both hold while the starting point was not permitted, and event-time tests show permission has to constrain the recovery action, not just the restored bytes.
Each link below shares sources, entities, or timing with this story.
The failure they target is specific and under-discussed: a cached error page or a negative price returns in the *expected schema* and gets consumed as fact, unlike a timeout the agent can see. Outcome Monitors check results against contracts mined from task-disjoint traces or...
The thread connecting Cherny's two-week Swift port (screenshot diff against the running Electron app) to the HANDBOOK.md compliance results (36.2% compliance with prose policy) is that agents drift against text and hold against executable checks. Concretely: replace "match the...
arXiv 2607.27080 traces malicious semantics through persistence, downstream consequence, and selective repair across 310 test cases in 48 contexts, under a 24-configuration matrix of 2 harnesses × 4 memory backends × 3 LLM backends. Malicious memory persists in 84.2% of cases,...
arXiv 2607.26998 flips the pentest agent's observation-action loop against it, replacing static honeytokens with a trajectory-adaptive policy that constructs new decoy artifacts conditioned on the agent's interaction history, folding validated ones into a factually consistent...
Every guardrail I've used adjudicates the current action, which means it can only stop the last step of a plan it never saw coming. JANUS trains a guard on partial trajectories to anticipate safety-relevant futures, then judges from both the observed prefix and the forecast, o...
In a GitHub Copilot SDK setup, an asynchronous memory-curator agent got read-only tools to check candidate memories against the current state before saving them. Pass rate on CLBench rose to 73% from 39% (arXiv 2609.11060). Queries per question fell to 4.7 from 8.8, and task-a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.