Fetching from the wire…
Security2026-09-23 · source-backed
arXiv 2609.26529 documents four ways an approved action gets rebuilt before execution: workflow reloads, transcript projection, argument rebinding, state lookups. A repair that validates every field can still leave an authorization bypass. Their APAS-Finder matched all 28 controlled outcomes, and five correct repairs blocked 60 of 60 out-of-scope effects. Bind approval to the exact object version consumed at execution, with a nonce or hash, not to the record the UI displayed.
Each link below shares sources, entities, or timing with this story.
113 non-developer participants ran an 18-action simulated agent day containing 7 overreach actions under three regimes (arXiv 2608.27443). User-authored consequence policies blocked 20.1 points less overreach than human-in-the-loop approval (95% CI [-32.1, -8.1]) and 14.5 poin...
StagedWorkspace (arXiv 2608.18050) targets a failure most knowledge-work stacks have and none measure: different components of the same agent referencing different versions of the same artifact. The fix is explicit version bindings between parsed records, review diffs, and nat...
arXiv 2608.07167 intercepts every tool call, validates against a SHA-256-locked Intent Contract using an isolated Judge model, then proves via EZKL that the safety check ran without exposing weights. F1 88.5% at a 1.1% false-positive rate on Agent-SafetyBench. Generation costs...
arXiv 2608.02764 targets agents that issue refunds, reserve inventory and move money, where budgets and approval status change between authorization and effect. The authors define policy-state serializability: committed effects must be explainable as authorized against the pol...
When an agent consolidates an external observation into long-term memory, attach platform-controlled metadata recording the source's trust level, then gate tool execution by matching action risk against supporting-memory authority. Laundered memories hit a 1.000 attack success...
arXiv 2607.26598 targets the failure where an agent recovers from an error within an episode but hits the identical failure in later tasks, because post-episode feedback never revises the persistent harness. Guided by a domain-level Evolution-SOP, it writes episodic memory rec...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.