Fetching from the wire…
Agents2026-09-22 · source-backed
Ask a privileged agent to do a new system task and it will take a file, process, socket, lock or capacity allocation from a healthy incumbent, since privilege decides whether an operation can run and never whether the requester may preempt the current owner. LeaseGuard is a deterministic admission layer before adapter execution, modeling preemption authority as resource leases with incumbent-health checks and coexistence limits. Across 60 conflict scenarios on two local model families, unauthorized preemption fell from 73.3% to zero, safe completion rose 70 points, and requested-task success cost 3.3 points. The authors flag their own hole: expiry-only reclamation can still expose a healthy incumbent after a missed renewal. (arXiv 2609.24077)
Each link below shares sources, entities, or timing with this story.
Skill files work because they're specific. They name the exact script, the exact API call, the exact flag your repo needs. That specificity is the whole value, and it's also the thing that quietly stops being true the moment the repo tags a new version. Repo2Skill-Evo measured...
arXiv 2609.20301 argues existing observability tools do per-execution debugging but not cross-run profiling, so nobody can answer where failures cluster or which tasks eat the budget at scale. The obstacle is that the responsible entity is a task intent like "diagnose authenti...
A fleet evaluation across 46 endpoints from six vendors found a recognition-enforcement gap: source-format features are linearly decodable from activations and models verbally identify forged authority when asked, but some configurations still emit the conflicting tool call. A...
ECP captures agent outputs, tool invocations, and audit context uniformly, with adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI so the same checks run against any of them. arXiv The authors explicitly label it work-in-progress with the method set expected to change....
When an agent consolidates an external observation into long-term memory, attach platform-controlled metadata recording the source's trust level, then gate tool execution by matching action risk against supporting-memory authority. Laundered memories hit a 1.000 attack success...
A 10,626-instance benchmark built on a 14-type error taxonomy, scoring structural integrity, diagnostic reasoning and recovery strategy during execution rather than at the end. Across ten-plus mainstream models, multi-turn error propagation and implicit tool-use failures remai...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.