Fetching from the wire…
Agents2026-09-09 · source-backed
CapScope derives a task-wide authority ceiling from trusted input before any repository content or tool output is read, then gives each sub-agent typed capabilities stored outside the model's context. Every tool call checks against the issuing agent's capabilities, so one sub-agent's permissions are never inherited. Across 300 runs it completed 68/75 repairs against baselines' 68-72/75. arXiv 2609.08371 Naming a resource shouldn't be sufficient authority to act on it, and the check belongs in the harness, not in a classifier hunting for injected text.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
The failure mode is a well-formed but policy-forbidden call, cancel a booking, change a passenger count, that neither the tool nor the agent's self-report flags (arXiv). In the airline domain tested, the fix wasn't more reasoning. It was cheap, read-only deterministic gates th...
arXiv 2608.26197 stacked finite-state control, forced tool selection, output validation and bounded retries on two open-weight models, and got mixed results across all four model-task cells. Adding structured planning, where the plan is checked against a fixed schema before an...
100 real frontier research tasks across seven scientific domains, full lifecycle, 800 annotated trajectories, 45-pattern failure taxonomy (arXiv 2608.14905). The headline isn't a leaderboard, it's a shared deficit: agents can't check what they produced against what they found,...
arXiv 2608.07167 intercepts every tool call, validates against a SHA-256-locked Intent Contract using an isolated Judge model, then proves via EZKL that the safety check ran without exposing weights. F1 88.5% at a 1.1% false-positive rate on Agent-SafetyBench. Generation costs...
The thread connecting Cherny's two-week Swift port (screenshot diff against the running Electron app) to the HANDBOOK.md compliance results (36.2% compliance with prose policy) is that agents drift against text and hold against executable checks. Concretely: replace "match the...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.