Fetching from the wire…
Agents2026-09-22 · source-backed
Runtime Authorization Consistency Checking names the failure where each MCP step is individually legal while the accumulated sequence exceeds what the session was granted. RAC treats authorization as runtime state carried by accepted steps, reconstructs a trusted authorization event from controller-observed metadata at the tool-call boundary, and admits a call only if it's no more permissive than the basis inherited through accepted lineage, with rejected steps dropped from lineage. On the 1,248-workflow TraceBench suite RAC had zero missed blocks where the strongest Static+History baseline missed 509 of 1,008, reached 92.8% block recall on blind LLM-generated plans against 68.8%, and ran at sub-millisecond p99. (arXiv 2609.23498)
Each link below shares sources, entities, or timing with this story.
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
A study of 3,171 GitHub repositories (2,660 multi-component setups, 511 skill collections) measured only byte-decidable defects in Claude Code, Cursor, Copilot and Codex configuration artifacts, validating every finding through independent re-derivation, an LLM adjudicator and...
Three separate Anthropic changes over about two weeks point the same direction, and none of them announced themselves as a strategy. Claude Code 2.1.238 added claude self-hosted-runner --defer-shutdown-max-min, which keeps serving attached sessions on SIGTERM, parks whatever's...
SkillsMetric evaluated 2,266 skills across 16 attack types, hitting F1 of 73.4%±0.5% overall (arXiv 2608.08468). Host destruction via shell commands: 0% detection. Natural-language prompt injection: 42%. If you lint third-party skills before install, this tells you precisely w...
An ArXiv study analyzing Claude Code's design space found something that should make every "auto-generate your context files" workflow uncomfortable. Human-curated CLAUDE.md files improved task success rates by roughly 4 percentage points. LLM-generated CLAUDE.md files reduced...
| # | Skill | Domain | Difficulty | |---|-------|--------|------------| | 1 | Claude Code /simplify + /batch — three-agent parallel review + codebase migrations | vibe-coding | intermediate | | 2 | Pipelock agent firewall — 9-layer DLP + MCP scanning inline proxy | agent-secur...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.