Fetching from the wire…
Policy2026-07-25 · source-backed
Llambí-Morillas and Fernández-Fernández formalize CVA as a relation jointly binding an agent principal, a concrete authorization request, an execution context, and policy satisfaction while keeping private authorization attributes confidential. They define authorization soundness, principal binding, request binding, policy binding and replay resistance, and ship an executable zero-knowledge proof of concept over Groth16. Their central claim is structural: existing agentic security frameworks don't explicitly separate identity binding from authorization-request binding from runtime execution binding, and that conflation is the open problem. If you're building agent auth, that three-way split is the design constraint.
Each link below shares sources, entities, or timing with this story.
Skill files work because they're specific. They name the exact script, the exact API call, the exact flag your repo needs. That specificity is the whole value, and it's also the thing that quietly stops being true the moment the repo tags a new version. Repo2Skill-Evo measured...
Schmitz, Hammond and Chan define agentic flooding as demand surges caused by systems that make interacting with government cheap, with LLM text generation as the primary enabling mechanism (arXiv 2608.16603). Accepted to AAAI AIES 2026. Their risk matrix places near-term risk...
Four stories about things going wrong. Here's one about something working, with actual numbers attached. In an August 7 disclosure covered by TechCrunch, Airbnb said AI now writes 60% of its new code, that concept-to-launch time on key initiatives has dropped by as much as 60%...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
Here's the experiment: a team of cooperating agents rebuilds SQLite in Rust from scratch, using only the 835-page manual. No source code. No test suites. No internet. Then it has to pass a held-out sqllogictest suite. It worked. Cursor published the research (Wilson Lin, July...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.