Policy dependency / Stack layer
Naive Retries Cut Success From 55.4% to 41.5% Under Correlated Failures; 1 of 113 Production Configs Randomizes Delay
arXiv 2608.25403
Policy dependency / Stack layer
Five runtime primitives for governing agents, and an explicit accounting of what enforcement costs
arXiv
Shared entity / Stack layer
Meta Tests a Consumer AI Agent, Reportedly Codenamed Hatch, Trained on DoorDash, Etsy, Reddit and Yelp
Finextra
Stack layer / Threat pattern
SkillShield defends coding agents from the system prompt alone, matching Llama Guard 3 with no runtime classifier
arXiv
Stack layer / Threat pattern
A survey formalizes AI slop in vulnerability assessment and proposes measuring the gap with a Deductive Coverage Score
arXiv
Stack layer / Threat pattern
MACGen splits secure code generation into planner, security advisor, coder and reviewer passing structured artifacts instead of shared dialogue
arXiv
Policy dependency / Stack layer
Borrowed Authority: Agent Skills Carry No Typed Way to Reject a Permission Claim, and Edge Skillguard Rejects 60/60 Attacks
arXiv 2608.25091
Stack layer / Threat pattern
ReDiR compresses the whole trajectory into a latent safety signal before each action, cutting multi-turn decomposition attacks below 8%
arXiv