Policy dependency / Stack layer
POLIS 5,280-Episode Study: Provenance-Aware Guards Block Authority Laundering That Local-State Guards Miss in 22 of 96 Episodes
arXiv 2608.09828
Stack layer / Contrast
ActBench: Attack Success Against Cowork Agents Ranges 10.1%–94.4% Across Models but Only 73.7%–94.4% Across Harnesses — the Model Matters More Than the Scaffold
arXiv 2608.09476
Stack layer / Threat pattern
Systematic Review of 85 Agentic-LLM Security Papers: Attack Work Outpaces Defense 3.9:1 and Action-Layer Risk Gets Only 4.7% of Attention
arXiv 2608.10530
Policy dependency / Stack layer
Split your agent's safety into four evolvable artifacts — system prompt, rule bank, safety memory, tool policy — for a 3.1x attack-success reduction
arXiv 2608.09885
Stack layer / Contrast
Deterministic rule-guided dispatch beats autonomous code-review agents by 2.17x SEM-F1 on 5–15x fewer tokens
arXiv 2608.09290
Stack layer / Contrast
MasDrift finds the centralization tradeoff: hierarchical agent teams finish more tasks and take up to 19.8% unauthorized actions doing it
arXiv
Stack layer
On-Policy Distillation Has a Degenerate-Agreement Failure Mode; TIDE Lifts Avg@8 From 6.9% to 20.3% and Cuts Response Length 3.6x
arXiv 2608.09836
Stack layer
Bengio and Goldwasser put probabilistic self-consistency in NP, with a verifier that checks exponentially many implied claims in polynomial time
arXiv