Stack layer / Threat pattern
How you lay out your repo changes prompt-injection success: highly modular workspaces measurably lower attack success rate
arXiv 2608.14876
Policy dependency / Stack layer
Compose agent guardrails as an algebra instead of a rule list: 94.8% of policy-violating events intercepted while keeping 86.9% task completion
arXiv 2608.16402
Policy dependency / Stack layer
Envs-FORGE Solves a Per-Seed MILP to Synthesize Agent-RL Environments, Hitting 77.1% on SWE-bench Verified vs 73.4% Base
arXiv 2608.14312
Stack layer / Threat pattern
MCP tools that write, not read, went from 27% to 65% of tool use - and measured defenses stop under 30% of attacks
arXiv 2608.17275
Stack layer / Contrast
'How Do Agents Fail on AutoResearch': 800 Trajectories Across 8 Harness-Model Combos Find One Failure Every Model Shares — They Never Check Output Against Evidence
arXiv (via HuggingFace Daily Papers)
Stack layer / Threat pattern
SkillWatermark: benign-looking skill descriptions turn agent network traffic into a covert exfiltration channel
arXiv
Stack layer / Contrast
Deliberately mixing low-relevance same-domain items into context improved relevance accuracy by +0.077 - and six production patterns cut tokens 60-70%
arXiv 2608.17188
Stack layer / Contrast
Agents predict their own failure well (0.8847 AUROC) but cannot tell you which collaboration protocol will fix it
arXiv 2608.14927