Policy dependency / Stack layer
A Fine-Tuned RoBERTa-Large Permission Gate Matches Claude Haiku 4.5 at Deciding What an Agent May Touch
arXiv 2609.15422
Policy dependency / Stack layer
RIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persisting
arXiv 2609.12127
Stack layer / Contrast
Two-gap framework recasts reward hacking and hallucination as symptoms of requirement and model gaps
arXiv
Stack layer / Contrast
Pre-registered ablation shows removing an LLM verifier stage from an offensive-security agent shifts median reported findings from 0 to 2 per run
arXiv
Stack layer / Contrast
ActGuard audits the planned action instead of filtering tool output, comparing each step against a local tool prior
arXiv
Policy dependency / Stack layer
HazardAuditor runs Claude Code, Codex, Hermes and OpenClaw in one harness and normalizes their events to train a guard model
arXiv / HuggingFace Daily Papers
Policy dependency / Stack layer
CodeBLEU Scored 91% for Both RAG Strategies While One of Them Hallucinated APIs 56.4% of the Time
arXiv 2609.12464
Stack layer / Threat pattern
Agent Frameworks Detect Dangerous Plan Steps and Then Execute Them Anyway; Fewer Than 20 Lines Closes the Gap
arXiv 2609.15293