Policy dependency / Stack layer
Pre-registered ablation shows removing an LLM verifier stage from an offensive-security agent shifts median reported findings from 0 to 2 per run
arXiv
Policy dependency / Stack layer
Two-gap framework recasts reward hacking and hallucination as symptoms of requirement and model gaps
arXiv
Policy dependency / Stack layer
RIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persisting
arXiv 2609.12127
Policy dependency / Stack layer
TRAIL pairs a translator agent against a challenger agent and gains 23.1% relative syntax accuracy on C-to-Rust translation
arXiv
Policy dependency / Stack layer
Difficulty-aware topology selection beats always-hierarchical multi-agent coding by 4.1 points at 40% of the cost
arXiv
Policy dependency / Stack layer
Agents Convert Tokens Into Progress Faster Than Independent Sampling at First, Then Fall Below It
arXiv 2609.15309
Policy dependency / Stack layer
Splitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops Fail
arXiv 2609.12839
Policy dependency / Stack layer
Atria Dawn's authors studied 769 task records from their own build, and participants called a third of the AI-assisted tasks infeasible without AI
arXiv / HuggingFace Daily Papers