Policy dependency / Stack layer
Self-Distillation Without Ground Truth Matches GRPO: U-OPSD Gains 8.5-10.7% on AIME/MATH Using Only the Model's Own Rollouts
arXiv 2608.06296
Stack layer / Contrast
Generative Reward Models Underperform in RL Because They Rank, Not Score — RRC Fixes the Mismatch
arXiv 2608.06310
Policy dependency / Stack layer
AgentOPSD: Critic-Free Turn-Level Credit Assignment via Bayesian Belief Updates Hits 89.1% on ALFWorld With a 7B Model
arXiv / HuggingFace Daily Papers (57 upvotes)
Policy dependency / Stack layer
ToolArtist Puts the Entire Image-Generation Workflow Under One Agentic Policy, Training It by Hiding the Generation Tool From Its Own SFT Traces
arXiv / HuggingFace Daily Papers
Stack layer / Contrast
LLMs Pick Python Even When It's the Wrong Language, and Fabricate Reasons — "Phantom Evidence" Named in 9,826 Reasoning Traces
arXiv 2608.06041
Stack layer / Contrast
Agent Signing Keys Move Into HSMs: PKCS#11 Keystore Plus Zero-Trust MCP Stack Drops Injection Success From 19.3% to 0%
arXiv 2608.06130
Stack layer / Contrast
Programmatic tool calling beats JSON tool calling on 11 of 14 models — the BFCL v4 head-to-head
arXiv
Policy dependency / Stack layer
CROSS-CATEGORY: Agent Containment Shipped in Three Unrelated Products in 48 Hours — While PromptArmor Showed Atlassian's Rovo Still Leaking
Multiple Sources (PromptArmor, Zed, Cloudflare, Mistral AI)