Policy dependency / Stack layer
A controlled ablation on a production science agent finds the model choice dominates topology and prompting, and a PPO policy nearly matches it for free
arXiv
Policy dependency / Stack layer
Attnlocate treats prompt injection as an object detection problem inside the attention matrix, hitting 0.934 TPR at 0.067 FPR
arXiv
Stack layer / Contrast
CyberFactory Turns CVEs From the Wild Into Executable Training Tasks; Aegis Hits 52.4% Pass@1 on CyberGym, +22.8 Over Its Base
arXiv 2608.23181
Stack layer / Update thread
IBM Released Granite 4.2 as Dense Reasoning Models at 3B, 8B and 30B Under Apache 2.0 With a 512K Context
Hugging Face (ibm-granite), via r/LocalLLaMA
Stack layer / Threat pattern
ReDiR compresses the whole trajectory into a latent safety signal before each action, cutting multi-turn decomposition attacks below 8%
arXiv
Stack layer / Threat pattern
SCALE-QA Tests Memory in Flat Unsegmented Threads; Episode Reconstruction Beats Long Context by 5.6-17.6 Points
arXiv 2608.25655
Policy dependency / Stack layer
Best Practice Critic Optimization Matches GRPO While Sampling One Response Per Prompt
arXiv 2608.23566
Stack layer / Update thread
Salesforce and Anthropic Ship Claudeforce, Putting 37 Sales Skills Inside Claude So Sellers Never Open Salesforce
Salesforce Newsroom (corroborated by VentureBeat, 2026-08-26)