ResearchCUDA Agent Agentic RL for CUDA Kernel GenerationarXiv·high signalXBlueskyLinkedInCopy linkAgentic RL system outperforms torch.compile by 100% and beats Claude Opus 4.5 by 40% on GPU kernelsSourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerSplit your agent's safety into four evolvable artifacts — system prompt, rule bank, safety memory, tool policy — for a 3.1x attack-success reductionarXiv 2608.09885Stack layer / ContrastDeterministic rule-guided dispatch beats autonomous code-review agents by 2.17x SEM-F1 on 5–15x fewer tokensarXiv 2608.09290Policy dependency / Stack layerPOLIS 5,280-Episode Study: Provenance-Aware Guards Block Authority Laundering That Local-State Guards Miss in 22 of 96 EpisodesarXiv 2608.09828Stack layer / Update threadSimWAM Hits 91.5 PDMS on NAVSIM by Throwing Away the Video Model After TrainingarXiv 2608.07468Stack layer / ContrastStop evaluating model routing by replaying logs: only 3% of replayed agent states are still valid, and replay mispredicted every success-relevant outcomearXiv 2608.08239Policy dependency / Update threadMatrAIx Builds a Simulated World of 8.3 Billion Persona Agents Across 1,290 Dimensions — and Releases a 1M Coreset With 91.5% Behavioral AdherencearXiv / HuggingFace Daily Papers (623 upvotes)Stack layer / Update threadKGCaRe Beats Think-on-Graph and Vanilla RAG by Traversing an LLM-Built Knowledge Graph Iteratively, Re-Entering With Clue Entities When Context Falls ShortarXiv 2608.09779Stack layer / ContrastPerplexity can't detect RAG poisoning — poisoned answers score *lower* perplexity; watch attention entropy collapse insteadarXiv 2608.06947