SourcesPapersWweb·medium signalXBlueskyLinkedInCopy link**SkillRL: Recursive Skill-Augmented Reinforcement Learning** — [arXiv](https://arxiv.org/abs/2602.08234) (medium). Agents learn reusable skills organized hierarchically. 10-20% token compression, 15.3% improvement on ALFWorld/WebShop. Most practical paper on making agents actually improve from experience.↳ Follow the threadPolicy dependency / Stack layerSAPO Shares One Autoregressive Backbone for Policy and Value, Beating PPO by 15.1 Points With No Separate CriticarXiv 2608.19842Stack layer / ContrastMeasured Across 9 Models: Telling an LLM to Be Concise Saves ~1.5x, Shortening Your Prompt Costs Up to 96% Morer/MachineLearning (49 upvotes)Stack layer / ContrastReward-guided graph generation cuts multi-agent communication tokens 20.5% at equal accuracyarXivPolicy dependency / Stack layerOpenRouter's Free Stealth Model 'Ox Alpha' Ships a 1,048,576-Token Context and Video Input, and HN Spent 142 Comments Trying to Unmask ItOpenRouter / Hacker NewsStack layer / Update threadOptimal MoE Learning Rates Extrapolate to 10 Trillion Tokens From Small Proxy Runs at R-squared 0.95arXiv 2608.20061Stack layer / ContrastCOPA Treats Prompt-Injection Defense as Lifelong Learning, Cutting Attack Success Up to 6.3xarXiv 2608.19982Stack layer / ContrastA 250M Model Trained From Scratch Ships in 60MB, Runs 400 Tokens/Sec on Laptop CPU, and Pages 1-Bit Context to Diskr/MachineLearning (112 upvotes)Stack layer / Threat patternWarp CEO Zach Lloyd launches Warp Factories and puts a number on it: 'we automate 30 to 35% of our tasks on a weekly basis'Warp