ResearchATLAS: 4B Model Matches Frontier on Agentic Tasks via Rubric-Based RLarXiv·high signalXBlueskyLinkedInCopy linkTreats context acquisition and tool selection as learnable behaviors. 4B params matches frontier performance. Solves eager tool loading and sparse rewards.SourceSource pagearXiv↳ Follow the threadStack layer / ContrastE2-Explainer turns multi-agent communication topology from black-box optimization into a causal attribution problem — then prunes the grapharXivStack layer / ContrastDFM Mimir v1: 1B-Parameter Hierarchical Reasoning Model Trained Entirely on Permissible Data, Competing With Qwen 3.5 4BarXiv 2608.13517Stack layer / Threat patternHARD Turns LLM Agent Runtime Defense Into a Self-Evolving Loop Instead of Hand-Written RulesarXiv 2608.12977Stack layer / ContrastTwo-Tier 'Librarian + Writer' Architecture Removes 6,845 Cross-Section Contradictions From LLM Research Reports to ZeroarXiv 2608.12984Stack layer / Threat patternSRE-Bench: 5,000 expert hours to build the first contamination-free reverse-engineering benchmark, and frontier agents solve only 31.5% of itarXivStack layer / ContrastERSkill makes memory retrieval itself an evolvable skill, improving agent-memory benchmarks by 31.3%arXivPolicy dependency / Stack layerExport Controls Turn Frontier Model Access Into a Cyber-Defence Dependency for Middle PowersarXiv 2608.13272Stack layer / Follow-up threadFaraday: a 27B agent trained specifically to replicate research beats Claude Opus 4.8 and GPT-5.5 on held-out replicationarXiv