ResearchAgentSkillOS Ecosystem Scale Skill Orchestration FrameworkarXiv·high signalXBlueskyLinkedInCopy linkFirst principled skill selection and orchestration for 200-200K skills via capability tree and DAG-based pipelines. Benchmarked on 30 tasks.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerRetrieval that crosses into your dependencies' source, not just your repo, adds up to 6.3% pass@1 and survives version changesarXiv 2609.09987Stack layer / Threat patternAgentAudit attaches to a running agent and scores its trace on ten dimensions, exposing 95.1 vs 22.6 trust spreads at similar task completionarXivStack layer / ContrastCapScope stops prompt injection by giving each coding subagent typed capabilities stored outside its contextarXivStack layer / ContrastA spec-first agent framework taxonomy: persuasion, front-loaded structure, or controls the agent cannot editarXiv 2609.09671Stack layer / ContrastEcdysis: fix the harness only for failure patterns that recur across tasks, not for each single failurearXiv 2609.11677Stack layer / Follow-up threadΦ-Bench tests whether LLMs can engineer their own serving and training stack, from kernels to end-to-end optimizationarXivStack layerReplacing the manager LLM in a compound system with a deterministic merge operator beat generative managers by 0.048-0.076 task-score pointsarXiv 2609.09815Stack layerShow-Harness gets frontier VLMs controlling robots zero-shot through discrete semantic action unitsarXiv