DeskSkillsBench Agent Skills BenchmarkWweb·high signalXBlueskyLinkedInCopy linkCurated skills +16.2pp pass rate. Self-generated skills zero benefit. Validates human-curated skill model.↳ Follow the threadStack layer / Threat patternBekchiAI Releases 2,057 Verifier-Checkable Agent Tasks Where Gold Answers Are Computed, Not WrittenarXiv 2608.26867Stack layer / Threat patternAgentDV Lifts RTL Testbench Generation From Zero Valid Environments to 100% Pass on Four DUTsarXiv 2608.27148Stack layer / Threat patternPentest Harness reaches 306 stars as a self-hosted, bring-your-own-key agent harness for authorized engagementsGitHubStack layer / Update threadSKILL.state throws away the conversation history: only the skill spec, current state, and latest observation reach the modelarXivStack layer / Threat patternheadcount packages Claude Code skills as 16 installable departments with 143 skills and namespaced addressingcbrock84/headcountStack layer / Threat patternGenIaC-SecBench gives LLM-generated infrastructure code a size-matched human baseline for the first timearXivStack layer / Threat patternOpenClaw 2.0 Ships as an Open-Source Agent Platform That Installs on Your Existing Subscription Instead of Selling OneOpenClawStack layer / Update threadSkill Routing Fails Because It Matches the Task and Ignores the User: SkillFeed Gains 35.1 Points Where the Profile DecidesarXiv 2608.28241