OSSbenchflow-ai/skillsbench Agent Skills BenchmarkGitHub·high signalXBlueskyLinkedInCopy linkFirst standardized benchmark for agent skills. Key finding: curated skills boost 16.2% but self-generated skills provide zero benefit.SourceSource pageGitHub↳ Follow the threadStack layer / Threat patternPentest Harness reaches 306 stars as a self-hosted, bring-your-own-key agent harness for authorized engagementsGitHubStack layer / Update threadDeepChat v1.1.1 introduces Code and Minimal tool modes with on-demand tool discovery and permission-scoped CLI accessGitHubStack layer / Threat patternBekchiAI Releases 2,057 Verifier-Checkable Agent Tasks Where Gold Answers Are Computed, Not WrittenarXiv 2608.26867Stack layer / Follow-up threadOpenSEO is a self-hostable Semrush alternative whose primary interface is MCP, at +517 stars todayGitHub TrendingPolicy dependency / Stack layerThe EU AI Office Sent Its First Formal Requests for Information to Frontier Model Providers on August 29TokensteadStack layer / Threat patternPromptfoo adds a Codex Security SDK provider and hardens its code-scan GitHub Action supply chainGitHubStack layer / Threat patternheadcount packages Claude Code skills as 16 installable departments with 143 skills and namespaced addressingcbrock84/headcountSame source domain / Semantic neighborWarp is publishing its internal agent skills as warpdotdev/common-skills, pushed again this morningGitHub Trending