Fetching from the wire…
Public story · 2026-09-22 · high
AgentPluginZoo scanned 68,072 bundles and found most work anyway, but 81% of capability-exporting plugins share a name with no rule for which one wins.
Why now: AgentPluginZoo's measurement gives the first bundle-scale check of the Agent Plugins v1.0.0 spec, published July 24, 2026.
AgentPluginZoo checked whether 68,072 agent plugin bundles conform to the Agent Plugins v1.0.0 spec, and found that only 6.2% do.
That gap matters for anyone stitching multiple plugins into one agent setup. The corpus spans 30,655 repositories, and the spec has no rule to stop two plugins from claiming the same capability name.
The failure rate looks worse than it is. 96.6% of bundles would validate after adding one missing boilerplate field, so most authors just skipped a formality rather than broke the spec.
The deeper problem sits one layer down. 40.2% of bundles would only load if the client silently drops fields their authors wrote into the manifest. Those dropped fields represent capabilities the authors declared that clients discard without telling anyone.
The sharper number is the collision rate. Among bundles that export capabilities, 81% share a name with at least one other plugin. The spec sets no namespace or precedence rule to decide which one a client loads. Install two plugins that both register a capability called search, and nothing in the spec names a winner.
The paper behind AgentPluginZoo, Packaged, But Not Portable, argues the packaging format is a distraction from what's missing. The spec standardized packaging. It didn't standardize composition. It names four missing concepts: qualified capability identity, a declared capability surface, a precedence rule, and inter-plugin relations.
If you're building on this spec, watch the collision rate, not the validation rate. A plugin ecosystem where 4 in 5 capability-exporting bundles can clobber each other's names needs a resolution rule before it needs more bundles.
Each link below shares sources, entities, or timing with this story.
The errors trace back to how the benchmark pairs pull requests with GitHub issues, not just to model quality.
A new benchmark of 203 real upgrade tasks shows breaking changes that never made it into a changelog trip up even the best agent setups.
It splits agent composition from runtime adaptation, and its GitHub repos are still active, not archived research code.
Missing argument logs hid why the agents failed for ten days across 8,199 runs testing 40 open and hosted models.
Cross-vendor AI review still shows up in just 1.6% of agent-authored pull requests, but reviewers grade outside code more harshly than their own.
Attackers who know only a target's role profile can chain marketplace skills into working attacks; success drops off after three hops.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.