Fetching from the wire…
Public story · 2026-03-08 · source-backed
537 test cases across 8 categories testing 6 commercial products. Composite scores range from ~39 to ~98. Critical finding: providers catching >95% of prompt injections miss most unauthorized tool calls — tool abuse detection is universally weak. Provenance verification is nearly absent. First empirical comparison of the agent security market. GitHub
Each link below shares sources, entities, or timing with this story.
Shared entity: AgentShield / Same source / Shared topic / Earlier coverage / Tension
Both cover AgentShield; cite the same source (GitHub); overlapping topics (abuse, agent, agentshield, security, tool).
Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Critical, GitHub; reported by the same outlet (github.com); overlapping topics (agent, benchmark, critical).
Shared entity: GitHub / Same source domain / Shared topic / Earlier coverage / Tension
Both cover GitHub; reported by the same outlet (github.com); overlapping topics (agent, category, critical, security).
Shared entities / Same source domain / Shared topic / What happened next
Both cover Critical, GitHub; reported by the same outlet (github.com); overlapping topics (agent, critical).
Both cover AgentShield, GitHub; reported by the same outlet (github.com); overlapping topics (agent, agentshield).
Both cover AgentShield, GitHub; reported by the same outlet (github.com); overlapping topics (agent, agentshield).
Shared entity: GitHub / Same source domain / Shared topic / What happened next / Tension
Both cover GitHub; reported by the same outlet (github.com); overlapping topics (agent, call, category).
Shared entity: GitHub / Same source domain / Shared topic / What happened next
Both cover GitHub; reported by the same outlet (github.com); overlapping topics (agent, call, category, commercial).