Fetching from the wire…
Agents2026-04-21 · source-backed
The Stanford HAI report shows OSWorld task success jumped from 12% to 66%, SWE-bench Verified climbed to near 100%, and cybersecurity agent solve rates hit 93%. But the same models that win Math Olympiad gold read analog clocks correctly only 50.1% of the time. 62% of enterprises cite security as the primary scaling blocker. The capability is there. The governance isn't.
Each link below shares sources, entities, or timing with this story.
Stanford HAI released AI Index / Shared entities / Same source / Shared topic / Earlier coverage
Linked by a graph relationship (Stanford HAI released AI Index); both cover SWE, Verified; cite the same source (Stanford HAI report).
Stanford HAI released AI Index / Shared entities / Same source / Earlier coverage
Linked by a graph relationship (Stanford HAI released AI Index); both cover Stanford HAI, SWE; cite the same source (Stanford HAI report).
Shared entities / Shared topic / Earlier coverage
Both cover OSWorld, SWE, Verified; overlapping topics (agent, task); earlier OSWorld coverage from 2026-04-12.
Both cover OSWorld, SWE, Verified; overlapping topics (agent, capability); earlier OSWorld coverage from 2026-03-06.
Shared entities / Earlier coverage / Tension
Both cover OSWorld, SWE, Verified; earlier OSWorld coverage from 2026-02-17; pushes against this story (vs).
Shared entities / What happened next
Both cover OSWorld, SWE, Verified; picks up the OSWorld thread on 2026-07-21.
Both cover OSWorld, SWE, Verified; picks up the OSWorld thread on 2026-06-08.
Shared entity: Stanford AI Index / Same source / Earlier coverage / Tension
Both cover Stanford AI Index; cite the same source (Stanford HAI report); earlier Stanford AI Index coverage from 2026-04-14.