Voices
Trajectory Labs Ran 720 Indirect Prompt Injection Attacks Against Claude Fable 5, Opus 5 and Sonnet 5 in Auto Mode — Zero Succeeded
Anthropic commissioned independent evaluator Trajectory Labs to test 72 indirect prompt injection scenarios, held out from Anthropic and each run 10 times, against models as of July 17, 2026. None of the 720 attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode. A clean sweep on a held-out third-party eval is a meaningfully stronger claim than internal red-teaming, though the scenario set is narrow enough that it doesn't cover the supply-chain vector Willison flags.
Source
↳ Follow the thread