Research
FuzzingBrain-Bench Scores Open-Ended Bug Discovery; Claude Opus 4.8 Crashes 60 of 77 Targets but Scores 196/579
Sheng et al. argue that predefined-target proof-of-concept benchmarks discard valid crashes an LLM actually found, so their benchmark hands the model an open-source project plus a sanitizer-instrumented harness in a self-contained Docker image and scores distinct crash signatures, capped per challenge and weighted by difficulty. V1 has 77 challenges from 43 projects (36 C, 32 C++, 9 Java/JVM). Claude Opus 4.8 leads with crashes in 60 of 77 challenges for 196 of a possible 579 points; none of Haiku 4.5, Sonnet 4.6, or Opus 4.8 crashed 13 of the challenges. Corpus and harnesses are public on GitHub.
↳ Follow the thread