Agents
Code Review Agent Benchmark (c-CRAB): State-of-the-Art Agents Including Devin, Claude Code, and Codex Solve Only ~40% of Review Tasks
Researchers from the National University of Singapore published c-CRAB, a benchmark evaluating AI agents on real-world code review tasks derived from actual human pull request reviews with auto-generated quality-gate tests. PR-agent, Devin, Claude Code, and Codex collectively solve only about 40% of c-CRAB tasks, revealing a hidden trade-off between issue resolution and spurious findings that constrains effective agent design.
Source
↳ Follow the thread