Skills
You need 20–50 tasks, not a benchmark suite, to qualify a new model — and DeepEval 4.0 now runs evals against Claude Code and Codex locally
Three frontier models shipped in a single week of July 2026, and teams with a standing harness had a routing decision in hours; Anthropic's own agent-eval guidance says 20–50 tasks drawn from your real usage and real failures is enough to detect issues. DeepEval 4.0 added a local evaluation harness aimed specifically at coding agents like Claude Code and Codex (~17k stars, 8M+ monthly PyPI downloads), and Promptfoo tagged v0.121.19 on July 14 (MIT, 23.3k stars) now under OpenAI while staying open source. The actionable move is to mine your own failure log for 20 tasks today rather than waiting to build a proper benchmark.
Source
↳ Follow the thread