Seed your eval golden set from 20–50 real production failures, not a curated wishlist
Digital Applied·medium signal
Start an agent eval suite with 20–50 actual production failures rather than hand-picked happy paths, applying the rule that 'two domain experts must independently reach the same pass/fail verdict,' then scale to 100+ for judge calibration and 200–500 for production gold sets. Early-stage agents show large effect sizes per change, so even a tiny failure-derived set gives real signal. This inverts the usual instinct to build big synthetic test suites that never reflect how the agent actually breaks.