A production run puts real numbers on the cheap-judge pattern: 9,081 matches, 32 cents, 13 minutes
paddo.dev published on 2026-09-19 a run of TypeSafe's Jev over Pricogni's backlog of 9,081 low-confidence product matches, using 150 lines of code, one bernoulli query for "do these match" and one choice query for the type of difference, with competitor descriptions capped at 1,500 characters. Total cost was $0.32 and wall time 13 minutes 22 seconds at 6 concurrent; the verdicts split 4,443 refuted (49%), 1,952 confirmed (21%), 2,686 abstained (30%), and a manual read of 50 verdicts found 48 defensible, with the one human-model disagreement resolving in the model's favor. The abstain rate is the number to copy: a third of the queue routed itself back to a human without anyone writing a confidence threshold.
Source
↳ Follow the thread