Markets
jevals Replaces LLM Judges With Decision Models and Publishes the Bill: $2.60 Per Sample Down to $0.03
Openlayer released jevals, a week-old alpha Python library that runs every eval for a trace in one request against a decision model instead of a frontier LLM. Its published comparison puts Ragas at 6 to 11 requests per sample, $2.60 per sample and 22 to 35 seconds for 20 samples, against jevals on Jev at one request, $0.03 per sample and 0.8 seconds, with p50 latency of 244ms. Local backends (Kev, Laya) drop the cost to zero at 1 to 6 seconds per sample. The eval-platform category has spent two years billing for frontier tokens it may no longer need to spend.
Source
↳ Follow the thread