Sources
JevBench v1.3.0 ranks 52 decision-model systems, and a Qwen3.5-4B open clone scores 73.1 against Jev's 74.4
Benchmark Heaven's JevBench scores 534 fixed decisions (220 of them hard) on four equally weighted axes: intelligence above chance, calibration, speed and cost. Jev 1.13.0 leads at 74.4 for $0.040. SemIf (formerly OpenJev, Qwen3.5-4B) follows at 73.1 for about $0.022, and the diffusion-Gemma djev comes third at 73.0. GPT-5.6 Luna on low effort posts the top intelligence score (95) but ranks 14th because it costs $0.242. The harness and public tasks are MIT-licensed at fstandhartinger/jevbench (created 09-19), and the results JSON is published with a sha256, so you can rerun it on your own routing or classification model before paying for Jev.
↳ Follow the thread