CROSS-CATEGORY: Small Typed Decision Models Showed Up in Evals, Moderation and Local Inference All in One Weekend
Three separate projects hit Hacker News within 48 hours attacking the same assumption, that a frontier model must be the thing making a yes/no call: jevals for agent evaluation gating, Kev as an open-weights decision-model family on Qwen3.5 (0.8B to 9B, 0.852 accuracy against TypeSafe Jev's 0.857), and a laya-coreml setup running a multilingual decision model fully offline on a Mac M4 at a 560MB physical footprint. Arize's September 18 analysis of Jev put the tradeoff at 68% accuracy for $0.0004 and 0.4 seconds per case versus Opus 5 at 73% for $0.18 and 38 seconds, and noted what you lose is the explanation. Any SaaS whose cost line is frontier-model classification should be re-pricing this week.
Source
↳ Follow the thread