Voices
Bespoke Nimble closes most of the Jev gap with a LoRA fine-tune of Qwen3.5-9B at 100ms on an H100
@madiator's clone uses contrastive data curation to fine-tune Qwen3.5-9B, lifting the base model from 66% to 90% against Jev's 93% on the same task, at 100ms latency on an H100. That is the most informative data point in the whole clone wave: it suggests most of Jev's accuracy is reachable with a public base model and a curated dataset, and that the remaining three points plus the pricing come from the architecture and the serving stack. For a builder, it means a self-hosted decision model is a weekend of data work, not a research program.
↳ Follow the thread