Skills
Bespoke Labs shipped Nimble with its own eval against Jev: 90.12% agreement versus Jev's 93.21%
Nimble takes text plus a schema of choice, boolean or score questions and returns typed answers with per-option probabilities, and cannot write free text at all. Its contribution over the other open typed-decision projects is the training recipe and an honest scoreboard: contrastive data curation builds paired examples differing in one critical fact so the correct answer flips, yielding 2,676 curated examples across 10 subject categories. On 324 held-out examples Bespoke-Nimble-9B agreed with reference labels 90.12% of the time against 66.36% for the Qwen3.5-9B base and 93.21% for TypeSafe's Jev, with median 106ms per example on an H100. Repo created 18 September, 1,567 stars, ~18GB unquantized.
↳ Follow the thread