Tools
Soup v0.73.2 Is a Post-Mortem on Its Own Eval Scorer: a Model Named the Right Tool 40/40 and Scored 0.225
MakazhanAlpamys/Soup — a YAML-driven fine-tuning tool whose pitch is layer streaming that trains an 8B model on a 4GB laptop GPU — shipped v0.73.2 on 2026-08-15 at 09:04 UTC fixing four defects in the regression half of `soup ship`. The headline case: the `mini_tool_call` suite was effectively ranking brace hygiene, so Llama-3.1-8B named the correct tool 40/40 yet scored 0.225 because it emitted three opening braces and two closing ones and the scorer rejected the bounded-scan result. Every defect was reproduced against shipped v0.73.1 before a line changed, using a stub emitting real model output shapes rather than a GPU — a reproducibility discipline more eval harnesses should adopt.
Source
↳ Follow the thread