Research
A Fine-Tuned 4B Qwen in 2.6 GB Beats GPT-5.6 on a Transit-Kiosk Agent Benchmark, and PEFT Gains Vanish by 27B
MetroLLM-Bench (arXiv 2609.10016, 2026-09-09) is a 955-case benchmark testing models as the policy layer of a transit kiosk across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split, a PEFT-trained 4B Qwen 3.5 student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, in a 2.6 GB Q4_K_M footprint, while 9B and 27B students add nothing further. The PEFT gain over base shrinks monotonically from +7.03 points at 2B to -0.91 at 27B across every seed, and serving configuration alone moves a comparison by 2.7 points. Released at github.com/continker/metrollm-bench.
↳ Follow the thread