Research
Music Tokenization Beats Model Scale: A 0.8B Performance-Resolution Model Outperforms a 27B Beat-Grid Model on Frechet Music Distance
Holding Qwen3.5 backbone (0.8B–27B), data, budget, and decoding fixed and swapping only the representation across seven tokenizations, the authors find representation is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves FMD, while switching representation halves it: their released PMT stream (10 ms timing, per-note velocity, multi-track texture, 609 symbols) reaches FMD 159 at 0.8B versus 272–286 for beat grids. They also release a 6.25M-caption corpus and an imprinting diagnostic showing published text-to-MIDI systems reproduce their training distribution nearly independent of the caption (72% vs 71% chord-time on disjoint domains).
↳ Follow the thread