Research
A 24 GiB MacBook Serves 196K Input Tokens Locally, a 6.93x Jump Over the mlx-vlm Baseline
JustFit (arXiv 2609.17475, submitted 15 Sep 2026) is an MLX inference runtime combining compressed KV execution, component residency swapping, and state-preserving serving transitions. On a 24 GiB M4 Pro MacBook running Qwen3.8-27B MXFP4, three independent runs completed 196,608 input and 16,384 output tokens, lifting single-request context from the mlx-vlm baseline's 30,720 positions to 212,992. A 32K-input probe reached 19.11 tokens/s with a median peak process footprint of 16,374 MiB, and the runtime answered 29 of 30 AIME 2026 problems correctly.
↳ Follow the thread