Research
Half of 36 Inference Configurations Sit on the Pareto Frontier, and AWQ 4-bit Misses the Quality Floor by 5.9%
arXiv 2609.17863 (15 Sep 2026) measures 54 configurations of Qwen2.5-7B-Instruct on vLLM 0.12 across L4, A100 and H100, then calibrates a simulator reproducing them with cross-campaign drift under 1.5%. On the calibrated grid 18 of 36 configurations reach the cost/quality/latency Pareto frontier, and combined optimizations reach it more often than single ones (9 of 15 versus 9 of 21). Quality testing on 200 GSM8K questions reorders the winners: AWQ 4-bit cuts per-token latency to 0.34x baseline on L4 but loses 5.9% strict accuracy, narrowly missing a 95% quality floor, while flexible answer extraction recovers FP16 parity, suggesting the loss is formatting rather than arithmetic.
↳ Follow the thread