Qwen3.8-27B at FP8 With xhigh Reasoning Scored 29/30 on AIME 2026, Matching BF16 and Beating It at Medium Effort
r/LocalLLaMA·medium signal
A practitioner benchmarked Qwen3.8-27B on the MathArena/aime_2026 dataset across BF16 and FP8 weights at medium and xhigh reasoning effort. FP8 at xhigh beat BF16 at medium and equalled BF16 at xhigh while running faster, which is the practically useful result: the quantization costs nothing on this task, the reasoning budget is what matters. On problem 7 both BF16 and FP8 at xhigh exhausted the generation budget without emitting an answer, so that miss is a context-limit failure rather than a wrong answer.