Reddit
A 100-task SWE-bench Verified sweep shows chat-template choice moves Qwen3.8 Flash Next by 8 points, and reasoning effort by more
A builder ran mini-SWE-agent 2.4.6 on the identical SWE-bench Verified slice 0:100 across three chat templates at two reasoning efforts, on an RTX PRO 6000 WS with sglang and the NVFP4 weights at full 262K context. Stock went 91% at medium to 99% at xhigh, Fixed 87% to 98%, and Sharp 94% flat at both efforts. The tradeoff is time and tokens: stock xhigh cost +143.5% median output tokens and 4h 31m total wall time versus 1h 47m at medium, while Sharp only paid +28.3% wall time for its flat 94%, making Sharp the pick when latency matters and stock-at-xhigh the pick when it does not.
↳ Follow the thread