Reddit
A hands-on Q8 comparison puts DeepSeek-V4-Flash-Vision at 40% slower token generation but 2x faster task completion than Qwen3.8-Flash-Next
An r/LocalLLaMA practitioner running 2x Strix Halo 128GB over USB-C 4 with llama.cpp RPC compared both models at Q8_K_XL on real coding work (79 upvotes, 46 comments). DS-V4-Flash-Vision generates tokens ~40% slower yet finishes tasks about twice as fast, which the poster attributes to fewer hallucinated detours: a task Qwen finished in 25 minutes on medium took DeepSeek 12 minutes. The sharpest number is a failure mode, not a win: Qwen3.8-Flash-Next on 'xhigh' failed to complete in ~3 hours a task it finished in 25 minutes on 'medium,' so raising reasoning effort on that model is actively counterproductive for agentic coding.
↳ Follow the thread