Reddit
llama.cpp Merges a DeepSeek V4 Flash 0731 Chat-Template Fix That Practitioners Say Ends Tool-Calling Loops
PR #26398 by tarruda, opened August 1, 2026, fixes the existing DeepSeek V4 preview Jinja template to match official encoder behavior and adds a dedicated 0731-variant template with distinct prompts per reasoning level, handling max-effort reasoning, structured output, and the `drop_thinking` default of True. The r/LocalLLaMA poster who flagged it (95 upvotes) reported looping and poor tool-calling behavior the previous day that stopped entirely after the fix landed. This is the concrete unblock for anyone who pulled V4 Flash GGUFs after last week's release and hit degraded agentic behavior — the model was fine, the chat template was not.
↳ Follow the thread