Fetching from the wire…
Public story · 2026-08-23 · high
The 2.7x tool-calling gain is self-reported from one benchmark the builder ran alone, with free weights published through Q4_K_M.
Why now: Posted to Reddit's LocalLLaMA forum as of the August 23 roundup, among builders still hunting for what actually fits in 16GB of VRAM.
One builder fine-tuned Gemma 4 12B for tool calling and CLI work, then reported a 2.7x improvement on tool calling, per a post on Reddit's LocalLLaMA forum.
The target was 16GB of VRAM, the point past which nothing bigger than a 12B model runs comfortably for coding work, per the post. Anyone running local models near that VRAM ceiling has to pick between a bigger, slower model and a smaller one built for the job. Free weights let them test this one for themselves.
The builder published the model in formats from fp16 down to Q4_K_M under the name TheOneWhoWill/Coding-Monkey-Gemma-GGUF. They say Q6 and higher quantization performs noticeably better for anyone with the VRAM to spare.
Tool calls attempted rose 15.7%, which the builder reads as less time lost in reasoning. In a test against GitHub Copilot on a fresh create-next-app project, the fine-tune correctly sequenced the npm installs that followed.
The 2.7x figure comes from one person's benchmark on their own model, with no outside reviewer, per the post. The post doesn't name a broader test suite, so treat the number as a ceiling, not an average. Whether the Q6-and-above versions hold the same gain, or whether it shrinks as quantization drops, isn't in the post yet.
Each link below shares sources, entities, or timing with this story.
Claude Code competes with GitHub Copilot / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Claude Code competes with GitHub Copilot); both cover GGUF, LocalLLaMA, Weights; reported by the same outlet (reddit.com).
GitHub Copilot uses Claude Sonnet / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (GitHub Copilot uses Claude Sonnet); both cover Gemma, LocalLLaMA, VRAM; reported by the same outlet (reddit.com).
GitHub Copilot supports OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GitHub Copilot supports OpenAI); both cover GGUF, LocalLLaMA; reported by the same outlet (reddit.com).
Claude Code competes with GitHub Copilot / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Claude Code competes with GitHub Copilot); both cover GGUF, LocalLLaMA; reported by the same outlet (reddit.com).
Gemma built by Google / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Gemma built by Google); both cover CLI, Gemma, LocalLLaMA; reported by the same outlet (reddit.com).
Ollama supports Gemma / Shared entities / Earlier coverage
Linked by a graph relationship (Ollama supports Gemma); both cover Gemma, GGUF, VRAM; earlier Gemma coverage from 2026-06-07.
Ollama supports Gemma / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Ollama supports Gemma); both cover Gemma, LocalLLaMA; reported by the same outlet (reddit.com).
Claude Code competes with GitHub Copilot / Shared entity: Tested / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code competes with GitHub Copilot); both cover Tested; overlapping topics (against, tool).