Fetching from the wire…
Models2026-08-18 · source-backed
A 179-upvote r/LocalLLaMA thread turned into an impromptu multi-hardware benchmark. OP reports 36 tok/s on a 4GB card where Qwen3.5-9B manages 5 at comparable quality. One commenter ran Q6 CPU-only on an i7-13700K at ~17-20 tok/s. Another ran a sub-4GB Q3 through 17 tool calls with zero errors at 1.7k tok/s prompt processing and 120 tok/s output with 70k context loaded. A third got 11-15 tok/s on an old i5 with 8GB single-channel DDR3 at 262,144 context. Limits flagged in comments: 120k-token documents failed, Portuguese counting breaks. A competent tool-calling subagent now fits in 4GB.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / Earlier coverage / Downstream implication
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (card, context).
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (benchmark, card, comment).
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (active, context).
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (commenter, tool).
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (benchmark, clean).
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (card, tool).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-08-13.
Shared entity: LocalLLaMA / Same source domain / Shared topic / Earlier coverage
Both cover LocalLLaMA; reported by the same outlet (reddit.com); overlapping topics (active, benchmark, comment, context).