Reddit
Ling 3.0 Tiny — 8B Total, 1.3B Active — Is Doing 36 tok/s on a 4GB-VRAM Card and 17–20 tok/s on Bare CPU, With Clean Tool Calling
A 179-upvote, 80-comment r/LocalLLaMA thread turned into an impromptu multi-hardware benchmark. OP reports 36 tok/s on a 4GB card where Qwen3.5-9B manages 5 tok/s at comparable quality; one commenter ran the Q6 quant CPU-only on an i7-13700K at ~17–20 tok/s; another ran a sub-4GB Q3 test through 17 tool calls with zero errors at 1.7k tok/s prompt processing and 120 tok/s output with 70k context loaded, and a third got 11–15 tok/s on an old i5 with 8GB single-channel DDR3 at 262,144 context. The comments also flag the limits — 120k-token documents failed, and Portuguese counting breaks — but the practical signal is that a competent tool-calling subagent now fits inside 4GB.
Source
↳ Follow the thread