Fetching from the wire…
Models2026-08-29 · source-backed
TerminalBytes measured about 14 tokens/sec at Q4_K_M (17GB file) on an M3 Ultra Mac Studio with 256GB, against roughly 28.6 tok/s for Qwen3.6 27B on the same machine. The reason it stays usable: the newer model answers in about a third fewer tokens. A 1-bit quant at 6.7GB reaches 27 tok/s and fits in 16GB. You need a llama.cpp build from the last few weeks or it fails with "unknown model architecture: qwen35." Tokens per second is now a misleading benchmark on its own, which is going to take the community a while to internalize. (TerminalBytes)
Each link below shares sources, entities, or timing with this story.
Shared entities / Shared topic / Earlier coverage
Both cover Qwen3, Tokens; overlapping topics (benchmark, model, token); earlier Qwen3 coverage from 2026-04-03.
Shared entity: Qwen3 / Shared topic / Earlier coverage / Tension
Both cover Qwen3; overlapping topics (architecture, benchmark, model, same); earlier Qwen3 coverage from 2026-04-21.
Both cover Qwen3; overlapping topics (against, benchmark, token); earlier Qwen3 coverage from 2026-08-25.
Shared entity: Qwen3 / Shared topic / Earlier coverage / Downstream implication
Both cover Qwen3; overlapping topics (architecture, model, same); earlier Qwen3 coverage from 2026-08-19.
Shared entity: Qwen3 / Shared topic / Earlier coverage / Tension
Both cover Qwen3; overlapping topics (against, benchmark, model); earlier Qwen3 coverage from 2026-08-18.
Both cover Qwen3; overlapping topics (against, benchmark, model); earlier Qwen3 coverage from 2026-08-17.
Shared entity: Tokens / Shared topic / Earlier coverage / Tension
Both cover Tokens; overlapping topics (fewer, model, token); earlier Tokens coverage from 2026-05-01.
Shared entity: Qwen3 / Shared topic / Earlier coverage / Tension
Both cover Qwen3; overlapping topics (community, model, qwen3); earlier Qwen3 coverage from 2026-03-20.