Fetching from the wire…
Models2026-08-27 · source-backed
Quesma ran the model across GPQA Diamond, IFBench and Terminal-Bench 2.1 (89 agentic coding tasks) on L40S, H100 and H200 via Modal. Q4_K_M at 17 GB matched BF16 at 55 GB within a point on all three. UD-Q2_K_XL at 10.7 GB held instruction-following but dropped Terminal-Bench from about 77% to about 72%. Both 1-bit quants at 6.2 GB fell to 15-20% on GPQA Diamond, which is around random chance, and degraded further on longer reasoning (Quesma). Q4_K_M remains the default answer and this is the cleanest evidence for it I've seen this month.
Each link below shares sources, entities, or timing with this story.
Shared entities / Shared topic / Earlier coverage / Tension
Both cover Bench, GPQA Diamond, Terminal; overlapping topics (agentic, diamond, terminal-bench); earlier Bench coverage from 2026-08-21.
Both cover Bench, Qwen3, Terminal; overlapping topics (agentic, benchmark, coding); earlier Bench coverage from 2026-04-21.
Both cover Bench, H100, Terminal; overlapping topics (agentic, coding); earlier Bench coverage from 2026-06-10.
Shared entities / Shared topic / Earlier coverage
Both cover Bench, GPQA Diamond, Terminal; overlapping topics (agentic, benchmark, coding); earlier Bench coverage from 2026-06-14.
Both cover Bench, Qwen3, Terminal; overlapping topics (benchmark, coding, terminal-bench); earlier Bench coverage from 2026-04-21.
Shared entities / Earlier coverage
Both cover Bench, GPQA Diamond, Qwen3, Terminal; earlier Bench coverage from 2026-08-20.
Both cover Bench, GPQA Diamond, Qwen3, Terminal; earlier Bench coverage from 2026-08-16.
Shared entities / Shared topic
Both cover Bench, H200, Terminal; overlapping topics (agentic, benchmark).