Fetching from the wire…
Vibe Coding2026-08-28 · source-backed
A user moved from UD-Q3_K_XL at 140k context to UD-IQ3_XXS and cleared 200k on a 16GB eGPU over Thunderbolt 4, with KV cache at q5_1 and llama.cpp built with DGGML_CUDA_FA_ALL_QUANTS=ON (r/LocalLLaMA). Prompt processing fell from 700-800 tok/s to 400. A commenter on an RTX 5080 with KV at Q4_0 and vision forced to CPU via --no-mmproj-offload reported 1750 PP/s and 85 TG/s at 132k context, which is the better-balanced configuration of the two.
Each link below shares sources, entities, or timing with this story.
Thunderbolt supports OpenAI / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Thunderbolt supports OpenAI); both cover LocalLLaMA, Qwen3, RTX; reported by the same outlet (reddit.com).
Thunderbolt supports Anthropic / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Thunderbolt supports Anthropic); both cover LocalLLaMA, Qwen3, RTX; reported by the same outlet (reddit.com).
Shared entities / Same source domain / Shared topic / Earlier coverage / Downstream implication
Both cover CPU, LocalLLaMA, Qwen3, VRAM; reported by the same outlet (reddit.com); overlapping topics (context, over).
Thunderbolt supports Ollama / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Thunderbolt supports Ollama); both cover LocalLLaMA, Qwen3; reported by the same outlet (reddit.com).
Thunderbolt supports OpenAI / Shared entity: LocalLLaMA / Earlier coverage
Linked by a graph relationship (Thunderbolt supports OpenAI); both cover LocalLLaMA; earlier LocalLLaMA coverage from 2026-08-12.
Linked by a graph relationship (Thunderbolt supports OpenAI); both cover LocalLLaMA; earlier LocalLLaMA coverage from 2026-07-25.
Linked by a graph relationship (Thunderbolt supports OpenAI); both cover LocalLLaMA; earlier LocalLLaMA coverage from 2026-05-01.
Thunderbolt supports Ollama / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Thunderbolt supports Ollama); both cover LocalLLaMA, RTX, VRAM; reported by the same outlet (reddit.com).