Fetching from the wire…
Tools2026-08-27 · source-backed
PR #26622 pushes a user-specified number of FFN sublayers to CPU while attention stays on the GPU, mirroring --n-cpu-moe (GitHub). The practical note from the r/LocalLLaMA thread is that the previous route was regex matching in -ot style, and for a model you plan to run for months a single flag beats both that and --fit. This is the concrete lever for running dense models on low-VRAM cards.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Shared topic / Earlier coverage / Downstream implication
Both cover CPU, GPU, LocalLLaMA, VRAM; reported by the same outlet (github.com); overlapping topics (card, llama, model).
Shared entities / Shared topic / Earlier coverage / Tension
Both cover GitHub, GPU, LocalLLaMA, MoE; overlapping topics (already, dense, model); earlier GitHub coverage from 2026-04-23.
Shared entities / Same source domain / Earlier coverage
Both cover GPU, LocalLLaMA, MoE, VRAM; reported by the same outlet (github.com); earlier GPU coverage from 2026-05-22.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover GitHub, LocalLLaMA, MoE; reported by the same outlet (github.com); overlapping topics (llama, model).
Both cover CPU, GPU, LocalLLaMA; reported by the same outlet (github.com); overlapping topics (localllama, model).
Both cover GPU, LocalLLaMA, MoE; reported by the same outlet (github.com); overlapping topics (dense, model).
Both cover CPU, GitHub, GPU; reported by the same outlet (github.com); overlapping topics (llama, model).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover CPU, GitHub, GPU; reported by the same outlet (github.com); earlier CPU coverage from 2026-02-22.