Fetching from the wire…
Public story · 2026-07-27 · high
The build packs 2.5TB of VRAM against a 594GB model file, but nobody's posted tokens per second yet.
Why now: The post surfaced in the July 27 LocalLLaMA coverage as the first reported consumer-GPU run of Kimi K3.
A Reddit user got Kimi K3 running across 80 RTX 5090 GPUs, linked over standard 25 gigabit Ethernet, according to a post on r/LocalLLaMA. That matters because NVLink and InfiniBand are the expensive part of any GPU cluster built for models this size. If Ethernet holds up, local-LLM builders get a cheaper path to running huge mixture-of-experts models without buying specialized interconnect hardware.
The numbers back up the scale. Eighty 5090s add up to roughly 2.5TB of aggregate VRAM, against a ~594GB MXFP4 weight file for K3. The leftover room covers activation memory and KV cache overhead, margin most local builds don't have.
Kimi K3 is a mixture-of-experts model. Expert-parallel MoE over commodity Ethernet is exactly the setup vLLM's engineering blog has warned about. Network bandwidth, not compute, ends up capping how fast any single user gets tokens back.
What's missing is a tokens-per-second number. The post is a single, unaudited report that nobody else has replicated yet.
Getting a 594GB model to load across 80 cards over Ethernet is real. Whether it's usable depends on throughput, and throughput is the one number missing from the thread. An 80-GPU Ethernet cluster and an 80-GPU Ethernet paperweight look identical without it.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Earlier coverage
Both cover LocalLLaMA, MoE, RTX, Someone; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-04-05.
Both cover LocalLLaMA, MoE, RTX, VRAM; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-05-22.
Shared entities / Same source domain / Earlier coverage / Tension
Both cover LocalLLaMA, MoE, RTX; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-04-23.
Shared entities / Same source domain / Shared topic
Both cover LocalLLaMA, MXFP4, Still; reported by the same outlet (reddit.com); overlapping topics (activation, over).
Shared entities / Same source domain / Earlier coverage
Both cover LocalLLaMA, MoE, Someone; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-04-06.
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover LocalLLaMA, MoE; reported by the same outlet (reddit.com); overlapping topics (deployment, over).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover LocalLLaMA, MoE; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-04-21.
Shared entities / Same source domain / Earlier coverage / Downstream implication
Both cover LocalLLaMA, VRAM; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-04-10.