Fetching from the wire…
Infra2026-07-12 · source-backed
Posted July 11 (293 points on HN), it distributes inference across a network built on iroh, using hole-punching to route requests directly between nodes with no central server (iroh). Models get partitioned by layer ranges, layers 0-15 on one machine, 16-31 on the next, so several modest GPUs run a model none could hold alone. For home labs and small teams hitting VRAM ceilings, this is the most interesting local-inference idea I've seen in a while. Whether the network latency stays acceptable across the hops is the open question.
Each link below shares sources, entities, or timing with this story.
Shared entities / Shared topic / Earlier coverage / Downstream implication
Both cover GPUs, VRAM; overlapping topics (gpus, model); earlier GPUs coverage from 2026-04-10.
Shared entities / Shared topic / What happened next
Both cover GPUs, Whether; overlapping topics (gpus, model); picks up the GPUs thread on 2026-07-24.
Shared entity: Models / Shared topic / What happened next / Tension
Both cover Models; overlapping topics (directly, model); picks up the Models thread on 2026-07-31.
Shared entity: GPUs / Shared topic / Earlier coverage / Tension
Both cover GPUs; overlapping topics (could, gpus); earlier GPUs coverage from 2026-06-26.
Both cover GPUs; overlapping topics (gpus, model); earlier GPUs coverage from 2026-05-09.
Shared entities / What happened next
Both cover Models, Whether; picks up the Models thread on 2026-07-29.
Shared entity: GPUs / Shared topic / Earlier coverage
Both cover GPUs; overlapping topics (directly, gpus, hitting); earlier GPUs coverage from 2026-05-11.
Both cover GPUs; overlapping topics (consumer, gpus, layer); earlier GPUs coverage from 2026-03-18.