Fetching from the wire…
Infra2026-08-20 · source-backed
Sharding by layer across ordinary Intel AI PCs on a normal network, pre-compiling each stage into an OpenVINO graph. The load-bearing detail: naive per-stage export runs well below monolithic inference because it misses OpenVINO's IndirectKVCache fusion. Injecting a beam_idx Gather into each shard triggers the fusion and restores parity. A two-node Llama 3.1 8B INT4 pipeline serves two concurrent users at 1.79x single-user throughput, and a four-node Lunar Lake deployment serves 70B at interactive speed with output token-for-token identical to non-speculative decoding. (arXiv 2608.19147)
Each link below shares sources, entities, or timing with this story.
Meta released Llama / Shared entity: Llama / Earlier coverage
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-08-16.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-07-19.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-06-15.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-05-02.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-05-01.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-04-02.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-03-12.
Meta released Llama / Shared topic / Tension
Linked by a graph relationship (Meta released Llama); overlapping topics (deployment, detail); pushes against this story (but).