Fetching from the wire…
Public story · 2026-02-23 · source-backed
Streams model layers through GPU memory via PCIe with an optional NVMe-to-GPU bypass. Runs Llama 3.1 70B on a single RTX 3090 (24GB VRAM). Warning: NVMe bypass mode can brick your drive — never use on your boot drive. Non-bypass mode alone is a serious consumer-hardware inference option.
Clone it: github.com/xaskasdf/ntransformer | C++, CUDA
Each link below shares sources, entities, or timing with this story.
Meta released Llama / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Meta released Llama); both cover CUDA, RTX; overlapping topics (inference, model).
Shared entities / Same source domain / Shared topic / What happened next / Downstream implication
Both cover CUDA, GPU, Llama, VRAM; reported by the same outlet (github.com); overlapping topics (inference, llama, model).
Meta released Llama / Shared entity: Llama / Shared topic / What happened next
Linked by a graph relationship (Meta released Llama); both cover Llama; overlapping topics (llama, model).
Meta released Llama / Shared entity: Show HN / Same source domain / What happened next
Linked by a graph relationship (Meta released Llama); both cover Show HN; reported by the same outlet (github.com).
Meta released Llama / Shared entity: NVMe / Same source domain / What happened next
Linked by a graph relationship (Meta released Llama); both cover NVMe; reported by the same outlet (github.com).
Meta released Llama / Shared entity: Llama / Shared topic / What happened next
Linked by a graph relationship (Meta released Llama); both cover Llama; overlapping topics (llama, model).
Linked by a graph relationship (Meta released Llama); both cover Llama; overlapping topics (llama, model).
Meta released Llama / Shared entity: GPU / What happened next / Tension
Linked by a graph relationship (Meta released Llama); both cover GPU; picks up the GPU thread on 2026-07-25.