OSSxaskasdf/ntransformer — Llama 70B on Single RTX 3090 via NVMe-to-GPU BypassGitHub·high signalXBlueskyLinkedInCopy linkC++ CUDA inference engine streams model layers through GPU via PCIe. Optional NVMe-to-GPU bypass. Show HN 380 pts.SourceSource pageGitHub↳ Follow the threadPolicy dependency / Stack layerFreeToken Serves 290B+ Parameter MoE Models on RTX 30/40/50 Consumer Cards Using CPU-GPU Co-ExecutionGitHub / Hacker NewsStack layer / ContrastApache Maka enters incubation as a local-first agent workspace where the event log is the runtimeGitHub TrendingStack layer / ContrastAPEX publishes a verification-first LLM inference tile in RTL with the KV-cache codec inside the datapath, 0.56 tok/s measured on FPGAGitHubStack layer / Update threadoMLX 0.6.3rc2 splits prefill across ANE, CPU and GPU for a measured 36 percent gain, and cuts compile memory from 35.8 GB to 4.7 GBGitHubPolicy dependency / Stack layerVendo (YC S26) Open-Sourced a Layer That Lets Your Customers Build Features on Top of Your SaaS Without Filing a TicketGitHub / Y Combinator (Launch HN, Aug 20, 2026)Stack layer / Update threadDeepSeek Harness v0.1.1-rc.1 patches a Bubblewrap sandbox escape and adds a vision model to the adapterGitHubStack layer / ContrastLeanCTX Compresses Agent File Reads to ~13 Tokens on a Cache HitGitHubStack layer / Update threadai-data-extractor normalizes chat history from nine coding assistants into one JSONL, including reverse-engineered Windsurf and Trae schemasGitHub