Hacker NewsBitNet 100B: 1-Bit Model Runs on Commodity CPUs — No GPU RequiredGitHub·high signalXBlueskyLinkedInCopy linkMicrosoft releases 100B parameter 1-bit model running on CPUs. 281pts HN. Paradigm shift for local inference without NVIDIA hardware.SourceSource pageGitHub↳ Follow the threadStack layer / Threat patterngoose v1.47.0 is mostly a resource-bounding security pass, and quietly stops paying the prompt-cache write premium on one-shot fast-model callsGitHubStack layer / Threat patternSimon Willison ships llm-openrouter 0.7 with three server-side tools: Shell, WebFetch and WebSearchGitHub (simonw/llm-openrouter)Policy dependency / Stack layerFreeToken Serves 290B+ Parameter MoE Models on RTX 30/40/50 Consumer Cards Using CPU-GPU Co-ExecutionGitHub / Hacker NewsStack layer / ContrastApache Maka enters incubation as a local-first agent workspace where the event log is the runtimeGitHub TrendingStack layer / ContrastAPEX publishes a verification-first LLM inference tile in RTL with the KV-cache codec inside the datapath, 0.56 tok/s measured on FPGAGitHubStack layer / Update threadoMLX 0.6.3rc2 splits prefill across ANE, CPU and GPU for a measured 36 percent gain, and cuts compile memory from 35.8 GB to 4.7 GBGitHubPolicy dependency / Stack layerVendo (YC S26) Open-Sourced a Layer That Lets Your Customers Build Features on Top of Your SaaS Without Filing a TicketGitHub / Y Combinator (Launch HN, Aug 20, 2026)Stack layer / ContrastKaku Forks WezTerm Into a Terminal Whose Defaults Assume You're Running Coding AgentsGitHub