RedditNVFP4 Support Coming to llama.cpp for Local InferenceGitHub·high signalXBlueskyLinkedInCopy linkTrue NVFP4 support arriving via GitHub PR 19769. Unlocks larger MoE models on 48GB or less memory at native FP4.SourceSource pageGitHub↳ Follow the threadShared entity / Recurring source clusterA llama.cpp bug had silently disabled Metal norm/MUL fusion, costing about 5% token-generation speedGitHubStack layer / Threat patternqwen-code 0.23.2 turns the CLI into a remotely accessible shell with one command and a QR pairing codeGitHubStack layer / ContrastEdge0 runs a 35B MoE on Apple Silicon in 2.9 GB of active memory by streaming experts off SSDGitHubPolicy dependency / Stack layerCROSS-CATEGORY: Three Independent Agent-Action Gates Shipped in 48 Hours, All Judging the Command Against Stated IntentProduct Hunt, github.com/AGGIB/Stroq and rewarelabs.com (three independent sources; the 72% figure is Reware's own)Stack layer / Contrastmesh-llm v0.76.0 makes the KV prefix cache survive eviction, restarts and cold nodesGitHubStack layer / Update threadcrewAI had gpt-4o-mini's context window recorded as 200,000 tokens instead of 128,000GitHubStack layer / Follow-up threadColibrì runs 744B to 2.8T MoE models on consumer hardware in pure C by streaming experts off diskGitHub TrendingStack layer / Update threadCline Desktop 0.0.25 lets Claude Code and Codex CLI providers start sessions with no API key, and caps Codex models at real backend budgetsGitHub Releases