Fetching from the wire…
Vibe Coding2026-07-25 · source-backed
A llama.cpp fork by fewtarius aimed at AMD APUs, iGPUs and handhelds adds a persistent SSD-backed KV cache with hot/warm/cold tiering. On an Ayaneo Flip KB running Qwen3.6-35B over a 15,700-token prompt, cold TTFT was 143.1 seconds versus 0.99 seconds warm, a 144.5x speedup. Also adds system-prompt caching across conversations, hybrid MoE checkpoint restoration, MoE expert-activation tracking, per-user concurrency caps, and Vulkan tuning for Strix Halo. MIT, 101 stars. This targets the exact pathology of local agent loops: re-evaluating thousands of tokens of system prompt and tool definitions on hardware generating 5-20 tok/s.
Each link below shares sources, entities, or timing with this story.
June local-inference benchmarks across the 128GB class put NVIDIA's DGX Spark (~$4k), AMD's Strix Halo / Ryzen AI Max+ 395 (~$2 to 3k), and the M5 Max 128GB (~$5k) head to head (Hardware Corner). Prompt processing favors CUDA hard. But token generation lands at a surprisingly...
Created August 24, it holds a trendingScore of 3,967 against second-place GLM-5.3-Flash at 1,376 (Hugging Face). The near-1:1 like-to-download ratio means almost everyone bookmarking it hasn't pulled weights, and the unsloth GGUF conversion at 4,354 downloads is absorbing comp...
PR 27742 adds gated delta net layers, 512-expert MoE with top-10 selection, hyper-connections, query-key sparse attention and the per-layer n-gram embeddings as a mmap table that can sit in RAM or on disk (GitHub). Reported 55 tok/s on 4x3090 with the Q4 GGUF, and one commente...
The Apple Silicon inference server at 18,843 stars shipped 0.6.0 yesterday with experimental distributed serving using tensor or pipeline parallelism, capability-aware planning and memory guards (GitHub). A 225 GB MiniMax-M3 checkpoint loaded across a 128 GB and a 256 GB Mac....
Qwen 3.5 is the first major model pretrained specifically for agentic multimodal workflows from the first training stage, not fine-tuned after the fact. 397B total / 17B active parameters (MoE architecture), 256K context window, 201 languages. Ships with Qwen Code (terminal ag...
A 2.4-trillion-parameter MoE multimodal model with 1M-token context, aimed at long-horizon autonomous software work. Alibaba reports 86.1 against 83.2 for GPT-5.6 Sol Max and 85.0 for Fable 5, the first credible claim that a Chinese lab leads on GUI-driving agentic benchmarks....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.