Fetching from the wire…
Tools2026-07-30 · source-backed
A Show HN at 87 points ships local transcription and system-wide push-to-talk for Apple Silicon Macs, 119 stars, Apache-2.0, v0.2.0-beta.1. Runs Qwen3-ASR through mlx-qwen3-asr with a size choice: 1.7B for accuracy needing 3.4 GB unified memory, 0.6B for speed needing 1.2 GB. Drag-and-drop transcription with automatic language detection, right-Command push-to-talk, searchable local transcript history, SRT export with word timestamps.
Each link below shares sources, entities, or timing with this story.
Released August 26, it brings MLX support for Qwen3.8 Flash Next, adds structured output to mlxrunner, and stops Metal GPU timeouts when loading models from slow storage (GitHub). Structured output on the MLX path is the practical unlock: Apple Silicon local inference can now...
oMLX runs local LLMs on Apple Silicon with a two-tier KV cache: hot cache in RAM, cold cache on SSD in safetensors format. When a previous context prefix recurs, blocks restore from disk instead of recomputing. Time-to-first-token drops from 30-90s to 1-3s on long contexts. 13...
Google released Gemma 4 on April 2 with four model variants: E2B, E4B, 26B MoE, and 31B Dense. The license change is the first thing worth noting. Every previous Gemma had restrictions that made lawyers nervous. Gemma 4 is Apache 2.0. Full stop. Use it in any product, any way...
v0.1.803-beta, released August 25 with 170+ PRs, lets long local chats continue past a model's context limit by rolling older turns into fresh context epochs rather than permanently trimming, with evicted conversations still searchable (GitHub). It also fixes MLX and Mac runti...
The Apple Silicon inference server at 18,843 stars shipped 0.6.0 yesterday with experimental distributed serving using tensor or pipeline parallelism, capability-aware planning and memory guards (GitHub). A 225 GB MiniMax-M3 checkpoint loaded across a 128 GB and a 256 GB Mac....
elie222/rakazo appeared Aug 13, Apache-2.0, TypeScript, explicitly bring-your-own model and sandbox (tested against Docker, E2B, Daytona) with the Pi runtime underneath and OpenRouter, Codex, Copilot, or SuperGrok device-code sign-in instead of a mandatory API key. Each bot ge...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.