Fetching from the wire…
OSS2026-07-28 · source-backed
+3,059 this week. Bundles Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo (350M), HumeAI TADA and Kokoro into one MIT-licensed local voice studio, with zero-shot cloning from audio samples plus 50+ preset voices across 23 languages, combining TTS output with dictation input. Runs on MLX on Apple Silicon, PyTorch on CUDA/ROCm, or CPU, with AMD ROCm and Intel Arc supported. An ElevenLabs and Whisper replacement in one app. Arriving alongside huggingface/speech-to-speech (+177 stars) and moeru-ai/airi's WebGPU realtime voice, three projects converging on voice as the agent supervision interface in one week. (GitHub)
Each link below shares sources, entities, or timing with this story.
The repo reached 9,740 stars with roughly double the next-fastest Python project's daily gain. It exposes a VAD → STT → LLM → TTS pipeline behind an OpenAI Realtime-compatible WebSocket API with every stage swappable: Parakeet TDT as default STT (Whisper, Paraformer alternativ...
Blaizzy/nativ (1,163 stars, Swift, MIT, macOS 26+) comes from the mlx-vlm author and bundles that server into a SwiftUI app that discovers MLX models already in your HF cache. It exposes OpenAI-compatible chat, Responses, image, audio and model endpoints plus Anthropic Message...
AlexsJones/llmfit released v1.1.10 today, adding RamaLama runtime discovery to its MCP server, the Qwen3.8 model family and MiniMax M3 vision capability exposure (GitHub). It also merged 32 MLX benchmark results on an Apple M4 Pro, the project's first MLX entries, giving an ap...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
An MIT-licensed, local-first AI workspace published around May 31 that reached roughly 54,000 stars and 6,300 forks by June 5. It bundles chat over local and remote models (vLLM, llama.cpp, Ollama, OpenRouter, OpenAI), autonomous agents with bash/files/web/memory tools plus MC...
Cross-platform desktop STT built with Tauri. Fully offline using Whisper and Parakeet models. GPU-accelerated on CUDA, or CPU-only via Parakeet V3. Designed to be *"the most forkable speech-to-text app."* Competes with paid tools like Wispr Flow with zero cost and full privacy...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.