Fetching from the wire…
Models2026-07-25 · source-backed
owensong released Inflect-Nano-v2 (3,966,721 deployable params) and Inflect-Micro-v2 (9,356,513) under Apache-2.0, VITS-family end-to-end text-to-waveform with 128 latent channels, 3 encoder layers, 4 flow coupling blocks, 24 kHz mono. Nano-v2 runs at 0.0933 RTF (10.72x real-time) on four CPU threads; Micro-v2 at 0.1593 RTF (6.28x). The trade-off is deliberate: English only, one fixed male voice, and the training-corpus pipeline stays private, so it's open-weight not open-source. 410 upvotes on r/LocalLLaMA. Usable on-device TTS now fits in under 10MB of parameters, which puts voice output inside embedded budgets that ruled it out a year ago.
Each link below shares sources, entities, or timing with this story.
owensong's VITS-family English TTS has 9,356,513 deployable parameters and a 37.53MB FP32 footprint producing 24kHz mono (Hugging Face). Reported: 66.2% preference in community blind listening tests, 4.395 UTMOS22 naturalness, 3.99% semantic error rate across multiple ASR eval...
The top trending HuggingFace paper (274 upvotes) introduces dots.tts, a 2B continuous autoregressive text-to-speech model hitting best average Seed-TTS-Eval (WER 0.94%/1.30% zh/en) with strong cloning and emotional range. CFG-aware MeanFlow distillation gives 85ms first-packet...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
PR #26062, "server: support MCP stdio," by ngxson, merged into ggml-org/llama.cpp on July 25 (r/LocalLLaMA). It landed alongside #26061 (vendored subprocess.h, merged July 24) and pwilkin's #26075 integration-and-tests PR. Until now, llama-server's web UI could only talk to MC...
Finally, a story about building something instead of worrying about something. Mistral released Voxtral TTS on March 26, an open-source text-to-speech model built on Ministral 3B. The numbers are striking: 90ms time-to-first-audio, 6x real-time factor (a 10-second clip generat...
Alibaba released Qwen3.5-0.8B, 2B, 4B, and 9B — all natively multimodal (text+image+video from same weights, no adapter), 262K context, Apache 2.0. The 9B beats last-gen Qwen3-30B across the board and outperforms GPT-5-Nano by 13 points on MMMU-Pro. Architecture uses Gated Del...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.