Fetching from the wire…
Models2026-09-15 · source-backed
Qwen3-TTS 1.7B and Qwen3-ASR 1.7B went to public beta September 14 with Coval placements: the ASR Fast endpoint ranks first on latency at 44ms p50 time-to-final-segment with a 3.6% WER for $0.12/hour, standard endpoint at $0.06 (Nari Labs). TTS Fast takes second on latency at 63ms p50 time-to-first-audio while ranking first on WER at 3.8%, $10 per million characters, standard at $5. Their comparison table puts AssemblyAI Universal 3.5 Pro at 3.75x the STT cost, Deepgram Nova 3 at 2.4x, ElevenLabs Eleven v3 at 5x and Cartesia Sonic 3.6 at 6.5x.
Each link below shares sources, entities, or timing with this story.
The repo reached 9,740 stars with roughly double the next-fastest Python project's daily gain. It exposes a VAD → STT → LLM → TTS pipeline behind an OpenAI Realtime-compatible WebSocket API with every stage swappable: Parakeet TDT as default STT (Whisper, Paraformer alternativ...
MAI-Transcribe-2 went to public preview September 3 through Azure Speech, first on FLEURS across 60 languages at 5.2% average WER and second on the independent Artificial Analysis leaderboard. Promotional price runs through December 31, 2026; the non-promotional rate is undisc...
Muse Voice Transcribe, released September 1, is a single real-time model doing streaming ASR, speaker diarization and endpointing at $3.00 per 1,000 audio minutes, roughly 80% below Google Cloud Speech-to-Text's standard $0.96/hour. It handles 20+ speakers natively with no pos...
Hugging Face's August 26 post introduces ColBERT-style late-interaction training with MultiVectorEncoder, MultiVectorEncoderTrainer, CachedMultiVectorMultipleNegativesRankingLoss and a MultiVectorInformationRetrievalEvaluator (Hugging Face). Their finetuned mLateOn-medical rea...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.