Fetching from the wire…
Research2026-07-26 · source-backed
Paul Azunre released twenty-one monolingual and five multilingual w2v-BERT 2.0 ASR base models spanning varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya, and Zimbabwe, languages with roughly a hundred million first-language speakers (arXiv 2607.21540). An annealed multi-step learning-rate schedule plus language conditioning via one-hot language identity prefixed to acoustic features brings the five multilingual families to 10-13% average WER, closing most of the gap to monolingual. Apache-2.0 attribution-only on the Hugging Face KhayaAI org, so commercial fine-tuning is permitted. Caveat: training data is primarily religious text, which constrains domain coverage hard.
Each link below shares sources, entities, or timing with this story.
granite-speech-5.0-470m-turboctc uses 16 conformer blocks trained with CTC on a 16,384 BPE head, temporal subsampling by 8, 128-frame block attention and self-conditioned CTC from the middle layer. IBM reports above 12,600 RTFx on a single H200, roughly 3.5 hours of audio per...
The top trending HuggingFace paper (274 upvotes) introduces dots.tts, a 2B continuous autoregressive text-to-speech model hitting best average Seed-TTS-Eval (WER 0.94%/1.30% zh/en) with strong cloning and emotional range. CFG-aware MeanFlow distillation gives 85ms first-packet...
Published August 25, it's IBM's first family of dense decoder-only reasoning models, with the 30B flagship claiming state-of-the-art resolve rates on SWE-Bench Pro and Terminal-Bench. The 8B and 30B went through an agentic training curriculum on real sandboxes for software eng...
A Show HN at 87 points ships local transcription and system-wide push-to-talk for Apple Silicon Macs, 119 stars, Apache-2.0, v0.2.0-beta.1. Runs Qwen3-ASR through mlx-qwen3-asr with a size choice: 1.7B for accuracy needing 3.4 GB unified memory, 0.6B for speed needing 1.2 GB....
arXiv 2608.19936 quantifies benchmark optimization in speech recognition by focusing on cases where the audio underdetermines the reference transcript, using three probe families: reference disagreement, masked-number recovery, and orthographic switching. The highest-scoring o...
arXiv 2607.28165 attacks always-listening multimodal agents by embedding malicious instructions in ambient audio that overlaps user speech, using instruction augmentation and scenario concealment so the injection is imperceptible. Eleven agents evaluated, 69.10% average ASR ag...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.