Reddit
Microsoft released VibeVoice-ASR-Streaming-7B, a 9B MIT-licensed streaming speaker-attributed ASR model
Microsoft published VibeVoice-ASR-Streaming-7B on Hugging Face under the MIT license, a 9-billion-parameter model that transcribes who said what continuously as speech arrives, supports custom hotwords for domain terms, and covers ten languages including Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian and Spanish. The accompanying technical report is arXiv 2609.02812, published roughly a day before the model page. The top comments are entirely about expected takedown risk, with one user already mirroring the weights after a previous VibeVoice release was pulled.
↳ Follow the thread