DONDO Releases 26 Apache-2.0 Speech Models Covering 27 African Language Varieties at 10-13% WER
Paul Azunre released DONDO, twenty-one monolingual and five multilingual w2v-BERT 2.0 ASR base models spanning twenty-seven language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya, and Zimbabwe — languages with on the order of a hundred million first-language speakers. An annealed multi-step learning-rate schedule plus language conditioning via one-hot language identity prefixed to acoustic features brings the five multilingual families to average word error rates of 10-13%, closing most of the gap to the monolingual models. Everything ships on the Hugging Face KhayaAI organisation under Apache-2.0 (attribution only), so commercial fine-tuning is permitted; the main caveat is that training data is primarily religious text, which constrains domain coverage.
↳ Follow the thread