r/LocalLLaMA / Mistral AI / TechCrunch / VentureBeat·medium signal
Mistral launched Voxtral TTS, a 4B-parameter open-weight text-to-speech model supporting 9 languages (English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, Arabic) with voice cloning from just 3 seconds of reference audio. Human evaluations show naturalness parity with ElevenLabs Flash v2.5 and the larger v3 model. Available on Hugging Face under Creative Commons and via API at $0.016 per 1K characters. At 4B parameters, it runs on consumer hardware including modern laptops.