Fetching from the wire…
Agents2026-09-15 · source-backed
A common pipeline shape converts a non-English request to English for inter-agent communication and back-translates the answer (arXiv 2609.15079). Testing a two-agent extraction-answer system on Aya-23-8B across Hindi, Chinese, Spanish and Arabic with 300 samples per language, English-forced routing lost 13.0 exact-match points for Spanish and 30.6 for Hindi against native-language routing. chrF overlap with English references correlated with failures, which points at translation loss. Native-language routing matters most when source and target are typologically distant.
Each link below shares sources, entities, or timing with this story.
Finally, a story about building something instead of worrying about something. Mistral released Voxtral TTS on March 26, an open-source text-to-speech model built on Ministral 3B. The numbers are striking: 90ms time-to-first-audio, 6x real-time factor (a 10-second clip generat...
English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, Modern Standard Arabic, Korean, Brazilian Portuguese, male and female voices each (HF). Latency by hardware: 32ms/239ms at 64 concurrent on B200, 47ms/275ms on H100, 79ms/395ms on A100. Character...
arXiv 2609.05141 uses 124 expert-authored, difficulty-screened questions across seven capability groups and 19 subtasks in five domains, each under four matched conditions (English or Chinese, all-images-first or interleaved), giving 496 instances. Weakest areas: document perc...
A 6B Diffusion Transformer trained from scratch paired with a frozen VLM understanding module on the LLaDA2.0-Mini backbone, building a visual generative prior through image-only pre-training and mid-training across a 220M-sample pipeline before touching paired image-text data...
English and Chinese, with reference-free voice design from a natural-language description, reference-guided cloning and low-latency streaming, currently first among open-weight models on the Artificial Analysis TTS leaderboard. The 6GB/4x-realtime figure comes from testers usi...
arXiv 2607.13125: 33 authors, a unified multimodal understanding-and-generation model under Apache 2.0 with weights, code, and recipes, trained on only 208.62 million unique images at a theoretical training cost around $400,000. Four variants (Base, Turbo, Edit, Edit-Turbo) co...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.