Nari Labs Got Qwen3-TTS Under 50ms Time-to-First-Audio at 10 RPS on a Single H100, at ~$2 per Million Characters
Nari Labs published the optimization writeup on August 19 (155 points): Qwen3-TTS 1.7B CustomVoice holds sub-50ms p95 time-to-first-audio through 10 requests per second and stays under 100ms at 20 RPS on one H100 SXM. Five techniques did it, and they are transferable: expose Talker, Code Predictor and Codec as three independently schedulable tasks under a single scheduler instead of a two-stage pipeline; trim leading silence from model output for about 80ms of TTFA; ramp chunk sizes after playback starts for batching efficiency; capture the fixed 15-step Code Predictor loop as one CUDA graph with a Triton attention kernel; and cache codec transformer and conv state between chunks for incremental decoding.
↳ Follow the thread