Inflect v2 Ships Two Complete TTS Models at 3.97M and 9.36M Parameters, Running 10.7x Real-Time on Four CPU Threads
Developer owensong released Inflect-Nano-v2 (3,966,721 deployable parameters) and Inflect-Micro-v2 (9,356,513) under Apache-2.0 — VITS-family end-to-end text-to-waveform generators with 128 latent channels, 3 encoder layers, 4 flow coupling blocks, and 24 kHz mono output. Nano-v2 runs at 0.0933 RTF (10.72x real-time) on four CPU threads; Micro-v2 at 0.1593 RTF (6.28x). The trade-off is deliberate and narrow: English only, one fixed male voice, and the training-corpus pipeline and filtering infrastructure stay private, so this is open-weight rather than open-source. The r/LocalLLaMA announcement drew 410 upvotes and 99 comments — the practical read is that usable on-device TTS now fits in under 10MB of parameters, which puts voice output inside embedded and edge budgets that previously ruled it out.
↳ Follow the thread