Reddit
NVIDIA Ships NemotronLabs VoiceChat 11B — First Open Full-Duplex Speech Model With Tool Calling, ~450ms Turn-Taking
NVIDIA uploaded NVIDIA-NemotronLabs-VoiceChat-11B to Hugging Face on August 3: a Fast Conformer speech encoder in front of Nemotron Nano v2 9B (hybrid Mamba/Transformer) with an NVIDIA TTS decoder behind it, collapsing the usual ASR→LLM→TTS cascade into one model. The card reports ~450ms latency on smooth turn-taking, 480ms on user interruption, and #2 among open full-duplex models on VoiceBench, released under the OpenMDW 1.1 license. The genuinely new part is a separate output channel that emits tool-call scripts while audio keeps flowing, with configurable 'on-hold' phrases the agent speaks during tool execution — the first open full-duplex model to do this.
↳ Follow the thread