NVIDIA Launches VoiceChat 11B, Bidirectional Speech Model with Ultra-Low Latency
NVIDIA released VoiceChat 11B, an 11-billion-parameter speech-to-speech model for real-time conversation. Rather than chaining separate speech recognition, language model, and text-to-speech components, the model performs bidirectional audio processing in a unified system, reducing latency to 448 milliseconds. It allows users to interrupt the agent and supports tool calling with high accuracy through separate output channels. The current version is publicly available but recommended for research only, with known limitations including a maximum audio context of two minutes and degradation after multiple conversation turns. The model was trained on approximately 550,000 hours of audio.