AI · August 10, 2026
NVIDIA's NemotronLabs VoiceChat 11B cuts voice AI latency to 448ms
NVIDIA has released NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model with ~448ms turn-taking latency and live tool calling for real-time voice actions.
What happened
NVIDIA has released NemotronLabs VoiceChat 11B, an open-source, full-duplex speech-to-speech AI model designed for natural, real-time voice conversations. The model reports turn-taking latency of roughly 448 milliseconds and supports live tool calling, allowing it to trigger actions or fetch information mid-conversation rather than waiting for a spoken exchange to finish.
Unlike conventional voice assistants that process speech in a rigid listen-then-respond sequence, a full-duplex architecture allows the model to listen and speak simultaneously, more closely mirroring how humans naturally interrupt, overlap and pace a conversation.
Why it matters
Turn-taking speed is one of the most under-discussed drivers of perceived quality in voice interactions. Delays much beyond half a second are typically enough for callers to feel they are talking to a machine rather than being heard; models that close that gap materially change how "human" an automated voice channel feels, without necessarily changing what the system is actually capable of doing.
Live tool calling is arguably the more consequential piece for service design: it means a voice agent could look up an order, check a policy or take an action while the conversation is still flowing, rather than parking the caller in a scripted hold pattern. For contact centres and voice-first customer service, that combination — low-latency turn-taking plus in-conversation action — is the technical building block for voice bots that feel less like IVR menus and more like a competent agent.
By the numbers
- 448 milliseconds — reported turn-taking latency for NemotronLabs VoiceChat 11B
- 11 billion — parameter count of the released model
The Renascence take
The headline number here is latency, but the more interesting story for CX teams is what an open, tool-calling voice model does to the build-versus-buy calculus for voice automation.
Most organisations still treat "the bot sounds robotic" as a scripting problem, when it is very often a latency problem — customers read a half-second silence as incompetence, not thoughtfulness. An open model that pushes turn-taking under half a second, and can act on live data mid-call, shifts the real design question away from "can it sound natural" toward "what should it be trusted to do unsupervised." Teams experimenting with voice AI should be piloting the tool-calling boundary now — which actions get handed to the model outright, which get confirmed, and which stay human-only — because that governance decision, not the latency benchmark, is what will determine whether faster voice bots earn trust or simply fail faster.
Sources
This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.