Nvidia released Nemotron 3 Diarization, an open-weight model with 100 million parameters capable of identifying up to eight speakers simultaneously with latency as low as 320 milliseconds. On the AISHELL-4 benchmark, this model achieved a Diarization Error Rate of 9.8%, compared to 27.2% for its predecessor, the Streaming Sortformer v2.1 released in July 2025. The system operates in both streaming and offline modes, with 10-millisecond resolution for speaker activity probabilities. Deployment costs can reach as low as $0.01 per audio hour through select inference providers. This model is part of Nvidia’s broader strategy to build a full stack for voice-driven applications, with use cases ranging from meeting transcription to call center analytics.
Source: Read the original article

