
Nvidia’s new platform adds software and hardware guardrails for AI agents.
Nvidia’s Nemotron 3 Diarization is a free 100M-parameter AI model that identifies up to eight speakers in recordings or live audio, with benchmark-leading results reported on VoiceArena.

Nvidia has released Nemotron 3 Diarization, an AI model designed to identify which speaker is talking at any given moment in a conversation. The model has about 100 million parameters, and its weights are freely available. It works with both recordings and live audio, making it relevant for real-time conversation analysis as well as post-processing.
Nemotron 3 Diarization can distinguish up to eight speakers and detect when multiple people talk at the same time. When paired with a speech recognition system such as Parakeet, it can produce transcripts with speaker labels, though the labels are anonymous, such as “speaker_2.” Accuracy can decline when there are more participants, heavy background noise, or reverb.

The model’s audio buffer can be set across four levels, ranging from 30.4 seconds down to 0.32 seconds. Shorter buffers generally reduce accuracy. On VoiceArena’s Diarization-Bench, Nemotron 3 Diarization is reported in first place with a 14.72 percent error rate, ahead of the next best system at 19.3 percent.
The release gives developers a freely available option for speaker diarization in live or recorded conversations. The benchmark notes are important: overlapping speech counts against systems, and even tiny misalignments at speaker transitions are scored as errors. Compared with its predecessor, Streaming Sortformer, Nemotron 3 cuts the error rate by an average of 41 percent across eight test scenarios when using a 1.04-second buffer.

Nvidia’s new platform adds software and hardware guardrails for AI agents.

Huawei plans to launch its Ascend 960DT AI chip in Q1 2027, ahead of its previous Q3 timeline.
Groq is rebuilding around AI cloud demand after a major Nvidia deal reshaped its business.

The AI infrastructure startup is pushing open networking fabric for GPUs and accelerators.