Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 1–10 of 328 papers

Title	Date	Tasks	Status	Hype
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models	Jun 23, 2025	Domain AdaptationGPU	CodeCode Available	3
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset	Jun 17, 2025	Domain Adaptationspeaker-diarization	—Unverified	0
Exploring Speaker Diarization with Mixture of Experts	Jun 17, 2025	Mixture-of-Expertsspeaker-diarization	—Unverified	0
Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models	Jun 16, 2025	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition	Jun 15, 2025	Decoderspeaker-diarization	—Unverified	0
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models	Jun 6, 2025	Automatic Speech Recognitionspeaker-diarization	—Unverified	0
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling	Jun 5, 2025	AttributeDecoder	—Unverified	0
Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm	Jun 3, 2025	Action DetectionActivity Detection	CodeCode Available	1
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization	May 30, 2025	GPUKnowledge Distillation	—Unverified	0
Pretraining Multi-Speaker Identification for Neural Speaker Diarization	May 30, 2025	speaker-diarizationSpeaker Diarization	—Unverified	0

Show:10 25 50

← PrevPage 1 of 33Next →

All datasets CALLHOME NIST-SRE 2000 AMI Lapel AMI MixHeadset CH109 DIHARD ETAPE AMI CALLHOME-109 AliMeeting DIHARD II Hub5'00 CallHome

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	Baseline	DER(%)	7.7	—	Unverified
2	pyannote (MFCC)	DER(%)	5.6	—	Unverified
3	pyannote (waveform)	DER(%)	4.9	—	Unverified