Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–75 of 328 papers

Title	Date	Tasks	Status	Hype	Score
Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors	Sep 25, 2023	Decoderspeaker-diarization	CodeCode Available	1	5
DiaCorrect: Error Correction Back-end For Speaker Diarization	Sep 15, 2023	Automatic Speech RecognitionDecoder	CodeCode Available	1	5
Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks	Jun 7, 2023	Audio ClassificationAudio Tagging	CodeCode Available	1	5
End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors	May 20, 2020	ClusteringDecoder	CodeCode Available	1	5
Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection	Sep 23, 2021	Clusteringspeaker-diarization	CodeCode Available	1	5
Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm	Jun 3, 2025	Action DetectionActivity Detection	CodeCode Available	1	5
CountNet: Estimating the Number of Concurrent Speakers Using Supervised Learning Speaker Count Estimation	Oct 28, 2018	blind source separationspeaker-diarization	CodeCode Available	0	5
Speaker Embedding-aware Neural Diarization for Flexible Number of Speakers with Textual Information	Nov 28, 2021	Action DetectionActivity Detection	CodeCode Available	0	5
Supervised Hierarchical Clustering using Graph Neural Networks for Speaker Diarization	Feb 24, 2023	ClusteringGraph Clustering	CodeCode Available	0	5
Speaker Diarization using Two-pass Leave-One-Out Gaussian PLDA Clustering of DNN Embeddings	Apr 6, 2021	Clusteringspeaker-diarization	CodeCode Available	0	5
A Comprehensive Evaluation of Incremental Speech Recognition and Diarization for Conversational AI	Dec 1, 2020	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	0	5
Compositional embedding models for speaker identification and diarization with simultaneous speech from 2+ speakers	Oct 22, 2020	speaker-diarizationSpeaker Diarization	CodeCode Available	0	5
Supervised online diarization with sample mean loss for multi-domain data	Nov 4, 2019	Clusteringspeaker-diarization	CodeCode Available	0	5
Compositional Clustering: Applications to Multi-Label Object Recognition and Speaker Identification	Sep 9, 2021	ClusteringFew-Shot Learning	CodeCode Available	0	5
Self-Supervised Metric Learning With Graph Clustering For Speaker Diarization	Sep 14, 2021	ClusteringGraph Clustering	CodeCode Available	0	5
Robust speaker recognition using unsupervised adversarial invariance	Nov 3, 2019	speaker-diarizationSpeaker Diarization	CodeCode Available	0	5
Scalable Adaptation of State Complexity for Nonparametric Hidden Markov Models	Dec 1, 2015	speaker-diarizationSpeaker Diarization	CodeCode Available	0	5
Self-supervised Representation Learning With Path Integral Clustering For Speaker Diarization	Apr 19, 2021	ClusteringRepresentation Learning	CodeCode Available	0	5
Probabilistic embeddings for speaker diarization	Apr 6, 2020	Clusteringspeaker-diarization	CodeCode Available	0	5
Self-Tuning Spectral Clustering for Speaker Diarization	Sep 16, 2024	Clusteringspeaker-diarization	CodeCode Available	0	5
On Out-of-Distribution Detection for Audio with Deep Nearest Neighbors	Oct 27, 2022	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	0	5
On the calibration of powerset speaker diarization models	Sep 24, 2024	speaker-diarizationSpeaker Diarization	CodeCode Available	0	5
LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization	Jul 23, 2019	Change DetectionClustering	CodeCode Available	0	5
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation	Jan 7, 2024	Audio-Visual Speech RecognitionAutomatic Speech Recognition	CodeCode Available	0	5
Automating Feedback Analysis in Surgical Training: Detection, Categorization, and Assessment	Dec 1, 2024	Action DetectionActivity Detection	CodeCode Available	0	5

Show:10 25 50

← PrevPage 3 of 14Next →

All datasets CALLHOME NIST-SRE 2000 AMI Lapel AMI MixHeadset CH109 DIHARD ETAPE AMI CALLHOME-109 AliMeeting DIHARD II Hub5'00 CallHome

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	COS+NJW-SC (Oracle SAD)	DER(%)	24.05	—	Unverified
2	EEND	DER(%)	23.07	—	Unverified
3	COS+AHC (Oracle SAD)	DER(%)	21.13	—	Unverified
4	SA-EEND (2-spk, no-adapt)	DER(%)	12.66	—	Unverified
5	EEND-OLA	DER(%)	12.57	—	Unverified
6	SA-EEND (2-spk, adapted)	DER(%)	10.76	—	Unverified
7	TOLD	DER(%)	10.14	—	Unverified
8	COS+B-SC (Oracle SAD)	DER(ig olp)	8.78	—	Unverified
9	PLDA+AHC (Oracle SAD)	DER(ig olp)	8.39	—	Unverified
10	COS+NME-SC (Oracle SAD)	DER(ig olp)	7.29	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	8.39	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	6.73	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	6.47	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	6.37	—	Unverified
5	x-vector (MCGAN)	DER(%)	5.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	ECAPA (SC)	DER(%)	2.36	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	2.03	—	Unverified
3	TitaNet-S (NME-SC)	DER(%)	2	—	Unverified
4	TitaNet-M (NME-SC)	DER(%)	1.99	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	TitaNet-S (NME-SC)	DER(%)	2.22	—	Unverified
2	TitaNet-M (NME-SC)	DER(%)	1.79	—	Unverified
3	ECAPA (SC)	DER(%)	1.78	—	Unverified
4	TitaNet-L (NME-SC)	DER(%)	1.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	9.72	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	1.19	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	1.13	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline (the best result in the literature as of Oct.2019)	DER(%)	11.2	—	Unverified
2	pyannote (MFCC)	DER(%)	10.5	—	Unverified
3	pyannote (waveform)	DER(%)	9.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline	DER(%)	7.7	—	Unverified
2	pyannote (MFCC)	DER(%)	5.6	—	Unverified
3	pyannote (waveform)	DER(%)	4.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	pyannote (MFCC)	DER(%)	6.3	—	Unverified
2	pyannote (waveform)	DER(%)	6	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	d-vector + spectral	DER(%)	12.54	—	Unverified
2	titanet-s	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SOND	DER(%)	4.46	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN-SML	DER(%)	27.3	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN	V	10.6	—	Unverified