Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 101–125 of 328 papers

Title	Date	Tasks	Status
Home monitoring for frailty detection through sound and speaker diarization analysis	Aug 17, 2023	Privacy Preservingspeaker-diarization	—Unverified
Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization	Jun 26, 2023	ClusteringCommunity Detection	—Unverified
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification	Apr 26, 2024	speaker-diarizationSpeaker Diarization	—Unverified
Generation of Speaker Representations Using Heterogeneous Training Batch Assembly	Mar 30, 2022	speaker-diarizationSpeaker Diarization	—Unverified
Chronological Self-Training for Real-Time Speaker Diarization	Aug 5, 2022	speaker-diarizationSpeaker Diarization	—Unverified
From Modular to End-to-End Speaker Diarization	Jun 27, 2024	speaker-diarizationSpeaker Diarization	—Unverified
GIST-AiTeR Speaker Diarization System for VoxCeleb Speaker Recognition Challenge (VoxSRC) 2023	Aug 15, 2023	speaker-diarizationSpeaker Diarization	—Unverified
Guided Speaker Embedding	Oct 16, 2024	speaker-diarizationSpeaker Diarization	—Unverified
CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings	Apr 20, 2020	speaker-diarizationSpeaker Diarization	—Unverified
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification.	Jun 1, 2022	speaker-diarizationSpeaker Diarization	—Unverified
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization	May 30, 2025	GPUKnowledge Distillation	—Unverified
Implicit Self-supervised Language Representation for Spoken Language Diarization	Aug 21, 2023	speaker-diarizationSpeaker Diarization	—Unverified
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection	Feb 13, 2024	Action DetectionActivity Detection	—Unverified
Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm	Oct 24, 2019	Clusteringspeaker-diarization	—Unverified
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling	Jun 5, 2025	AttributeDecoder	—Unverified
Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications	Mar 11, 2024	Action DetectionActivity Detection	—Unverified
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation	Sep 19, 2023	speaker-diarizationSpeaker Diarization	—Unverified
Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech	Jun 13, 2024	Language Identificationspeaker-diarization	—Unverified
Improving Transformer-based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads	Mar 2, 2023	Action DetectionActivity Detection	—Unverified
Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings	Sep 25, 2024	Clusteringspeaker-diarization	—Unverified
Indigenous language technologies in Canada: Assessment, challenges, and successes	Aug 1, 2018	Machine TranslationOptical Character Recognition	—Unverified
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization	Aug 22, 2024	speaker-diarizationSpeaker Diarization	—Unverified
BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers	Nov 5, 2020	ClusteringDecoder	—Unverified
Interrelate Training and Searching: A Unified Online Clustering Framework for Speaker Diarization	Jun 28, 2022	ClusteringOnline Clustering	—Unverified
A Review of Speaker Diarization: Recent Advances with Deep Learning	Jan 24, 2021	Deep LearningRetrieval	—Unverified

Show:10 25 50

← PrevPage 5 of 14Next →

All datasets CALLHOME NIST-SRE 2000 AMI Lapel AMI MixHeadset CH109 DIHARD ETAPE AMI CALLHOME-109 AliMeeting DIHARD II Hub5'00 CallHome

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	COS+NJW-SC (Oracle SAD)	DER(%)	24.05	—	Unverified
2	EEND	DER(%)	23.07	—	Unverified
3	COS+AHC (Oracle SAD)	DER(%)	21.13	—	Unverified
4	SA-EEND (2-spk, no-adapt)	DER(%)	12.66	—	Unverified
5	EEND-OLA	DER(%)	12.57	—	Unverified
6	SA-EEND (2-spk, adapted)	DER(%)	10.76	—	Unverified
7	TOLD	DER(%)	10.14	—	Unverified
8	COS+B-SC (Oracle SAD)	DER(ig olp)	8.78	—	Unverified
9	PLDA+AHC (Oracle SAD)	DER(ig olp)	8.39	—	Unverified
10	COS+NME-SC (Oracle SAD)	DER(ig olp)	7.29	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	8.39	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	6.73	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	6.47	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	6.37	—	Unverified
5	x-vector (MCGAN)	DER(%)	5.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	ECAPA (SC)	DER(%)	2.36	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	2.03	—	Unverified
3	TitaNet-S (NME-SC)	DER(%)	2	—	Unverified
4	TitaNet-M (NME-SC)	DER(%)	1.99	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	TitaNet-S (NME-SC)	DER(%)	2.22	—	Unverified
2	TitaNet-M (NME-SC)	DER(%)	1.79	—	Unverified
3	ECAPA (SC)	DER(%)	1.78	—	Unverified
4	TitaNet-L (NME-SC)	DER(%)	1.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	9.72	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	1.19	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	1.13	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline (the best result in the literature as of Oct.2019)	DER(%)	11.2	—	Unverified
2	pyannote (MFCC)	DER(%)	10.5	—	Unverified
3	pyannote (waveform)	DER(%)	9.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline	DER(%)	7.7	—	Unverified
2	pyannote (MFCC)	DER(%)	5.6	—	Unverified
3	pyannote (waveform)	DER(%)	4.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	pyannote (MFCC)	DER(%)	6.3	—	Unverified
2	pyannote (waveform)	DER(%)	6	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	d-vector + spectral	DER(%)	12.54	—	Unverified
2	titanet-s	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SOND	DER(%)	4.46	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN-SML	DER(%)	27.3	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN	V	10.6	—	Unverified