Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 126–150 of 328 papers

Title	Date	Tasks	Status
Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains	Jul 24, 2023	Multi-class Classificationspeaker-diarization	—Unverified
An Infinite Hidden Markov Model With Similarity-Biased Transitions	Jul 21, 2017	speaker-diarizationSpeaker Diarization	—Unverified
A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio	Jul 6, 2021	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified
An approach to optimize inference of the DIART speaker diarization pipeline	Aug 5, 2024	Inference OptimizationKnowledge Distillation	—Unverified
Multimodal Clustering with Role Induced Constraints for Speaker Diarization	Apr 1, 2022	Clusteringspeaker-diarization	—Unverified
DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding	May 28, 2021	speaker-diarizationSpeaker Diarization	—Unverified
DISPLACE Challenge: DIarization of SPeaker and LAnguage in Conversational Environments	Mar 1, 2023	speaker-diarizationSpeaker Diarization	—Unverified
Autoapprentissage pour le regroupement en locuteurs : premi\`eres investigations (First investigations on self trained speaker diarization )	Jul 1, 2016	Domain Adaptationspeaker-diarization	—Unverified
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation	Nov 26, 2024	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified
Audiovisual speaker diarization of TV series	Dec 18, 2018	speaker-diarizationSpeaker Diarization	—Unverified
An Experimental Review of Speaker Diarization methods with application to Two-Speaker Conversational Telephone Speech recordings	May 29, 2023	Clusteringspeaker-diarization	—Unverified
Meta-learning for robust child-adult classification from speech	Oct 28, 2019	ClassificationMeta-Learning	—Unverified
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models	Jun 6, 2025	Automatic Speech Recognitionspeaker-diarization	—Unverified
Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion	Mar 31, 2016	Clusteringspeaker-diarization	—Unverified
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset	Jun 17, 2025	Domain Adaptationspeaker-diarization	—Unverified
Long-Term Conversation Analysis: Privacy-Utility Trade-off under Noise and Reverberation	Aug 1, 2024	Action DetectionActivity Detection	—Unverified
Audio-Visual Approach For Multimodal Concurrent Speaker Detection	Jul 1, 2024	Multimodal Deep Learningspeaker-diarization	—Unverified
An Effortless Way To Create Large-Scale Datasets For Famous Speakers	May 1, 2014	Person IdentificationSpeaker Diarization	—Unverified
Advances in Online Audio-Visual Meeting Transcription	Dec 10, 2019	Sound Source Localizationspeaker-diarization	—Unverified
Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge	May 22, 2025	speaker-diarizationSpeaker Diarization	—Unverified
Matics Software Suite: New Tools for Evaluation and Data Exploration	May 1, 2018	Optical Character Recognition (OCR)Speaker Diarization	—Unverified
Meeting Transcription Using Virtual Microphone Arrays	May 3, 2019	speaker-diarizationSpeaker Diarization	—Unverified
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR	Sep 18, 2024	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified
An automated medical scribe for documenting clinical encounters	Jun 1, 2018	speaker-diarizationSpeaker Diarization	—Unverified
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives	Feb 14, 2024	GPUspeaker-diarization	—Unverified

Show:10 25 50

← PrevPage 6 of 14Next →

All datasets CALLHOME NIST-SRE 2000 AMI Lapel AMI MixHeadset CH109 DIHARD ETAPE AMI CALLHOME-109 AliMeeting DIHARD II Hub5'00 CallHome

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	COS+NJW-SC (Oracle SAD)	DER(%)	24.05	—	Unverified
2	EEND	DER(%)	23.07	—	Unverified
3	COS+AHC (Oracle SAD)	DER(%)	21.13	—	Unverified
4	SA-EEND (2-spk, no-adapt)	DER(%)	12.66	—	Unverified
5	EEND-OLA	DER(%)	12.57	—	Unverified
6	SA-EEND (2-spk, adapted)	DER(%)	10.76	—	Unverified
7	TOLD	DER(%)	10.14	—	Unverified
8	COS+B-SC (Oracle SAD)	DER(ig olp)	8.78	—	Unverified
9	PLDA+AHC (Oracle SAD)	DER(ig olp)	8.39	—	Unverified
10	COS+NME-SC (Oracle SAD)	DER(ig olp)	7.29	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	8.39	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	6.73	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	6.47	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	6.37	—	Unverified
5	x-vector (MCGAN)	DER(%)	5.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	ECAPA (SC)	DER(%)	2.36	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	2.03	—	Unverified
3	TitaNet-S (NME-SC)	DER(%)	2	—	Unverified
4	TitaNet-M (NME-SC)	DER(%)	1.99	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	TitaNet-S (NME-SC)	DER(%)	2.22	—	Unverified
2	TitaNet-M (NME-SC)	DER(%)	1.79	—	Unverified
3	ECAPA (SC)	DER(%)	1.78	—	Unverified
4	TitaNet-L (NME-SC)	DER(%)	1.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	9.72	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	1.19	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	1.13	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline (the best result in the literature as of Oct.2019)	DER(%)	11.2	—	Unverified
2	pyannote (MFCC)	DER(%)	10.5	—	Unverified
3	pyannote (waveform)	DER(%)	9.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline	DER(%)	7.7	—	Unverified
2	pyannote (MFCC)	DER(%)	5.6	—	Unverified
3	pyannote (waveform)	DER(%)	4.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	pyannote (MFCC)	DER(%)	6.3	—	Unverified
2	pyannote (waveform)	DER(%)	6	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	d-vector + spectral	DER(%)	12.54	—	Unverified
2	titanet-s	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SOND	DER(%)	4.46	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN-SML	DER(%)	27.3	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN	V	10.6	—	Unverified