Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–75 of 328 papers

Title	Date	Tasks	Status	Hype
Speaker Diarization with Region Proposal Network	Feb 14, 2020	Region Proposalspeaker-diarization	CodeCode Available	1
Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection	Sep 23, 2021	Clusteringspeaker-diarization	CodeCode Available	1
Utterance-by-utterance overlap-aware neural diarization with Graph-PIT	Jul 28, 2022	ClusteringSegmentation	CodeCode Available	1
VoxLingua107: a Dataset for Spoken Language Recognition	Nov 25, 2020	Action DetectionActivity Detection	CodeCode Available	1
DiariST: Streaming Speech Translation with Speaker Diarization	Sep 14, 2023	speaker-diarizationSpeaker Diarization	CodeCode Available	1
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors	Dec 7, 2023	Decoderspeaker-diarization	CodeCode Available	1
All-neural online source separation, counting, and diarization for meeting analysis	Feb 21, 2019	AllAutomatic Speech Recognition	—Unverified	0
Constrained speaker diarization of TV series based on visual patterns	Dec 18, 2018	Clusteringspeaker-diarization	—Unverified	0
Computer-assisted Speaker Diarization: How to Evaluate Human Corrections	May 1, 2018	Active LearningFace Recognition	—Unverified	0
Assessing the Robustness of Spectral Clustering for Deep Speaker Diarization	Mar 21, 2024	Clusteringspeaker-diarization	—Unverified	0
A sticky HDP-HMM with application to speaker diarization	May 15, 2009	speaker-diarizationSpeaker Diarization	—Unverified	0
Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for M2MeT Challenge	Feb 6, 2022	Action DetectionActivity Detection	—Unverified	0
Domain-Dependent Speaker Diarization for the Third DIHARD Challenge	Jan 25, 2021	ClusteringDimensionality Reduction	—Unverified	0
Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding	Dec 5, 2024	Audio GenerationAutomatic Speech Recognition	—Unverified	0
Compositional Embeddings: Joint Perception and Comparison of Class Label Sets	Sep 25, 2019	object-detectionObject Detection	—Unverified	0
ASR Error Correction and Domain Adaptation Using Machine Translation	Mar 13, 2020	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
Compositional Embeddings for Multi-Label One-Shot Learning	Feb 11, 2020	Object DetectionObject Recognition	—Unverified	0
ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings	Jun 5, 2024	speaker-diarizationSpeaker Diarization	—Unverified	0
Aligning Speakers: Evaluating and Visualizing Text-based Diarization Using Efficient Multiple Sequence Alignment (Extended Version)	Sep 14, 2023	Multiple Sequence Alignmentspeaker-diarization	—Unverified	0
3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization	Mar 29, 2024	Self-Supervised Learningspeaker-diarization	—Unverified	0
ECAPA-TDNN Embeddings for Speaker Diarization	Apr 3, 2021	speaker-diarizationSpeaker Diarization	—Unverified	0
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification	Apr 26, 2024	speaker-diarizationSpeaker Diarization	—Unverified	0
Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization	Jun 26, 2023	ClusteringCommunity Detection	—Unverified	0
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification.	Jun 1, 2022	speaker-diarizationSpeaker Diarization	—Unverified	0
Chronological Self-Training for Real-Time Speaker Diarization	Aug 5, 2022	speaker-diarizationSpeaker Diarization	—Unverified	0

Show:10 25 50

← PrevPage 3 of 14Next →

All datasets CALLHOME NIST-SRE 2000 AMI Lapel AMI MixHeadset CH109 DIHARD ETAPE AMI CALLHOME-109 AliMeeting DIHARD II Hub5'00 CallHome

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	COS+NJW-SC (Oracle SAD)	DER(%)	24.05	—	Unverified
2	EEND	DER(%)	23.07	—	Unverified
3	COS+AHC (Oracle SAD)	DER(%)	21.13	—	Unverified
4	SA-EEND (2-spk, no-adapt)	DER(%)	12.66	—	Unverified
5	EEND-OLA	DER(%)	12.57	—	Unverified
6	SA-EEND (2-spk, adapted)	DER(%)	10.76	—	Unverified
7	TOLD	DER(%)	10.14	—	Unverified
8	COS+B-SC (Oracle SAD)	DER(ig olp)	8.78	—	Unverified
9	PLDA+AHC (Oracle SAD)	DER(ig olp)	8.39	—	Unverified
10	COS+NME-SC (Oracle SAD)	DER(ig olp)	7.29	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	8.39	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	6.73	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	6.47	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	6.37	—	Unverified
5	x-vector (MCGAN)	DER(%)	5.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	ECAPA (SC)	DER(%)	2.36	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	2.03	—	Unverified
3	TitaNet-S (NME-SC)	DER(%)	2	—	Unverified
4	TitaNet-M (NME-SC)	DER(%)	1.99	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	TitaNet-S (NME-SC)	DER(%)	2.22	—	Unverified
2	TitaNet-M (NME-SC)	DER(%)	1.79	—	Unverified
3	ECAPA (SC)	DER(%)	1.78	—	Unverified
4	TitaNet-L (NME-SC)	DER(%)	1.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	9.72	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	1.19	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	1.13	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline (the best result in the literature as of Oct.2019)	DER(%)	11.2	—	Unverified
2	pyannote (MFCC)	DER(%)	10.5	—	Unverified
3	pyannote (waveform)	DER(%)	9.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline	DER(%)	7.7	—	Unverified
2	pyannote (MFCC)	DER(%)	5.6	—	Unverified
3	pyannote (waveform)	DER(%)	4.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	pyannote (MFCC)	DER(%)	6.3	—	Unverified
2	pyannote (waveform)	DER(%)	6	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	d-vector + spectral	DER(%)	12.54	—	Unverified
2	titanet-s	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SOND	DER(%)	4.46	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN-SML	DER(%)	27.3	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN	V	10.6	—	Unverified