Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 251–300 of 328 papers

Title	Date	Tasks	Status	Hype
End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors	May 20, 2020	ClusteringDecoder	CodeCode Available	1
A Thousand Words are Worth More Than One Recording: NLP Based Speaker Change Point Detection	May 18, 2020	Change Point Detectionspeaker-diarization	—Unverified	0
Speech Recognition and Multi-Speaker Diarization of Long Conversations	May 16, 2020	Data Augmentationspeaker-diarization	CodeCode Available	1
Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario	May 14, 2020	Action DetectionActivity Detection	—Unverified	0
Preparation of Bangla Speech Corpus from Publicly Available Audio \& Text	May 1, 2020	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
Semi-supervised Acoustic Modelling for Five-lingual Code-switched ASR using Automatically-segmented Soap Opera Speech	May 1, 2020	Acoustic ModellingAction Detection	—Unverified	0
CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings	Apr 20, 2020	speaker-diarizationSpeaker Diarization	—Unverified	0
Speaker Diarization with Lexical Information	Apr 13, 2020	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
Semi-supervised acoustic modelling for five-lingual code-switched ASR using automatically-segmented soap opera speech	Apr 8, 2020	Acoustic ModellingAction Detection	—Unverified	0
Probabilistic embeddings for speaker diarization	Apr 6, 2020	Clusteringspeaker-diarization	CodeCode Available	0
ASR Error Correction and Domain Adaptation Using Machine Translation	Mar 13, 2020	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
Tackling real noisy reverberant meetings with all-neural source separation, counting, and diarization system	Mar 9, 2020	Allspeaker-diarization	—Unverified	0
Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap	Mar 5, 2020	Clusteringspeaker-diarization	CodeCode Available	1
End-to-End Neural Diarization: Reformulating Speaker Diarization as Simple Multi-label Classification	Feb 24, 2020	ClusteringGeneral Classification	CodeCode Available	1
Speaker Diarization with Region Proposal Network	Feb 14, 2020	Region Proposalspeaker-diarization	CodeCode Available	1
Self-supervised learning for audio-visual speaker diarization	Feb 13, 2020	Self-Supervised Learningspeaker-diarization	—Unverified	0
Compositional Embeddings for Multi-Label One-Shot Learning	Feb 11, 2020	Object DetectionObject Recognition	—Unverified	0
Phoneme Boundary Detection using Learnable Segmental Features	Feb 11, 2020	Boundary DetectionKeyword Spotting	CodeCode Available	1
Advances in Online Audio-Visual Meeting Transcription	Dec 10, 2019	Sound Source Localizationspeaker-diarization	—Unverified	0
The Speed Submission to DIHARD II: Contributions & Lessons Learned	Nov 6, 2019	Action DetectionActivity Detection	—Unverified	0
pyannote.audio: neural building blocks for speaker diarization	Nov 4, 2019	Action DetectionActivity Detection	CodeCode Available	3
Supervised online diarization with sample mean loss for multi-domain data	Nov 4, 2019	Clusteringspeaker-diarization	CodeCode Available	0
Robust speaker recognition using unsupervised adversarial invariance	Nov 3, 2019	speaker-diarizationSpeaker Diarization	CodeCode Available	0
Meta-learning for robust child-adult classification from speech	Oct 28, 2019	ClassificationMeta-Learning	—Unverified	0
Speaker diarization using latent space clustering in generative adversarial network	Oct 24, 2019	ClusteringDiagnostic	—Unverified	0
Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm	Oct 24, 2019	Clusteringspeaker-diarization	—Unverified	0
Compositional Embeddings: Joint Perception and Comparison of Class Label Sets	Sep 25, 2019	object-detectionObject Detection	—Unverified	0
Simultaneous Speech Recognition and Speaker Diarization for Monaural Dialogue Recordings with Target-Speaker Acoustic Models	Sep 17, 2019	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
End-to-End Neural Speaker Diarization with Self-attention	Sep 13, 2019	Clusteringspeaker-diarization	CodeCode Available	1
End-to-End Neural Speaker Diarization with Permutation-Free Objectives	Sep 12, 2019	ClusteringDomain Adaptation	CodeCode Available	0
LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization	Jul 23, 2019	Change DetectionClustering	CodeCode Available	0
Toeplitz Inverse Covariance based Robust Speaker Clustering for Naturalistic Audio Streams	Jul 12, 2019	Clusteringspeaker-diarization	—Unverified	0
Joint Speech Recognition and Speaker Diarization via Sequence Transduction	Jul 9, 2019	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
Ultrasound tongue imaging for diarization and alignment of child speech therapy sessions	Jul 1, 2019	speaker-diarizationSpeaker Diarization	CodeCode Available	0
Large-Scale Speaker Diarization of Radio Broadcast Archives	Jun 19, 2019	speaker-diarizationSpeaker Diarization	—Unverified	0
The Second DIHARD Diarization Challenge: Dataset, task, and baselines	Jun 18, 2019	Action DetectionActivity Detection	CodeCode Available	0
UWB-NTIS Speaker Diarization System for the DIHARD II 2019 Challenge	May 27, 2019	Clusteringspeaker-diarization	—Unverified	0
Meeting Transcription Using Virtual Microphone Arrays	May 3, 2019	speaker-diarizationSpeaker Diarization	—Unverified	0
Latent Class Model with Application to Speaker Diarization	Apr 25, 2019	modelspeaker-diarization	—Unverified	0
All-neural online source separation, counting, and diarization for meeting analysis	Feb 21, 2019	AllAutomatic Speech Recognition	—Unverified	0
AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection	Jan 5, 2019	Active Speaker DetectionAudio-Visual Active Speaker Detection	CodeCode Available	1
Détection de locuteurs dans les séries TV	Dec 18, 2018	Clusteringspeaker-diarization	—Unverified	0
Audiovisual speaker diarization of TV series	Dec 18, 2018	speaker-diarizationSpeaker Diarization	—Unverified	0
Constrained speaker diarization of TV series based on visual patterns	Dec 18, 2018	Clusteringspeaker-diarization	—Unverified	0
Speaker Diarization With Lexical Information	Nov 27, 2018	Clusteringspeaker-diarization	—Unverified	0
Designing an Effective Metric Learning Pipeline for Speaker Diarization	Nov 1, 2018	Metric Learningspeaker-diarization	—Unverified	0
CountNet: Estimating the Number of Concurrent Speakers Using Supervised Learning Speaker Count Estimation	Oct 28, 2018	blind source separationspeaker-diarization	CodeCode Available	0
Semi-supervised acoustic model training for speech with code-switching	Oct 23, 2018	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
Fully Supervised Speaker Diarization	Oct 10, 2018	Clusteringspeaker-diarization	CodeCode Available	0
The EURECOM Submission to the First DIHARD Challenge	Sep 6, 2018	Clusteringspeaker-diarization	CodeCode Available	0

Show:10 25 50

← PrevPage 6 of 7Next →

All datasets CALLHOME NIST-SRE 2000 AMI Lapel AMI MixHeadset CH109 DIHARD ETAPE AMI CALLHOME-109 AliMeeting DIHARD II Hub5'00 CallHome

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	COS+NJW-SC (Oracle SAD)	DER(%)	24.05	—	Unverified
2	EEND	DER(%)	23.07	—	Unverified
3	COS+AHC (Oracle SAD)	DER(%)	21.13	—	Unverified
4	SA-EEND (2-spk, no-adapt)	DER(%)	12.66	—	Unverified
5	EEND-OLA	DER(%)	12.57	—	Unverified
6	SA-EEND (2-spk, adapted)	DER(%)	10.76	—	Unverified
7	TOLD	DER(%)	10.14	—	Unverified
8	COS+B-SC (Oracle SAD)	DER(ig olp)	8.78	—	Unverified
9	PLDA+AHC (Oracle SAD)	DER(ig olp)	8.39	—	Unverified
10	COS+NME-SC (Oracle SAD)	DER(ig olp)	7.29	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	8.39	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	6.73	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	6.47	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	6.37	—	Unverified
5	x-vector (MCGAN)	DER(%)	5.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	ECAPA (SC)	DER(%)	2.36	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	2.03	—	Unverified
3	TitaNet-S (NME-SC)	DER(%)	2	—	Unverified
4	TitaNet-M (NME-SC)	DER(%)	1.99	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	TitaNet-S (NME-SC)	DER(%)	2.22	—	Unverified
2	TitaNet-M (NME-SC)	DER(%)	1.79	—	Unverified
3	ECAPA (SC)	DER(%)	1.78	—	Unverified
4	TitaNet-L (NME-SC)	DER(%)	1.73	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	x-vector (PLDA + AHC)	DER(%)	9.72	—	Unverified
2	TitaNet-L (NME-SC)	DER(%)	1.19	—	Unverified
3	TitaNet-M (NME-SC)	DER(%)	1.13	—	Unverified
4	TitaNet-S (NME-SC)	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline (the best result in the literature as of Oct.2019)	DER(%)	11.2	—	Unverified
2	pyannote (MFCC)	DER(%)	10.5	—	Unverified
3	pyannote (waveform)	DER(%)	9.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Baseline	DER(%)	7.7	—	Unverified
2	pyannote (MFCC)	DER(%)	5.6	—	Unverified
3	pyannote (waveform)	DER(%)	4.9	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	pyannote (MFCC)	DER(%)	6.3	—	Unverified
2	pyannote (waveform)	DER(%)	6	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	d-vector + spectral	DER(%)	12.54	—	Unverified
2	titanet-s	DER(%)	1.11	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SOND	DER(%)	4.46	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN-SML	DER(%)	27.3	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	UIS-RNN	V	10.6	—	Unverified