SOTAVerified

Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Showing 276–300 of 328 papers

TitleStatusHype
Robust speaker recognition using unsupervised adversarial invarianceCode0
Meta-learning for robust child-adult classification from speech—0
Speaker diarization using latent space clustering in generative adversarial network—0
Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm—0
Compositional Embeddings: Joint Perception and Comparison of Class Label Sets—0
Simultaneous Speech Recognition and Speaker Diarization for Monaural Dialogue Recordings with Target-Speaker Acoustic Models—0
End-to-End Neural Speaker Diarization with Permutation-Free Objectives—0
LSTM based Similarity Measurement with Spectral Clustering for Speaker DiarizationCode0
Toeplitz Inverse Covariance based Robust Speaker Clustering for Naturalistic Audio Streams—0
Joint Speech Recognition and Speaker Diarization via Sequence Transduction—0
Ultrasound tongue imaging for diarization and alignment of child speech therapy sessionsCode0
Large-Scale Speaker Diarization of Radio Broadcast Archives—0
The Second DIHARD Diarization Challenge: Dataset, task, and baselinesCode0
UWB-NTIS Speaker Diarization System for the DIHARD II 2019 Challenge—0
Meeting Transcription Using Virtual Microphone Arrays—0
Latent Class Model with Application to Speaker Diarization—0
All-neural online source separation, counting, and diarization for meeting analysis—0
Constrained speaker diarization of TV series based on visual patterns—0
Audiovisual speaker diarization of TV series—0
Détection de locuteurs dans les séries TV—0
Speaker Diarization With Lexical Information—0
Designing an Effective Metric Learning Pipeline for Speaker Diarization—0
CountNet: Estimating the Number of Concurrent Speakers Using Supervised Learning Speaker Count EstimationCode0
Semi-supervised acoustic model training for speech with code-switching—0
Fully Supervised Speaker DiarizationCode0
Show:102550
← PrevPage 12 of 14Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1COS+NJW-SC (Oracle SAD)DER(%)24.05—Unverified
2EENDDER(%)23.07—Unverified
3COS+AHC (Oracle SAD)DER(%)21.13—Unverified
4SA-EEND (2-spk, no-adapt)DER(%)12.66—Unverified
5EEND-OLADER(%)12.57—Unverified
6SA-EEND (2-spk, adapted)DER(%)10.76—Unverified
7TOLDDER(%)10.14—Unverified
8COS+B-SC (Oracle SAD)DER(ig olp)8.78—Unverified
9PLDA+AHC (Oracle SAD)DER(ig olp)8.39—Unverified
10COS+NME-SC (Oracle SAD)DER(ig olp)7.29—Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)8.39—Unverified
2TitaNet-L (NME-SC)DER(%)6.73—Unverified
3TitaNet-M (NME-SC)DER(%)6.47—Unverified
4TitaNet-S (NME-SC)DER(%)6.37—Unverified
5x-vector (MCGAN)DER(%)5.73—Unverified
#ModelMetricClaimedVerifiedStatus
1ECAPA (SC)DER(%)2.36—Unverified
2TitaNet-L (NME-SC)DER(%)2.03—Unverified
3TitaNet-S (NME-SC)DER(%)2—Unverified
4TitaNet-M (NME-SC)DER(%)1.99—Unverified
#ModelMetricClaimedVerifiedStatus
1TitaNet-S (NME-SC)DER(%)2.22—Unverified
2TitaNet-M (NME-SC)DER(%)1.79—Unverified
3ECAPA (SC)DER(%)1.78—Unverified
4TitaNet-L (NME-SC)DER(%)1.73—Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)9.72—Unverified
2TitaNet-L (NME-SC)DER(%)1.19—Unverified
3TitaNet-M (NME-SC)DER(%)1.13—Unverified
4TitaNet-S (NME-SC)DER(%)1.11—Unverified
#ModelMetricClaimedVerifiedStatus
1Baseline (the best result in the literature as of Oct.2019)DER(%)11.2—Unverified
2pyannote (MFCC)DER(%)10.5—Unverified
3pyannote (waveform)DER(%)9.9—Unverified
#ModelMetricClaimedVerifiedStatus
1BaselineDER(%)7.7—Unverified
2pyannote (MFCC)DER(%)5.6—Unverified
3pyannote (waveform)DER(%)4.9—Unverified
#ModelMetricClaimedVerifiedStatus
1pyannote (MFCC)DER(%)6.3—Unverified
2pyannote (waveform)DER(%)6—Unverified
#ModelMetricClaimedVerifiedStatus
1d-vector + spectralDER(%)12.54—Unverified
2titanet-sDER(%)1.11—Unverified
#ModelMetricClaimedVerifiedStatus
1SONDDER(%)4.46—Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNN-SMLDER(%)27.3—Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNNV10.6—Unverified