SOTAVerified

Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Showing 201250 of 328 papers

TitleStatusHype
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge0
The USTC-Ximalaya system for the ICASSP 2022 multi-channel multi-party meeting transcription (M2MeT) challenge0
The Volcspeech system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge0
The xmuspeech system for multi-channel multi-party meeting transcription challenge0
Third DIHARD Challenge Evaluation Plan0
"This is Houston. Say again, please". The Behavox system for the Apollo-11 Fearless Steps Challenge (phase II)0
Three-class Overlapped Speech Detection using a Convolutional Recurrent Neural Network0
Tight integration of neural- and clustering-based diarization through deep unfolding of infinite Gaussian mixture model0
Toeplitz Inverse Covariance based Robust Speaker Clustering for Naturalistic Audio Streams0
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch0
Towards end-2-end learning for predicting behavior codes from spoken utterances in psychotherapy conversations0
Late Audio-Visual Fusion for In-The-Wild Speaker Diarization0
Towards Measuring and Scoring Speaker Diarization Fairness0
Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled Family Audio0
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders0
Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network0
Training Speaker Embedding Extractors Using Multi-Speaker Audio with Unknown Speaker Boundaries0
Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR0
Triplet Network with Attention for Speaker Diarization0
TSUP Speaker Diarization System for Conversational Short-phrase Speaker Diarization Challenge0
Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification0
Unified Audio Event Detection0
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection0
UniX-Encoder: A Universal X-Channel Speech Encoder for Ad-Hoc Microphone Array Speech Processing0
Unsupervised Adaptation of SPLDA0
Unsupervised Speaker Diarization in Distributed IoT Networks Using Federated Learning0
Unsupervised Speaker Diarization that is Agnostic to Language, Overlap-Aware, and Tuning Free0
Using Active Speaker Faces for Diarization in TV shows0
Utterance-Wise Meeting Transcription System Using Asynchronous Distributed Microphones0
UWB-NTIS Speaker Diarization System for the DIHARD II 2019 Challenge0
VOXLINGUA107: A DATASET FOR SPOKEN LANGUAGE RECOGNITION0
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering0
Weakly Supervised Training of Speaker Identification Models0
An approach to optimize inference of the DIART speaker diarization pipeline0
X-Vectors with Multi-Scale Aggregation for Speaker Diarization0
A Benchmark for Multi-speaker Anonymization0
A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio0
A Comparative Study on Multichannel Speaker-Attributed Automatic Speech Recognition in Multi-party Meetings0
Advances in Online Audio-Visual Meeting Transcription0
A framework for the automatic inference of stochastic turn-taking styles0
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond0
AG-LSEC: Audio Grounded Lexical Speaker Error Correction0
Aligning Speakers: Evaluating and Visualizing Text-based Diarization Using Efficient Multiple Sequence Alignment (Extended Version)0
All-neural online source separation, counting, and diarization for meeting analysis0
An Alternative to Low-level-Sychrony-Based Methods for Speech Detection0
An automated medical scribe for documenting clinical encounters0
An Effortless Way To Create Large-Scale Datasets For Famous Speakers0
An Experimental Review of Speaker Diarization methods with application to Two-Speaker Conversational Telephone Speech recordings0
An Infinite Hidden Markov Model With Similarity-Biased Transitions0
基於i-vector與PLDA並使用GMM-HMM強制對位之自動語者分段標記系統 (Speaker Diarization based on I-vector PLDA Scoring and using GMM-HMM Forced Alignment) [In Chinese]0
Show:102550
← PrevPage 5 of 7Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1COS+NJW-SC (Oracle SAD)DER(%)24.05Unverified
2EENDDER(%)23.07Unverified
3COS+AHC (Oracle SAD)DER(%)21.13Unverified
4SA-EEND (2-spk, no-adapt)DER(%)12.66Unverified
5EEND-OLADER(%)12.57Unverified
6SA-EEND (2-spk, adapted)DER(%)10.76Unverified
7TOLDDER(%)10.14Unverified
8COS+B-SC (Oracle SAD)DER(ig olp)8.78Unverified
9PLDA+AHC (Oracle SAD)DER(ig olp)8.39Unverified
10COS+NME-SC (Oracle SAD)DER(ig olp)7.29Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)8.39Unverified
2TitaNet-L (NME-SC)DER(%)6.73Unverified
3TitaNet-M (NME-SC)DER(%)6.47Unverified
4TitaNet-S (NME-SC)DER(%)6.37Unverified
5x-vector (MCGAN)DER(%)5.73Unverified
#ModelMetricClaimedVerifiedStatus
1ECAPA (SC)DER(%)2.36Unverified
2TitaNet-L (NME-SC)DER(%)2.03Unverified
3TitaNet-S (NME-SC)DER(%)2Unverified
4TitaNet-M (NME-SC)DER(%)1.99Unverified
#ModelMetricClaimedVerifiedStatus
1TitaNet-S (NME-SC)DER(%)2.22Unverified
2TitaNet-M (NME-SC)DER(%)1.79Unverified
3ECAPA (SC)DER(%)1.78Unverified
4TitaNet-L (NME-SC)DER(%)1.73Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)9.72Unverified
2TitaNet-L (NME-SC)DER(%)1.19Unverified
3TitaNet-M (NME-SC)DER(%)1.13Unverified
4TitaNet-S (NME-SC)DER(%)1.11Unverified
#ModelMetricClaimedVerifiedStatus
1Baseline (the best result in the literature as of Oct.2019)DER(%)11.2Unverified
2pyannote (MFCC)DER(%)10.5Unverified
3pyannote (waveform)DER(%)9.9Unverified
#ModelMetricClaimedVerifiedStatus
1BaselineDER(%)7.7Unverified
2pyannote (MFCC)DER(%)5.6Unverified
3pyannote (waveform)DER(%)4.9Unverified
#ModelMetricClaimedVerifiedStatus
1pyannote (MFCC)DER(%)6.3Unverified
2pyannote (waveform)DER(%)6Unverified
#ModelMetricClaimedVerifiedStatus
1d-vector + spectralDER(%)12.54Unverified
2titanet-sDER(%)1.11Unverified
#ModelMetricClaimedVerifiedStatus
1SONDDER(%)4.46Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNN-SMLDER(%)27.3Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNNV10.6Unverified