SOTAVerified

Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Showing 101150 of 328 papers

TitleStatusHype
All-neural online source separation, counting, and diarization for meeting analysis0
Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding0
Compositional Embeddings: Joint Perception and Comparison of Class Label Sets0
ASR Error Correction and Domain Adaptation Using Machine Translation0
Implicit spoken language diarization0
HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification0
Compositional Embeddings for Multi-Label One-Shot Learning0
ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings0
Aligning Speakers: Evaluating and Visualizing Text-based Diarization Using Efficient Multiple Sequence Alignment (Extended Version)0
Home monitoring for frailty detection through sound and speaker diarization analysis0
Guided Speaker Embedding0
Implicit Self-supervised Language Representation for Spoken Language Diarization0
GIST-AiTeR Speaker Diarization System for VoxCeleb Speaker Recognition Challenge (VoxSRC) 20230
Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm0
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling0
Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications0
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation0
Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization0
Improving Transformer-based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads0
Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings0
Indigenous language technologies in Canada: Assessment, challenges, and successes0
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization0
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification0
Interrelate Training and Searching: A Unified Online Clustering Framework for Speaker Diarization0
Investigating Confidence Estimation Measures for Speaker Diarization0
Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains0
Generation of Speaker Representations Using Heterogeneous Training Batch Assembly0
Chronological Self-Training for Real-Time Speaker Diarization0
From Modular to End-to-End Speaker Diarization0
CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings0
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification.0
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization0
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection0
Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech0
BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers0
A Review of Speaker Diarization: Recent Advances with Deep Learning0
Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization0
Exploring Speaker Diarization with Mixture of Experts0
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis0
Bi-LSTM Scoring Based Similarity Measurement with Agglomerative Hierarchical Clustering (AHC) for Speaker Diarization0
A Review of Common Online Speaker Diarization Methods0
AG-LSEC: Audio Grounded Lexical Speaker Error Correction0
A Comparative Study on Multichannel Speaker-Attributed Automatic Speech Recognition in Multi-party Meetings0
End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap Detection0
End-to-End Speaker Diarization as Post-Processing0
A Reinforcement Learning Framework for Online Speaker Diarization0
End-to-end Online Speaker Diarization with Target Speaker Tracking0
Bazinga! A Dataset for Multi-Party Dialogues Structuring0
A Real-time Speaker Diarization System Based on Spatial Spectrum0
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond0
Show:102550
← PrevPage 3 of 7Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1COS+NJW-SC (Oracle SAD)DER(%)24.05Unverified
2EENDDER(%)23.07Unverified
3COS+AHC (Oracle SAD)DER(%)21.13Unverified
4SA-EEND (2-spk, no-adapt)DER(%)12.66Unverified
5EEND-OLADER(%)12.57Unverified
6SA-EEND (2-spk, adapted)DER(%)10.76Unverified
7TOLDDER(%)10.14Unverified
8COS+B-SC (Oracle SAD)DER(ig olp)8.78Unverified
9PLDA+AHC (Oracle SAD)DER(ig olp)8.39Unverified
10COS+NME-SC (Oracle SAD)DER(ig olp)7.29Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)8.39Unverified
2TitaNet-L (NME-SC)DER(%)6.73Unverified
3TitaNet-M (NME-SC)DER(%)6.47Unverified
4TitaNet-S (NME-SC)DER(%)6.37Unverified
5x-vector (MCGAN)DER(%)5.73Unverified
#ModelMetricClaimedVerifiedStatus
1ECAPA (SC)DER(%)2.36Unverified
2TitaNet-L (NME-SC)DER(%)2.03Unverified
3TitaNet-S (NME-SC)DER(%)2Unverified
4TitaNet-M (NME-SC)DER(%)1.99Unverified
#ModelMetricClaimedVerifiedStatus
1TitaNet-S (NME-SC)DER(%)2.22Unverified
2TitaNet-M (NME-SC)DER(%)1.79Unverified
3ECAPA (SC)DER(%)1.78Unverified
4TitaNet-L (NME-SC)DER(%)1.73Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)9.72Unverified
2TitaNet-L (NME-SC)DER(%)1.19Unverified
3TitaNet-M (NME-SC)DER(%)1.13Unverified
4TitaNet-S (NME-SC)DER(%)1.11Unverified
#ModelMetricClaimedVerifiedStatus
1Baseline (the best result in the literature as of Oct.2019)DER(%)11.2Unverified
2pyannote (MFCC)DER(%)10.5Unverified
3pyannote (waveform)DER(%)9.9Unverified
#ModelMetricClaimedVerifiedStatus
1BaselineDER(%)7.7Unverified
2pyannote (MFCC)DER(%)5.6Unverified
3pyannote (waveform)DER(%)4.9Unverified
#ModelMetricClaimedVerifiedStatus
1pyannote (MFCC)DER(%)6.3Unverified
2pyannote (waveform)DER(%)6Unverified
#ModelMetricClaimedVerifiedStatus
1d-vector + spectralDER(%)12.54Unverified
2titanet-sDER(%)1.11Unverified
#ModelMetricClaimedVerifiedStatus
1SONDDER(%)4.46Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNN-SMLDER(%)27.3Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNNV10.6Unverified