SOTAVerified

Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Showing 101150 of 328 papers

TitleStatusHype
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization0
Auxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarization0
基於i-vector與PLDA並使用GMM-HMM強制對位之自動語者分段標記系統 (Speaker Diarization based on I-vector PLDA Scoring and using GMM-HMM Forced Alignment) [In Chinese]0
EEND-DEMUX: End-to-End Neural Speaker Diarization via Demultiplexed Speaker Embeddings0
Chronological Self-Training for Real-Time Speaker Diarization0
Generation of Speaker Representations Using Heterogeneous Training Batch Assembly0
GIST-AiTeR Speaker Diarization System for VoxCeleb Speaker Recognition Challenge (VoxSRC) 20230
Guided Speaker Embedding0
A framework for the automatic inference of stochastic turn-taking styles0
Home monitoring for frailty detection through sound and speaker diarization analysis0
ECAPA-TDNN Embeddings for Speaker Diarization0
Implicit Self-supervised Language Representation for Spoken Language Diarization0
Implicit spoken language diarization0
Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm0
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling0
Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications0
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation0
Domain-Dependent Speaker Diarization for the Third DIHARD Challenge0
Improving Transformer-based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads0
Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings0
Indigenous language technologies in Canada: Assessment, challenges, and successes0
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization0
Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis0
Interrelate Training and Searching: A Unified Online Clustering Framework for Speaker Diarization0
Investigating Confidence Estimation Measures for Speaker Diarization0
Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains0
An Infinite Hidden Markov Model With Similarity-Biased Transitions0
A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio0
An approach to optimize inference of the DIART speaker diarization pipeline0
Multimodal Clustering with Role Induced Constraints for Speaker Diarization0
DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding0
DISPLACE Challenge: DIarization of SPeaker and LAnguage in Conversational Environments0
Autoapprentissage pour le regroupement en locuteurs : premi\`eres investigations (First investigations on self trained speaker diarization )0
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation0
Audiovisual speaker diarization of TV series0
An Experimental Review of Speaker Diarization methods with application to Two-Speaker Conversational Telephone Speech recordings0
Meta-learning for robust child-adult classification from speech0
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models0
Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion0
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset0
Long-Term Conversation Analysis: Privacy-Utility Trade-off under Noise and Reverberation0
Audio-Visual Approach For Multimodal Concurrent Speaker Detection0
An Effortless Way To Create Large-Scale Datasets For Famous Speakers0
Advances in Online Audio-Visual Meeting Transcription0
Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge0
Matics Software Suite: New Tools for Evaluation and Data Exploration0
Meeting Transcription Using Virtual Microphone Arrays0
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR0
An automated medical scribe for documenting clinical encounters0
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives0
Show:102550
← PrevPage 3 of 7Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1COS+NJW-SC (Oracle SAD)DER(%)24.05Unverified
2EENDDER(%)23.07Unverified
3COS+AHC (Oracle SAD)DER(%)21.13Unverified
4SA-EEND (2-spk, no-adapt)DER(%)12.66Unverified
5EEND-OLADER(%)12.57Unverified
6SA-EEND (2-spk, adapted)DER(%)10.76Unverified
7TOLDDER(%)10.14Unverified
8COS+B-SC (Oracle SAD)DER(ig olp)8.78Unverified
9PLDA+AHC (Oracle SAD)DER(ig olp)8.39Unverified
10COS+NME-SC (Oracle SAD)DER(ig olp)7.29Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)8.39Unverified
2TitaNet-L (NME-SC)DER(%)6.73Unverified
3TitaNet-M (NME-SC)DER(%)6.47Unverified
4TitaNet-S (NME-SC)DER(%)6.37Unverified
5x-vector (MCGAN)DER(%)5.73Unverified
#ModelMetricClaimedVerifiedStatus
1ECAPA (SC)DER(%)2.36Unverified
2TitaNet-L (NME-SC)DER(%)2.03Unverified
3TitaNet-S (NME-SC)DER(%)2Unverified
4TitaNet-M (NME-SC)DER(%)1.99Unverified
#ModelMetricClaimedVerifiedStatus
1TitaNet-S (NME-SC)DER(%)2.22Unverified
2TitaNet-M (NME-SC)DER(%)1.79Unverified
3ECAPA (SC)DER(%)1.78Unverified
4TitaNet-L (NME-SC)DER(%)1.73Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)9.72Unverified
2TitaNet-L (NME-SC)DER(%)1.19Unverified
3TitaNet-M (NME-SC)DER(%)1.13Unverified
4TitaNet-S (NME-SC)DER(%)1.11Unverified
#ModelMetricClaimedVerifiedStatus
1Baseline (the best result in the literature as of Oct.2019)DER(%)11.2Unverified
2pyannote (MFCC)DER(%)10.5Unverified
3pyannote (waveform)DER(%)9.9Unverified
#ModelMetricClaimedVerifiedStatus
1BaselineDER(%)7.7Unverified
2pyannote (MFCC)DER(%)5.6Unverified
3pyannote (waveform)DER(%)4.9Unverified
#ModelMetricClaimedVerifiedStatus
1pyannote (MFCC)DER(%)6.3Unverified
2pyannote (waveform)DER(%)6Unverified
#ModelMetricClaimedVerifiedStatus
1d-vector + spectralDER(%)12.54Unverified
2titanet-sDER(%)1.11Unverified
#ModelMetricClaimedVerifiedStatus
1SONDDER(%)4.46Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNN-SMLDER(%)27.3Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNNV10.6Unverified