SOTAVerified

Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Showing 176–200 of 328 papers

TitleStatusHype
asya: Mindful verbal communication using deep learning—0
A Thousand Words are Worth More Than One Recording: NLP Based Speaker Change Point Detection—0
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR—0
Audio-video fusion strategies for active speaker detection in meetings—0
Audio-Visual Approach For Multimodal Concurrent Speaker Detection—0
Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion—0
Audiovisual speaker diarization of TV series—0
Autoapprentissage pour le regroupement en locuteurs : premi\`eres investigations (First investigations on self trained speaker diarization )—0
Auxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarization—0
Bazinga! A Dataset for Multi-Party Dialogues Structuring—0
Bi-LSTM Scoring Based Similarity Measurement with Agglomerative Hierarchical Clustering (AHC) for Speaker Diarization—0
BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers—0
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection—0
CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings—0
Chronological Self-Training for Real-Time Speaker Diarization—0
Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization—0
Compositional Embeddings for Multi-Label One-Shot Learning—0
Compositional Embeddings: Joint Perception and Comparison of Class Label Sets—0
Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding—0
Computer-assisted Speaker Diarization: How to Evaluate Human Corrections—0
Constrained speaker diarization of TV series based on visual patterns—0
Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for M2MeT Challenge—0
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions—0
Designing an Effective Metric Learning Pipeline for Speaker Diarization—0
Détection de locuteurs dans les séries TV—0
Show:102550
← PrevPage 8 of 14Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1COS+NJW-SC (Oracle SAD)DER(%)24.05—Unverified
2EENDDER(%)23.07—Unverified
3COS+AHC (Oracle SAD)DER(%)21.13—Unverified
4SA-EEND (2-spk, no-adapt)DER(%)12.66—Unverified
5EEND-OLADER(%)12.57—Unverified
6SA-EEND (2-spk, adapted)DER(%)10.76—Unverified
7TOLDDER(%)10.14—Unverified
8COS+B-SC (Oracle SAD)DER(ig olp)8.78—Unverified
9PLDA+AHC (Oracle SAD)DER(ig olp)8.39—Unverified
10COS+NME-SC (Oracle SAD)DER(ig olp)7.29—Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)8.39—Unverified
2TitaNet-L (NME-SC)DER(%)6.73—Unverified
3TitaNet-M (NME-SC)DER(%)6.47—Unverified
4TitaNet-S (NME-SC)DER(%)6.37—Unverified
5x-vector (MCGAN)DER(%)5.73—Unverified
#ModelMetricClaimedVerifiedStatus
1ECAPA (SC)DER(%)2.36—Unverified
2TitaNet-L (NME-SC)DER(%)2.03—Unverified
3TitaNet-S (NME-SC)DER(%)2—Unverified
4TitaNet-M (NME-SC)DER(%)1.99—Unverified
#ModelMetricClaimedVerifiedStatus
1TitaNet-S (NME-SC)DER(%)2.22—Unverified
2TitaNet-M (NME-SC)DER(%)1.79—Unverified
3ECAPA (SC)DER(%)1.78—Unverified
4TitaNet-L (NME-SC)DER(%)1.73—Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)9.72—Unverified
2TitaNet-L (NME-SC)DER(%)1.19—Unverified
3TitaNet-M (NME-SC)DER(%)1.13—Unverified
4TitaNet-S (NME-SC)DER(%)1.11—Unverified
#ModelMetricClaimedVerifiedStatus
1Baseline (the best result in the literature as of Oct.2019)DER(%)11.2—Unverified
2pyannote (MFCC)DER(%)10.5—Unverified
3pyannote (waveform)DER(%)9.9—Unverified
#ModelMetricClaimedVerifiedStatus
1BaselineDER(%)7.7—Unverified
2pyannote (MFCC)DER(%)5.6—Unverified
3pyannote (waveform)DER(%)4.9—Unverified
#ModelMetricClaimedVerifiedStatus
1pyannote (MFCC)DER(%)6.3—Unverified
2pyannote (waveform)DER(%)6—Unverified
#ModelMetricClaimedVerifiedStatus
1d-vector + spectralDER(%)12.54—Unverified
2titanet-sDER(%)1.11—Unverified
#ModelMetricClaimedVerifiedStatus
1SONDDER(%)4.46—Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNN-SMLDER(%)27.3—Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNNV10.6—Unverified