SOTAVerified

Speaker Diarization

Speaker Diarization is the task of segmenting and co-indexing audio recordings by speaker. The way the task is commonly defined, the goal is not to identify known speakers, but to co-index segments that are attributed to the same speaker; in other words, diarization implies finding speaker boundaries and grouping segments that belong to the same speaker, and, as a by-product, determining the number of distinct speakers. In combination with speech recognition, diarization enables speaker-attributed speech-to-text transcription.

Source: Improving Diarization Robustness using Diversification, Randomization and the DOVER Algorithm

Papers

Showing 126–150 of 328 papers

TitleStatusHype
Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications—0
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives—0
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection—0
The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models—0
Spatial-Temporal Activity-Informed Diarization and Separation—0
End-to-End Supervised Hierarchical Graph Clustering for Speaker DiarizationCode0
NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription—0
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization—0
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech RepresentationCode0
Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification—0
Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition—0
EEND-DEMUX: End-to-End Neural Speaker Diarization via Demultiplexed Speaker Embeddings—0
Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization—0
Summary of the DISPLACE Challenge 2023 -- DIarization of SPeaker and LAnguage in Conversational Environments—0
UniX-Encoder: A Universal X-Channel Speech Encoder for Ad-Hoc Microphone Array Speech Processing—0
EmoDiarize: Speaker Diarization and Emotion Identification from Speech Signals using Convolutional Neural Networks—0
Powerset multi-class cross entropy loss for neural speaker diarization—0
The CHiME-7 Challenge: System Description and Performance of NeMo Team's DASR System—0
Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation—0
End-to-end Online Speaker Diarization with Target Speaker Tracking—0
One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition—0
NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization—0
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation—0
Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network—0
Aligning Speakers: Evaluating and Visualizing Text-based Diarization Using Efficient Multiple Sequence Alignment (Extended Version)—0
Show:102550
← PrevPage 6 of 14Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1COS+NJW-SC (Oracle SAD)DER(%)24.05—Unverified
2EENDDER(%)23.07—Unverified
3COS+AHC (Oracle SAD)DER(%)21.13—Unverified
4SA-EEND (2-spk, no-adapt)DER(%)12.66—Unverified
5EEND-OLADER(%)12.57—Unverified
6SA-EEND (2-spk, adapted)DER(%)10.76—Unverified
7TOLDDER(%)10.14—Unverified
8COS+B-SC (Oracle SAD)DER(ig olp)8.78—Unverified
9PLDA+AHC (Oracle SAD)DER(ig olp)8.39—Unverified
10COS+NME-SC (Oracle SAD)DER(ig olp)7.29—Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)8.39—Unverified
2TitaNet-L (NME-SC)DER(%)6.73—Unverified
3TitaNet-M (NME-SC)DER(%)6.47—Unverified
4TitaNet-S (NME-SC)DER(%)6.37—Unverified
5x-vector (MCGAN)DER(%)5.73—Unverified
#ModelMetricClaimedVerifiedStatus
1ECAPA (SC)DER(%)2.36—Unverified
2TitaNet-L (NME-SC)DER(%)2.03—Unverified
3TitaNet-S (NME-SC)DER(%)2—Unverified
4TitaNet-M (NME-SC)DER(%)1.99—Unverified
#ModelMetricClaimedVerifiedStatus
1TitaNet-S (NME-SC)DER(%)2.22—Unverified
2TitaNet-M (NME-SC)DER(%)1.79—Unverified
3ECAPA (SC)DER(%)1.78—Unverified
4TitaNet-L (NME-SC)DER(%)1.73—Unverified
#ModelMetricClaimedVerifiedStatus
1x-vector (PLDA + AHC)DER(%)9.72—Unverified
2TitaNet-L (NME-SC)DER(%)1.19—Unverified
3TitaNet-M (NME-SC)DER(%)1.13—Unverified
4TitaNet-S (NME-SC)DER(%)1.11—Unverified
#ModelMetricClaimedVerifiedStatus
1Baseline (the best result in the literature as of Oct.2019)DER(%)11.2—Unverified
2pyannote (MFCC)DER(%)10.5—Unverified
3pyannote (waveform)DER(%)9.9—Unverified
#ModelMetricClaimedVerifiedStatus
1BaselineDER(%)7.7—Unverified
2pyannote (MFCC)DER(%)5.6—Unverified
3pyannote (waveform)DER(%)4.9—Unverified
#ModelMetricClaimedVerifiedStatus
1pyannote (MFCC)DER(%)6.3—Unverified
2pyannote (waveform)DER(%)6—Unverified
#ModelMetricClaimedVerifiedStatus
1d-vector + spectralDER(%)12.54—Unverified
2titanet-sDER(%)1.11—Unverified
#ModelMetricClaimedVerifiedStatus
1SONDDER(%)4.46—Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNN-SMLDER(%)27.3—Unverified
#ModelMetricClaimedVerifiedStatus
1UIS-RNNV10.6—Unverified