SOTAVerified

Speaker Identification

Papers

Showing 151–200 of 248 papers

TitleStatusHype
Representation Learning to Classify and Detect Adversarial Attacks against Speaker and Speech Recognition Systems—0
QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus—0
Fusion of Embeddings Networks for Robust Combination of Text Dependent and Independent Speaker Recognition—0
Graph-based Label Propagation for Semi-Supervised Speaker Identification—0
PF-Net: Personalized Filter for Speaker Recognition from Raw WaveformCode0
End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings—0
Streaming Multi-talker Speech Recognition with Joint Speaker Identification—0
End-to-End Speaker-Attributed ASR with Transformer—0
A Survey on Paralinguistics in Tamil Speech Processing—0
Voice Privacy with Smart Digital Assistants in Educational Settings—0
Triplet loss based embeddings for forensic speaker identification in Spanish—0
CASA-Based Speaker Identification Using Cascaded GMM-CNN Classifier in Noisy and Emotional Talking Conditions—0
Speaker attribution with voice profiles by graph-based semi-supervised learning—0
Attention-based multi-task learning for speech-enhancement and speaker-identification in multi-speaker dialogue scenarioCode0
Hypothesis Stitcher for End-to-End Speaker-attributed ASR on Long-form Multi-talker Recordings—0
A Study of Few-Shot Audio Classification—0
How Far Are We from Robust Voice Conversion: A Survey—0
Multi-Modal Emotion Detection with Transfer Learning—0
T-vectors: Weakly Supervised Speaker Identification Using Hierarchical Transformer Model—0
Compositional embedding models for speaker identification and diarization with simultaneous speech from 2+ speakersCode0
Contrastive Learning of General-Purpose Audio RepresentationsCode0
A Lightweight Speaker Recognition System Using Timbre Properties—0
Remarks on Optimal Scores for Speaker Recognition—0
SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification Systems—0
Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers—0
Integrated Replay Spoofing-aware Text-independent Speaker Verification—0
Speaker and Posture Classification using Instantaneous Intraspeech Breathing Features—0
Identify Speakers in Cocktail Parties with End-to-End AttentionCode0
Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio RepresentationCode0
Weakly Supervised Training of Hierarchical Attention Networks for Speaker Identification—0
Speaker Recognition in Bengali Language from Nonlinear Features—0
End-to-end Recurrent Denoising Autoencoder Embeddings for Speaker Identification—0
Deep Neural Networks for Automatic Speech Processing: A Survey from Large Corpora to Limited Data—0
Speaker Identification using EEG—0
Multi-Task Learning with Auxiliary Speaker Identification for Conversational Emotion Recognition—0
Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention—0
Supervised Speaker Embedding De-Mixing in Two-Speaker Environment—0
Robust Speaker Recognition Using Speech Enhancement And Attention Model—0
The Deterministic plus Stochastic Model of the Residual Signal and its Applications—0
Advances in Online Audio-Visual Meeting Transcription—0
Privacy-Preserving Adversarial Representation Learning in ASR: Reality or Illusion?—0
Supervised Initialization of LSTM Networks for Fundamental Frequency Detection in Noisy Speech Signals—0
Reducing audio membership inference attack accuracy to chance: 4 defenses—0
Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors—0
Delving into VoxCeleb: environment invariant speaker recognition—0
Word-level Embeddings for Cross-Task Transfer Learning in Speech ProcessingCode0
H-VECTORS: Utterance-level Speaker Embedding Using A Hierarchical Attention Model—0
Latent space representation for multi-target speaker detection and identification with a sparse dataset using Triplet neural networksCode0
Emirati-Accented Speaker Identification in Stressful Talking Conditions—0
Improving Noise Robustness In Speaker Identification Using A Two-Stage Attention Model—0
Show:102550
← PrevPage 4 of 5Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MSM-MAETop-1 (%)96.6—Unverified
2M2D/0.6Top-1 (%)96.5—Unverified
3M2D/0.7Top-1 (%)96.3—Unverified
4M2D ratio=0.6Top-1 (%)94.8—Unverified
5AudioMAE (local)Top-1 (%)94.8—Unverified
6ATST Base (ours)Top-1 (%)94.3—Unverified
7AudioMAE (global)Top-1 (%)94.1—Unverified
8AutoSpeech (N=8,C=128)Top-1 (%)87.66—Unverified
9SSAST-FRAMETop-1 (%)80.8—Unverified
10SSAMBATop-1 (%)70.1—Unverified
#ModelMetricClaimedVerifiedStatus
1Fuzzy RetrievalTop-1 (%)67.77—Unverified
#ModelMetricClaimedVerifiedStatus
1Fuzzy RetrievalTop-1 (%)80.83—Unverified
#ModelMetricClaimedVerifiedStatus
1Fuzzy RetrievalTop-1 (%)95.13—Unverified