SOTAVerified

Speaker Recognition

Speaker Recognition is the process of identifying or confirming the identity of a person given his speech segments.

Source: Margin Matters: Towards More Discriminative Deep Neural Network Embeddings for Speaker Recognition

Papers

Showing 76–100 of 435 papers

TitleStatusHype
The OCON model: an old but green solution for distributable supervised classification for acoustic monitoring in smart cities—0
Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample—0
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection—0
Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models—0
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels—0
oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models—0
Text-To-Speech Synthesis In The Wild—0
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings—0
The VoxCeleb Speaker Recognition Challenge: A Retrospective—0
Convexity-based Pruning of Speech Representation Models—0
Long-Term Conversation Analysis: Privacy-Utility Trade-off under Noise and Reverberation—0
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning—0
Team HYU ASML ROBOVOX SP Cup 2024 System Description—0
Phonetic Richness for Improved Automatic Speaker Verification—0
A voice and speech corpus of patients who underwent upper airway surgery in pre- and post-operative statesCode0
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation—0
We Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings—0
Prosody-Driven Privacy-Preserving Dementia DetectionCode0
Open-Source Conversational AI with SpeechBrain 1.0—0
CEC: A Noisy Label Detection Method for Speaker Recognition—0
Challenging margin-based speaker embedding extractors by using the variational information bottleneck—0
PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation—0
The Reasonable Effectiveness of Speaker Embeddings for Violence Detection—0
Fill in the Gap! Combining Self-supervised Representation Learning with Neural Audio Synthesis for Speech Inpainting—0
Speaker Characterization by means of Attention Pooling—0
Show:102550
← PrevPage 4 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1w2v2-aamEER1.88—Unverified
2WavLM+ECAPA-TDNNEER0.39—Unverified