SOTAVerified

Speaker Identification

Papers

Showing 151–200 of 248 papers

TitleStatusHype
Streaming Multi-talker Speech Recognition with Joint Speaker Identification—0
Supervised Initialization of LSTM Networks for Fundamental Frequency Detection in Noisy Speech Signals—0
Many-to-Many Voice Conversion with Out-of-Dataset Speaker Support—0
Symmetric Saliency-based Adversarial Attack To Speaker Identification—0
Test-Time Training for Speech—0
Text-based Speaker Identification on Multiparty Dialogues Using Multi-document Convolutional Neural Networks—0
Text Independent Speaker Identification System for Access Control—0
The Deterministic plus Stochastic Model of the Residual Signal and its Applications—0
The DIRHA simulated corpus—0
The exploitation of Multiple Feature Extraction Techniques for Speaker Identification in Emotional States under Disguised Voices—0
SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification Systems—0
The RATS Collection: Supporting HLT Research with Degraded Audio Data—0
TIMIT Speaker Profiling: A Comparison of Multi-task learning and Single-task learning Approaches—0
Towards Advanced Speech Signal Processing: A Statistical Perspective on Convolution-Based Architectures and its Applications—0
Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR—0
Triplet loss based embeddings for forensic speaker identification in Spanish—0
T-vectors: Weakly Supervised Speaker Identification Using Hierarchical Transformer Model—0
Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction—0
Unraveling Adversarial Examples against Speaker Identification -- Techniques for Attack Detection and Victim Model Classification—0
VAST: A Corpus of Video Annotation for Speech Technologies—0
VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-Resolution—0
Voice Privacy with Smart Digital Assistants in Educational Settings—0
Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices—0
VoxWatch: An open-set speaker recognition benchmark on VoxCeleb—0
WaBERT: A Low-resource End-to-end Model for Spoken Language Understanding and Speech-to-BERT Alignment—0
Weakly Supervised Training of Hierarchical Attention Networks for Speaker Identification—0
Weakly Supervised Training of Speaker Identification Models—0
Supervised Speaker Embedding De-Mixing in Two-Speaker Environment—0
A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement—0
Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors—0
Advanced accent/dialect identification and accentedness assessment with multi-embedding models and automatic speech recognition—0
Advanced Rich Transcription System for Estonian Speech—0
Advances in Online Audio-Visual Meeting Transcription—0
AdvEst: Adversarial Perturbation Estimation to Classify and Detect Adversarial Attacks against Speaker Identification—0
A Joint Model for Quotation Attribution and Coreference Resolution—0
A Lightweight Speaker Recognition System Using Timbre Properties—0
A Multi Level Data Fusion Approach for Speaker Identification on Telephone Speech—0
A Novel Minimum Divergence Approach to Robust Speaker Identification—0
An Unsupervised Speaker Clustering Technique based on SOM and I-vectors for Speech Recognition Systems—0
基於聽覺感知模型之類神經網路及其在語者識別上之應用 (Two-stage Attentional Auditory Model Inspired Neural Network and Its Application to Speaker Identification) [In Chinese]—0
A Preliminary Exploration with GPT-4o Voice Mode—0
A Real-time Speaker Diarization System Based on Spatial Spectrum—0
A Study of Acoustic Features in Arabic Speaker Identification under Noisy Environmental Conditions—0
A Study of Few-Shot Audio Classification—0
A Survey on Paralinguistics in Tamil Speech Processing—0
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR—0
A user study to compare two conversational assistants designed for people with hearing impairments—0
Target Speech Extraction: Independent Vector Extraction Guided by Supervised Speaker Identification—0
Can Musical Emotion Be Quantified With Neural Jitter Or Shimmer? A Novel EEG Based Study With Hindustani Classical Music—0
CASA-Based Speaker Identification Using Cascaded GMM-CNN Classifier in Noisy and Emotional Talking Conditions—0
Show:102550
← PrevPage 4 of 5Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MSM-MAETop-1 (%)96.6—Unverified
2M2D/0.6Top-1 (%)96.5—Unverified
3M2D/0.7Top-1 (%)96.3—Unverified
4M2D ratio=0.6Top-1 (%)94.8—Unverified
5AudioMAE (local)Top-1 (%)94.8—Unverified
6ATST Base (ours)Top-1 (%)94.3—Unverified
7AudioMAE (global)Top-1 (%)94.1—Unverified
8AutoSpeech (N=8,C=128)Top-1 (%)87.66—Unverified
9SSAST-FRAMETop-1 (%)80.8—Unverified
10SSAMBATop-1 (%)70.1—Unverified
#ModelMetricClaimedVerifiedStatus
1Fuzzy RetrievalTop-1 (%)67.77—Unverified
#ModelMetricClaimedVerifiedStatus
1Fuzzy RetrievalTop-1 (%)80.83—Unverified
#ModelMetricClaimedVerifiedStatus
1Fuzzy RetrievalTop-1 (%)95.13—Unverified