SOTAVerified

Automatic Speech Recognition

Papers

Showing 25012550 of 3174 papers

TitleStatusHype
Wav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition0
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models0
W-CTC: a Connectionist Temporal Classification Loss with Wild Cards0
Weak Alignment Supervision from Hybrid Model Improves End-to-end ASR0
Weak-Attention Suppression For Transformer Based Speech Recognition0
Weakly Supervised Construction of ASR Systems with Massive Video Data0
Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition0
Weighted-Sampling Audio Adversarial Example Attack0
WER-BERT: Automatic WER Estimation with BERT in a Balanced Ordinal Classification Paradigm0
WERd: Using Social Text Spelling Variants for Evaluating Dialectal Speech Recognition0
WER we are and WER we think we are0
WER We Stand: Benchmarking Urdu ASR Models0
WEST: Word Encoded Sequence Transducers0
What Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model0
What has LeBenchmark Learnt about French Syntax?0
What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations0
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation0
Where are we in Named Entity Recognition from Speech?0
Where are we in semantic concept extraction for Spoken Language Understanding?0
Which French speech recognition system for assistant robots?0
Whisper Finetuning on Nepali Language0
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language0
Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges0
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers0
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages0
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision0
Whither the Priors for (Vocal) Interactivity?0
Who Are We Talking About? Handling Person Names in Speech Translation0
Who Are We Talking About? Handling Person Names in Speech Translation0
Who Needs Decoders? Efficient Estimation of Sequence-level Attributes0
Why Does Decentralized Training Outperform Synchronous Training In The Large Batch Setting?0
Wiki-En-ASR-Adapt: Large-scale synthetic dataset for English ASR Customization0
Without Further Ado: Direct and Simultaneous Speech Translation by AppTek in 20210
WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition0
Word-Free Spoken Language Understanding for Mandarin-Chinese0
Word-level confidence estimation for RNN transducers0
Word Level Timestamp Generation for Automatic Speech Recognition and Translation0
Word Order Does Not Matter For Speech Recognition0
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese0
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models0
You Do Not Need More Data: Improving End-To-End Speech Recognition by Text-To-Speech Data Augmentation0
You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish0
Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication0
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus0
Zero-resource Speech Translation and Recognition with LLMs0
Zero-Shot Automatic Pronunciation Assessment0
Zero-Shot Cross-lingual Aphasia Detection using Automatic Speech Recognition0
Zero-shot Disfluency Detection for Indian Languages0
Zero-Shot Joint Modeling of Multiple Spoken-Text-Style Conversion Tasks using Switching Tokens0
Zero-shot Speech Translation0
Show:102550
← PrevPage 51 of 64Next →

No leaderboard results yet.