SOTAVerified

Automatic Speech Recognition

Papers

Showing 19512000 of 3174 papers

TitleStatusHype
Visual Information Matters for ASR Error Correction0
Visualizing Automatic Speech Recognition -- Means for a Better Understanding?0
VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis0
Voice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer0
Voice Privacy with Smart Digital Assistants in Educational Settings0
Voice Quality and Pitch Features in Transformer-Based Speech Recognition0
Voice Query Auto Completion0
VoxArabica: A Robust Dialect-Aware Arabic Speech Recognition System0
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka0
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing0
WaBERT: A Low-resource End-to-end Model for Spoken Language Understanding and Speech-to-BERT Alignment0
Warped Language Models for Noise Robust Language Understanding0
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR0
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning0
wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts0
Wav2vec-S: Semi-Supervised Pre-Training for Low-Resource ASR0
Wav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition0
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models0
W-CTC: a Connectionist Temporal Classification Loss with Wild Cards0
Weak Alignment Supervision from Hybrid Model Improves End-to-end ASR0
Weak-Attention Suppression For Transformer Based Speech Recognition0
Weakly Supervised Construction of ASR Systems with Massive Video Data0
Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition0
Weighted-Sampling Audio Adversarial Example Attack0
WER-BERT: Automatic WER Estimation with BERT in a Balanced Ordinal Classification Paradigm0
WERd: Using Social Text Spelling Variants for Evaluating Dialectal Speech Recognition0
WER we are and WER we think we are0
WER We Stand: Benchmarking Urdu ASR Models0
WEST: Word Encoded Sequence Transducers0
What Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model0
What has LeBenchmark Learnt about French Syntax?0
What is lost in Normalization? Exploring Pitfalls in Multilingual ASR Model Evaluations0
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation0
Where are we in Named Entity Recognition from Speech?0
Where are we in semantic concept extraction for Spoken Language Understanding?0
Which French speech recognition system for assistant robots?0
Whisper Finetuning on Nepali Language0
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language0
Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges0
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers0
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages0
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision0
Whither the Priors for (Vocal) Interactivity?0
Who Are We Talking About? Handling Person Names in Speech Translation0
Who Are We Talking About? Handling Person Names in Speech Translation0
Who Needs Decoders? Efficient Estimation of Sequence-level Attributes0
Why Does Decentralized Training Outperform Synchronous Training In The Large Batch Setting?0
Wiki-En-ASR-Adapt: Large-scale synthetic dataset for English ASR Customization0
Without Further Ado: Direct and Simultaneous Speech Translation by AppTek in 20210
WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition0
Show:102550
← PrevPage 40 of 64Next →

No leaderboard results yet.