SOTAVerified

Speech-to-Text

Papers

Showing 251–275 of 403 papers

TitleStatusHype
VR-GPT: Visual Language Model for Intelligent Virtual Reality Applications—0
WaBERT: A Low-resource End-to-end Model for Spoken Language Understanding and Speech-to-BERT Alignment—0
wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts—0
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition—0
WER-BERT: Automatic WER Estimation with BERT in a Balanced Ordinal Classification Paradigm—0
What shall we do with an hour of data? Speech recognition for the un- and under-served languages of Common Voice—0
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation—0
Which French speech recognition system for assistant robots?—0
Whisper Finetuning on Nepali Language—0
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages—0
With One Voice: Composing a Travel Voice Assistant from Re-purposed Models—0
Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question Answering—0
XTREME-S: Evaluating Cross-lingual Speech Representations—0
Learning Adaptive Segmentation Policy for End-to-End Simultaneous Translation—0
Learnings from Technological Interventions in a Low Resource Language: A Case-Study on Gondi—0
Leveraging Virtual Reality and AI Tutoring for Language Learning: A Case Study of a Virtual Campus Environment with OpenAI GPT Integration with Unity 3D—0
Leveraging Weakly Supervised Data to Improve End-to-End Speech-to-Text Translation—0
LIA-RAG: a system based on graphs and divergence of probabilities applied to Speech-To-Text Summarization—0
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect—0
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization—0
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation—0
Low-Resource Speech-to-Text Translation—0
M3ST: Mix at Three Levels for Speech Translation—0
MAM: Masked Acoustic Modeling for End-to-End Speech-to-Text Translation—0
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction—0
Show:102550
← PrevPage 11 of 17Next →

No leaderboard results yet.