SOTAVerified

Speech-to-Text Translation

Translate audio signals of speech in one language into text in a foreign language, either in an end-to-end or cascade manner.

Papers

Showing 101–146 of 146 papers

TitleStatusHype
Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?—0
Strategies for improving low resource speech to text translation relying on pre-trained ASR models—0
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection—0
Subtitles to Segmentation: Improving Low-Resource Speech-to-TextTranslation Pipelines—0
Subtitles to Segmentation: Improving Low-Resource Speech-to-Text Translation Pipelines—0
TASK AWARE MULTI-TASK LEARNING FOR SPEECH TO TEXT TASKS—0
The USFD Spoken Language Translation System for IWSLT 2014—0
Towards Measuring Fairness in AI: the Casual Conversations Dataset—0
Towards speech-to-text translation without speech recognition—0
Towards the evaluation of automatic simultaneous speech translation from a communicative perspective—0
Towards Unsupervised Speech-to-Text Translation—0
Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces—0
Unveiling the Role of Pretraining in Direct Speech Translation—0
Using of heterogeneous corpora for training of an ASR system—0
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation—0
Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling—0
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases—0
Recent Advances in Direct Speech-to-text Translation—0
Representation Purification for End-to-End Speech Translation—0
Revisiting End-to-End Speech-to-Text Translation From Scratch—0
Robust Semantic Communications for Speech Transmission—0
S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Streaming Speech Generation—0
SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation—0
Self-Supervised Representations Improve End-to-End Speech Translation—0
Simple and Effective Unsupervised Speech Translation—0
M-Adapter: Modality Adaptation for End-to-End Speech-to-Text TranslationCode0
SparQLe: Speech Queries to Text Translation Through LLMsCode0
Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text TranslationCode0
Pre-training on high-resource speech recognition improves low-resource speech-to-text translationCode0
Voices Unheard: NLP Resources and Models for Yorùbá Regional DialectsCode0
Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language ModelsCode0
Speechformer: Reducing Information Loss in Direct Speech TranslationCode0
LibriS2S: A German-English Speech-to-Speech Translation CorpusCode0
Synchronous Speech Recognition and Speech-to-Text Translation with Interactive DecodingCode0
CoVoSwitch: Machine Translation of Synthetic Code-Switched Text Based on Intonation UnitsCode0
Direct speech-to-speech translation with a sequence-to-sequence modelCode0
WACO: Word-Aligned Contrastive Learning for Speech TranslationCode0
Efficient Speech Translation with Dynamic Latent PerceiversCode0
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak SupervisionCode0
Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language UnderstandingCode0
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation EvaluationCode0
fairseq S2T: Fast Speech-to-Text Modeling with fairseqCode0
End-to-End Automatic Speech Translation of AudiobooksCode0
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation SystemCode0
Don't Discard Fixed-Window Audio Segmentation in Speech-to-Text TranslationCode0
An Empirical Study of Consistency Regularization for End-to-End Speech-to-Text TranslationCode0
Show:102550
← PrevPage 3 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Task Modulation + Multitask Learning(ASR/MT) + Data AugmentationCase-sensitive sacreBLEU28.88—Unverified
2Wav2Vec2.0+mBART+AdaptorsCase-sensitive sacreBLEU28.22—Unverified
3Transformer + Meta Learning(ASR/MT) + Data AugmentationCase-sensitive sacreBLEU27.51—Unverified
4Transformer with AdaptersCase-sensitive sacreBLEU24.63—Unverified
5Dual-decoder TransformerCase-sensitive sacreBLEU23.63—Unverified
6SpeechformerCase-sensitive sacreBLEU23.6—Unverified
7Transformer + ASR PretrainCase-sensitive sacreBLEU22.8—Unverified
8Transformer + ASR PretrainCase-sensitive sacreBLEU22.7—Unverified
#ModelMetricClaimedVerifiedStatus
1Transformer with AdaptersCase-sensitive sacreBLEU28.73—Unverified
2SpeechformerCase-sensitive sacreBLEU28.5—Unverified
3Dual-decoder TransformerCase-sensitive sacreBLEU28.12—Unverified
4Transformer + ASR Pretrain + SpecAugCase-sensitive sacreBLEU27.4—Unverified
5Transformer + ASR PretrainCase-sensitive sacreBLEU26.8—Unverified
#ModelMetricClaimedVerifiedStatus
1Dual-decoder TransformerCase-sensitive sacreBLEU33.45—Unverified
2Transformer + ASR Pretrain + SpecAugCase-sensitive sacreBLEU33.3—Unverified
3Transformer + ASR PretrainCase-sensitive sacreBLEU32.3—Unverified
#ModelMetricClaimedVerifiedStatus
1SeamlessM4T LargeBLEU30.6—Unverified
2SeamlessM4T MediumBLEU26.6—Unverified
#ModelMetricClaimedVerifiedStatus
1SeamlessM4T LargeBLEU34.1—Unverified
2SeamlessM4T MediumBLEU29.8—Unverified
#ModelMetricClaimedVerifiedStatus
1SeamlessM4T LargeBLEU21.5—Unverified
2SeamlessM4T MediumBLEU19.2—Unverified
#ModelMetricClaimedVerifiedStatus
1SeamlessM4T LargeBLEU24—Unverified
2SeamlessM4T MediumBLEU20.9—Unverified
#ModelMetricClaimedVerifiedStatus
1Transformer + ASR Pretrain + SpecAugCase-insensitive sacreBLEU17.2—Unverified
2Transformer + ASR PretrainCase-insensitive sacreBLEU16.5—Unverified
#ModelMetricClaimedVerifiedStatus
1MediBeng Whisper TinyBleu0.98—Unverified
2Whisper TinyBleu0.3—Unverified
#ModelMetricClaimedVerifiedStatus
1Transformer with AdaptersSacreBLEU26.61—Unverified
2Dual-decoder TransformerSacreBLEU25.62—Unverified
#ModelMetricClaimedVerifiedStatus
1SpeechformerCase-sensitive sacreBLEU27.7—Unverified