SOTAVerified

Speech-to-Speech Translation

Speech-to-speech translation (S2ST) consists on translating speech from one language to speech in another language. This can be done with a cascade of automatic speech recognition (ASR), text-to-text machine translation (MT), and text-to-speech (TTS) synthesis sub-systems, which is text-centric. Recently, works on S2ST without relying on intermediate text representation is emerging.

Papers

Showing 1–50 of 117 papers

TitleStatusHype
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs—0
S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Streaming Speech Generation—0
Phi-Omni-ST: A multimodal language model for direct speech-to-speech translation—0
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing—0
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech TranslationCode0
Language translation, and change of accent for speech-to-speech task using diffusion model—0
Using Phonemes in cascaded S2S translation pipelineCode0
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation—0
Direct Speech to Speech Translation: A Review—0
Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus—0
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM—0
Speech to Speech Translation with Translatotron: A State of the Art Review—0
High-Fidelity Simultaneous Speech-To-Speech TranslationCode5
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation—0
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation—0
Direct Speech-to-Speech Neural Machine Translation: A Survey—0
Findings of the IWSLT 2024 Evaluation Campaign—0
Phonology-Guided Speech-to-Speech Translation for African Languages—0
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens—0
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data SelectionCode0
What does it take to get state of the art in simultaneous speech-to-speech translation?—0
PolySinger: Singing-Voice to Singing-Voice Translation from English to Japanese—0
Preset-Voice Matching for Privacy Regulated Speech-to-Speech Translation Systems—0
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-SpeechCode1
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation—0
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMsCode11
NAIST Simultaneous Speech Translation System for IWSLT 2024—0
Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation—0
CTC-based Non-autoregressive Textless Speech-to-Speech TranslationCode1
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech TranslationCode2
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?—0
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task LearningCode5
Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing—0
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation—0
SimulTron: On-Device Simultaneous Speech to Speech Translation—0
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought—0
TransVIP: Speech to Speech Translation System with Voice and Isochrony PreservationCode2
CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning—0
DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech TranslationCode0
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation—0
Direct Punjabi to English speech translation using discrete units—0
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine TranslatorsCode2
A Case Study on Filtering for End-to-End Speech Translation—0
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data—0
TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation—0
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech ModelsCode1
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech RepresentationCode1
DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation—0
DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech TranslationCode1
Enhancing expressivity transfer in textless speech-to-speech translation—0
Show:102550
← PrevPage 1 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Hokkien→En (Two-pass decoding)ASR-BLEU (Dev)13.6—Unverified
2Hokkien→En (Two-stage)ASR-BLEU (Dev)12.5—Unverified
3Hokkien→En (Three-stage)ASR-BLEU (Dev)12.5—Unverified
4Hokkien→En (Single-pass decoding)ASR-BLEU (Dev)8.8—Unverified
5En→Hokkien (Two-pass decoding)ASR-BLEU (Dev)7.8—Unverified
6En→Hokkien (Three-stage)ASR-BLEU (Dev)7.5—Unverified
7En→Hokkien (Two-stage)ASR-BLEU (Dev)7.1—Unverified
8En→Hokkien (Single-pass decoding)ASR-BLEU (Dev)6.6—Unverified
#ModelMetricClaimedVerifiedStatus
1GenTranslateV2ASR-BLEU32.3—Unverified
2GenTranslateV1ASR-BLEU30.1—Unverified
3SeamlessM4T LargeV2ASR-BLEU29.4—Unverified
4SeamlessM4T LargeASR-BLEU25.8—Unverified
5AudioPaLM2ASR-BLEU24—Unverified
6WhisperV2ASR-BLEU23.5—Unverified
7SeamlessM4T MediumASR-BLEU20.4—Unverified
#ModelMetricClaimedVerifiedStatus
1SeamlessM4T LargeASR-BLEU36.5—Unverified
2SeamlessM4T MediumASR-BLEU28.1—Unverified