Speech Synthesis

Speech synthesis is the task of generating speech from some other modality like text, lip movements etc.

Please note that the leaderboards here are not really comparable between studies - as they use mean opinion score as a metric and collect different samples from Amazon Mechnical Turk.

( Image credit: WaveNet: A generative model for raw audio )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 451–500 of 1249 papers

Title	Date	Tasks	Status
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?	Jun 11, 2024	Contrastive LearningSpeech Synthesis	—Unverified
Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data	Feb 29, 2024	Representation LearningSpeech Synthesis	—Unverified
Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data	Sep 17, 2024	Speech Synthesis	—Unverified
Enhancing audio quality for expressive Neural Text-to-Speech	Aug 13, 2021	Acoustic ModellingSpeech Synthesis	—Unverified
A Preliminary Study on Mandarin-Hakka neural machine translation using small-sized data	Nov 1, 2022	Machine TranslationSpeech Synthesis	—Unverified
Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping	Sep 25, 2023	Speech Synthesistext-to-speech	—Unverified
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech	Sep 24, 2024	Emotional Speech SynthesisSpeech Synthesis	—Unverified
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning	Apr 22, 2025	Deep LearningSpeaker Verification	—Unverified
FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder	Jul 5, 2024	Generative Adversarial NetworkSpeech Synthesis	—Unverified
A Flow-Based Neural Network for Time Domain Speech Enhancement	Jun 16, 2021	Density EstimationSpeech Enhancement	—Unverified
fairseq Sˆ2: A Scalable and Integrable Speech Synthesis Toolkit	Nov 1, 2021	Speech Synthesistext-to-speech	—Unverified
Fast and Accurate Decision Trees for Natural Language Processing Tasks	Sep 1, 2017	AttributeBIG-bench Machine Learning	—Unverified
A comparison of Vietnamese Statistical Parametric Speech Synthesis Systems	May 26, 2020	GPUSpeech Synthesis	—Unverified
Fast and small footprint Hybrid HMM-HiFiGAN based system for speech synthesis in Indian languages	Feb 13, 2023	Speech Synthesistext-to-speech	—Unverified
Fast Bootstrapping of Grapheme to Phoneme System for Under-resourced Languages - Application to the Iban Language	Oct 1, 2013	Speech RecognitionSpeech Synthesis	—Unverified
Fast, Compact, and High Quality LSTM-RNN Based Statistical Parametric Speech Synthesizers for Mobile Devices	Jun 20, 2016	QuantizationSpeech Synthesis	—Unverified
A Bengali Speech Synthesizer on Android OS	Jul 1, 2012	Speech Synthesis	—Unverified
Energy-Based Models For Speech Synthesis	Oct 19, 2023	Speech Synthesis	—Unverified
Can large-scale vocoded spoofed data improve speech spoofing countermeasure with a self-supervised front end?	Sep 12, 2023	Self-Supervised LearningSpeech Synthesis	—Unverified
End-to-End Video-To-Speech Synthesis using Generative Adversarial Networks	Apr 27, 2021	Lip ReadingSpeech Synthesis	—Unverified
End-to-End Text-to-Speech using Latent Duration based on VQ-VAE	Oct 19, 2020	Speech Synthesistext-to-speech	—Unverified
Fast Spectrogram Inversion using Multi-head Convolutional Neural Networks	Aug 20, 2018	speech-recognitionSpeech Recognition	—Unverified
A Preliminary Study on Deep Learning-based Chinese Text to Taiwanese Speech Synthesis System	Sep 1, 2020	Speech Synthesis	—Unverified
CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center	May 23, 2023	Speech Synthesis	—Unverified
End-to-End Feedback Loss in Speech Chain Framework via Straight-Through Estimator	Oct 31, 2018	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified
End-to-End Emotional Speech Synthesis Using Style Tokens and Semi-Supervised Training	Jun 26, 2019	Emotional Speech SynthesisEmotion Recognition	—Unverified
Bytes are All You Need: End-to-End Multilingual Speech Recognition and Synthesis with Bytes	Nov 22, 2018	Allspeech-recognition	—Unverified
A Practical Guide to Logical Access Voice Presentation Attack Detection	Jan 10, 2022	Artifact DetectionSpeaker Verification	—Unverified
Fine-grained Noise Control for Multispeaker Speech Synthesis	Apr 11, 2022	Expressive Speech SynthesisSpeech Synthesis	—Unverified
AffectEcho: Speaker Independent and Language-Agnostic Emotion and Affect Transfer for Speech Synthesis	Aug 16, 2023	AttributeSpeech Synthesis	—Unverified
Fine-grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement	Nov 8, 2020	DisentanglementSpeech Synthesis	—Unverified
Fitting New Speakers Based on a Short Untranscribed Sample	Feb 20, 2018	Speech Synthesistext-to-speech	—Unverified
End-to-End Binaural Speech Synthesis	Jul 8, 2022	DecoderSpeech Synthesis	—Unverified
Flavored Tacotron: Conditional Learning for Prosodic-linguistic Features	Apr 8, 2021	DecoderSpeech Synthesis	—Unverified
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts	Jan 8, 2025	Speech Synthesis	—Unverified
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis	Nov 20, 2023	Speech Synthesis	—Unverified
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech	Mar 13, 2024	GPUSpeech Synthesis	—Unverified
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis	Jun 30, 2024	CPUDecoder	—Unverified
BU-TTS: An Open-Source, Bilingual Welsh-English, Text-to-Speech Corpus	Jun 1, 2022	Speech Synthesistext-to-speech	—Unverified
Applying Syntaxx2013Prosody Mapping Hypothesis and Prosodic Well-Formedness Constraints to Neural Sequence-to-Sequence Speech Synthesis	Mar 29, 2022	Speech Synthesistext-to-speech	—Unverified
Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis	Apr 10, 2025	Speech Synthesistext-to-speech	—Unverified
Forward Attention in Sequence-to-sequence Acoustic Modelling for Speech Synthesis	Jul 18, 2018	Acoustic ModellingDecoder	—Unverified
FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model	Mar 6, 2023	Language ModelingLanguage Modelling	—Unverified
From `Solved Problems' to New Challenges: A Report on LDC Activities	May 1, 2018	Dialogue ManagementLanguage Identification	—Unverified
Building Text-To-Speech Voices in the Cloud	May 1, 2012	Speech RecognitionSpeech Synthesis	—Unverified
From Flat to Feeling: A Feasibility and Impact Study on Dynamic Facial Emotions in AI-Generated Avatars	Jun 16, 2025	GPUSpeech Synthesis	—Unverified
From Speaker Identification to Affective Analysis: A Multi-Step System for Analyzing Children's Stories	Apr 1, 2014	Age EstimationSpeaker Identification	—Unverified
Empirical Analysis of Oral and Nasal Vowels of Konkani	May 17, 2023	Speech Synthesis	—Unverified
Full Attention Bidirectional Deep Learning Structure for Single Channel Speech Enhancement	Aug 27, 2021	Audio Signal ProcessingSpeech Enhancement	—Unverified
Emphasized Accent Phrase Prediction from Text for Advertisement Text-To-Speech Synthesis	Dec 1, 2014	Speech Synthesistext-to-speech	—Unverified

Show:10 25 50

← PrevPage 10 of 25Next →

All datasets LibriTTS North American English LJSpeech Mandarin Chinese Blizzard Challenge 2013

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	PeriodWave-Turbo-L	PESQ	4.45	—	Unverified
2	BigVGAN-v2	PESQ	4.36	—	Unverified
3	EVA-GAN-big	PESQ	4.35	—	Unverified
4	PeriodWave + FreeU	PESQ	4.25	—	Unverified
5	RFWave	PESQ	4.23	—	Unverified
6	BigVSAN (w/ snakebeta)	PESQ	4.12	—	Unverified
7	BigVSAN	PESQ	4.12	—	Unverified
8	EVA-GAN-base	PESQ	4.03	—	Unverified
9	BigVGAN	PESQ	4.03	—	Unverified
10	Vocos	PESQ	3.7	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Tacotron 2	Mean Opinion Score	4.53	—	Unverified
2	WaveNet (Linguistic)	Mean Opinion Score	4.34	—	Unverified
3	WaveNet (L+F)	Mean Opinion Score	4.21	—	Unverified
4	Tacotron	Mean Opinion Score	4	—	Unverified
5	HMM-driven concatenative	Mean Opinion Score	3.86	—	Unverified
6	LSTM-RNN parametric	Mean Opinion Score	3.67	—	Unverified
7	means	Mean Opinion Score	0	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	BDDM vocoder	Mean Opinion Score	4.48	—	Unverified
2	DiffWave LARGE	Mean Opinion Score	4.44	—	Unverified
3	Neural HMM	Mean Opinion Score	3.24	—	Unverified
4	Neural HMM Ablation with 1 state per phone	Mean Opinion Score	2.68	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	WaveNet (L+F)	Mean Opinion Score	4.08	—	Unverified
2	LSTM-RNN parametric	Mean Opinion Score	3.79	—	Unverified
3	HMM-driven concatenative	Mean Opinion Score	3.47	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SampleRNN (2-tier)	NLL	1.39	—	Unverified
2	SampleRNN (3-tier)	NLL	1.39	—	Unverified