Speech Synthesis

Speech synthesis is the task of generating speech from some other modality like text, lip movements etc.

Please note that the leaderboards here are not really comparable between studies - as they use mean opinion score as a metric and collect different samples from Amazon Mechnical Turk.

( Image credit: WaveNet: A generative model for raw audio )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 201–250 of 1249 papers

Title	Date	Tasks	Status	Hype
PhaseAug: A Differentiable Augmentation for Speech Synthesis to Simulate One-to-Many Mapping	Nov 8, 2022	Generative Adversarial NetworkSpeech Synthesis	CodeCode Available	1
Byakto Speech: Real-time long speech synthesis with convolutional neural network: Transfer learning from English to Bangla	May 31, 2021	Deep Learningspeech-recognition	CodeCode Available	1
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models	May 21, 2025	Bayesian OptimizationSpeech Synthesis	CodeCode Available	1
Digital Voicing of Silent Speech	Oct 6, 2020	Electromyography (EMG)Speech Synthesis	CodeCode Available	1
RAD-TTS: Parallel Flow-Based TTS with Robust Alignment Learning and Diverse Synthesis	Jun 2, 2021	DiversityRhythm	CodeCode Available	1
Disentanglement in a GAN for Unconditional Speech Synthesis	Jul 4, 2023	DisentanglementGenerative Adversarial Network	CodeCode Available	1
Effective Deep Learning Models for Automatic Diacritization of Arabic Text	Nov 1, 2020	Arabic Text DiacritizationDecoder	CodeCode Available	1
Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis	Jan 11, 2025	AttributeBenchmarking	CodeCode Available	1
Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme	Sep 28, 2021	Speech SynthesisVoice Conversion	CodeCode Available	1
Diffusion-Based Mel-Spectrogram Enhancement for Personalized Speech Synthesis with Found Data	May 18, 2023	Speech EnhancementSpeech Synthesis	CodeCode Available	1
DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding	Aug 15, 2023	Speech Synthesis	CodeCode Available	1
SAMO: Speaker Attractor Multi-Center One-Class Learning for Voice Anti-Spoofing	Nov 4, 2022	DiversitySpeaker Verification	CodeCode Available	1
CDPAM: Contrastive learning for perceptual audio similarity	Feb 9, 2021	Contrastive LearningSpeech Enhancement	CodeCode Available	1
A Resource for Computational Experiments on Mapudungun	Dec 4, 2019	Machine Translationspeech-recognition	CodeCode Available	1
ADAPTERMIX: Exploring the Efficacy of Mixture of Adapters for Low-Resource TTS Adaptation	May 29, 2023	Speech Synthesistext-to-speech	CodeCode Available	1
Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding	Aug 12, 2020	Speech Synthesistext-to-speech	CodeCode Available	1
SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers	Nov 15, 2024	Image GenerationSpeech Synthesis	CodeCode Available	1
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training	Jul 31, 2023	DenoisingExpressive Speech Synthesis	CodeCode Available	1
DiffWave: A Versatile Diffusion Model for Audio Synthesis	Sep 21, 2020	Audio SynthesisDiversity	CodeCode Available	1
Articulation GAN: Unsupervised modeling of articulatory learning	Oct 27, 2022	Generative Adversarial NetworkSpeech Synthesis	CodeCode Available	1
Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis	Jun 8, 2019	Expressive Speech SynthesisSpeech Synthesis	CodeCode Available	1
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing	Oct 14, 2021	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	1
Deep Speech Synthesis from MRI-Based Articulatory Representations	Jul 5, 2023	Computational EfficiencyDenoising	CodeCode Available	1
Deep Speech Synthesis from Articulatory Representations	Sep 13, 2022	Speech Synthesis	CodeCode Available	1
Generative Expressive Conversational Speech Synthesis	Jul 31, 2024	Speech Synthesis	CodeCode Available	1
Detection of Prosodic Boundaries in Speech Using Wav2Vec 2.0	Sep 29, 2022	SentenceSpeech Synthesis	CodeCode Available	1
ArTST: Arabic Text and Speech Transformer	Oct 25, 2023	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	1
Deep Learning Based Assessment of Synthetic Speech Naturalness	Apr 23, 2021	Deep LearningPrediction	CodeCode Available	1
Synthesizing Speech from Intracranial Depth Electrodes using an Encoder-Decoder Framework	Nov 2, 2021	DecoderEEG	CodeCode Available	1
Tacotron: Towards End-to-End Speech Synthesis	Mar 29, 2017	Audio SynthesisSpeech Synthesis	CodeCode Available	1
Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech	Oct 8, 2021	Emotion InterpretationExpressive Speech Synthesis	CodeCode Available	1
Deep Learning Enabled Semantic Communications with Speech Recognition and Synthesis	May 9, 2022	Deep LearningSemantic Communication	CodeCode Available	1
A Spectral Energy Distance for Parallel Speech Synthesis	Aug 3, 2020	scoring ruleSpeech Synthesis	CodeCode Available	1
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet	Feb 4, 2025	Speech Synthesistext-to-speech	CodeCode Available	1
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed	Sep 23, 2022	Pitch controlSpeech Synthesis	CodeCode Available	1
Cross-modal information fusion for voice spoofing detection	Feb 1, 2023	Automatic Speech Recognitionfake voice detection	CodeCode Available	1
EfficientNet-Absolute Zero for Continuous Speech Keyword Spotting	Dec 31, 2020	Keyword SpottingKeyword Spotting CSS	CodeCode Available	1
Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition	Mar 29, 2022	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	CodeCode Available	1
From Speaker Verification to Multispeaker Speech Synthesis, Deep Transfer with Feedback Constraint	May 10, 2020	Speaker VerificationSpeech Synthesis	CodeCode Available	1
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions	Dec 16, 2017	Speech Synthesis	CodeCode Available	1
A Machine of Few Words -- Interactive Speaker Recognition with Reinforcement Learning	Aug 7, 2020	Decision Makingreinforcement-learning	—Unverified	0
A Survey on Bridging EEG Signals and Generative AI: From Image and Text to Beyond	Feb 17, 2025	Contrastive LearningEEG	—Unverified	0
Constructive Interaction for Talking about Interesting Topics	May 1, 2012	ManagementSpeech Recognition	—Unverified	0
Construction of English-French Multimodal Affective Conversational Corpus from TV Dramas	May 1, 2018	Emotion RecognitionSpeech Recognition	—Unverified	0
A Survey of Voice Translation Methodologies - Acoustic Dialect Decoder	Oct 13, 2016	DecoderSentence	—Unverified	0
Alternate Endings: Improving Prosody for Incremental Neural TTS with Predicted Future Text Input	Feb 19, 2021	Language ModelingLanguage Modelling	—Unverified	0
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding	Oct 17, 2024	Speech Synthesis	—Unverified	0
Conditioning Sequence-to-sequence Networks with Learned Activations	Sep 29, 2021	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified	0
Contextual Expressive Text-to-Speech	Nov 26, 2022	Speech Synthesistext-to-speech	—Unverified	0
Conditional Spoken Digit Generation with StyleGAN	Sep 15, 2020	Image GenerationSpeech Synthesis	—Unverified	0

Show:10 25 50

← PrevPage 5 of 25Next →

All datasets LibriTTS North American English LJSpeech Mandarin Chinese Blizzard Challenge 2013

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	PeriodWave-Turbo-L	PESQ	4.45	—	Unverified
2	BigVGAN-v2	PESQ	4.36	—	Unverified
3	EVA-GAN-big	PESQ	4.35	—	Unverified
4	PeriodWave + FreeU	PESQ	4.25	—	Unverified
5	RFWave	PESQ	4.23	—	Unverified
6	BigVSAN (w/ snakebeta)	PESQ	4.12	—	Unverified
7	BigVSAN	PESQ	4.12	—	Unverified
8	EVA-GAN-base	PESQ	4.03	—	Unverified
9	BigVGAN	PESQ	4.03	—	Unverified
10	Vocos	PESQ	3.7	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Tacotron 2	Mean Opinion Score	4.53	—	Unverified
2	WaveNet (Linguistic)	Mean Opinion Score	4.34	—	Unverified
3	WaveNet (L+F)	Mean Opinion Score	4.21	—	Unverified
4	Tacotron	Mean Opinion Score	4	—	Unverified
5	HMM-driven concatenative	Mean Opinion Score	3.86	—	Unverified
6	LSTM-RNN parametric	Mean Opinion Score	3.67	—	Unverified
7	means	Mean Opinion Score	0	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	BDDM vocoder	Mean Opinion Score	4.48	—	Unverified
2	DiffWave LARGE	Mean Opinion Score	4.44	—	Unverified
3	Neural HMM	Mean Opinion Score	3.24	—	Unverified
4	Neural HMM Ablation with 1 state per phone	Mean Opinion Score	2.68	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	WaveNet (L+F)	Mean Opinion Score	4.08	—	Unverified
2	LSTM-RNN parametric	Mean Opinion Score	3.79	—	Unverified
3	HMM-driven concatenative	Mean Opinion Score	3.47	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SampleRNN (2-tier)	NLL	1.39	—	Unverified
2	SampleRNN (3-tier)	NLL	1.39	—	Unverified