Speech Synthesis

Speech synthesis is the task of generating speech from some other modality like text, lip movements etc.

Please note that the leaderboards here are not really comparable between studies - as they use mean opinion score as a metric and collect different samples from Amazon Mechnical Turk.

( Image credit: WaveNet: A generative model for raw audio )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–100 of 1249 papers

Title	Date	Tasks	Status	Hype
RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction	Mar 8, 2024	Audio GenerationComputational Efficiency	CodeCode Available	2
EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks	Jan 31, 2024	Audio GenerationSpeech Synthesis	CodeCode Available	2
Generative Adversarial Training for Text-to-Speech Synthesis Based on Raw Phonetic Input and Explicit Prosody Modelling	Oct 14, 2023	Speech Synthesistext-to-speech	CodeCode Available	2
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT	Oct 7, 2023	Audio captioningAutomatic Speech Recognition	CodeCode Available	2
P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting	Sep 22, 2023	DecoderSpeech Synthesis	CodeCode Available	2
HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform	Sep 18, 2023	Speech Synthesis	CodeCode Available	2
FunCodec: A Fundamental, Reproducible and Integrable Open-source Toolkit for Neural Speech Codec	Sep 14, 2023	Automatic Speech Recognitionspeech-recognition	CodeCode Available	2
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network	Sep 6, 2023	Generative Adversarial NetworkSpeech Synthesis	CodeCode Available	2
CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency Model	May 11, 2023	DenoisingGPU	CodeCode Available	2
Source-Filter-Based Generative Adversarial Neural Vocoder for High Fidelity Speech Synthesis	Apr 26, 2023	Speech Synthesistext-to-speech	CodeCode Available	2
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers	Apr 18, 2023	In-Context LearningSpeech Synthesis	CodeCode Available	2
A Vector Quantized Approach for Text to Speech Synthesis on Real-World Spontaneous Speech	Feb 8, 2023	Code GenerationDiversity	CodeCode Available	2
Towards Building Text-To-Speech Systems for the Next Billion Users	Nov 17, 2022	DiversitySpeech Synthesis	CodeCode Available	2
StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis	May 30, 2022	Data AugmentationSelf-Supervised Learning	CodeCode Available	2
GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech	May 15, 2022	Speech SynthesisStyle Transfer	CodeCode Available	2
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality	May 9, 2022	SentenceSpeech Synthesis	CodeCode Available	2
FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis	Apr 21, 2022	DenoisingGPU	CodeCode Available	2
BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis	Mar 25, 2022	Image GenerationSpeech Synthesis	CodeCode Available	2
iSTFTNet: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform	Mar 4, 2022	Speech Synthesistext-to-speech	CodeCode Available	2
Generative Modeling for Low Dimensional Speech Attributes with Neural Spline Flows	Mar 3, 2022	Speech Synthesistext-to-speech	CodeCode Available	2
Conditional Diffusion Probabilistic Model for Speech Enhancement	Feb 10, 2022	modelSpeech Enhancement	CodeCode Available	2
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis	Oct 12, 2020	CPUGPU	CodeCode Available	2
Improving Opus Low Bit Rate Quality with Neural Speech Synthesis	Aug 10, 2020	DecoderSpeech Synthesis	CodeCode Available	2
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram	Oct 25, 2019	Generative Adversarial NetworkGPU	CodeCode Available	2
Using Speech Synthesis to Train End-to-End Spoken Language Understanding Models	Oct 21, 2019	Data AugmentationNatural Language Understanding	CodeCode Available	2
FastSpeech: Fast, Robust and Controllable Text to Speech	May 22, 2019	DecoderSpeech Synthesis	CodeCode Available	2
A Real-Time Wideband Neural Vocoder at 1.6 kb/s Using LPCNet	Mar 28, 2019	Speech Synthesis	CodeCode Available	2
LPCNet: Improving Neural Speech Synthesis Through Linear Prediction	Oct 28, 2018	PredictionSpeech Synthesis	CodeCode Available	2
Neural Speech Synthesis with Transformer Network	Sep 19, 2018	DecoderMachine Translation	CodeCode Available	2
Efficient Neural Audio Synthesis	Feb 23, 2018	Audio SynthesisCPU	CodeCode Available	2
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems	Jun 19, 2025	BenchmarkingDescriptive	CodeCode Available	1
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models	May 21, 2025	Bayesian OptimizationSpeech Synthesis	CodeCode Available	1
SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis	Apr 14, 2025	Face SwappingSpeech Synthesis	CodeCode Available	1
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet	Feb 4, 2025	Speech Synthesistext-to-speech	CodeCode Available	1
Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis	Jan 11, 2025	AttributeBenchmarking	CodeCode Available	1
AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder	Jan 9, 2025	Pitch ClassificationPitch control	CodeCode Available	1
Region-Based Optimization in Continual Learning for Audio Deepfake Detection	Dec 16, 2024	Audio Deepfake DetectionContinual Learning	CodeCode Available	1
SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers	Nov 15, 2024	Image GenerationSpeech Synthesis	CodeCode Available	1
Mitigating Unauthorized Speech Synthesis for Voice Protection	Oct 28, 2024	Data AugmentationFace Swapping	CodeCode Available	1
STTATTS: Unified Speech-To-Text And Text-To-Speech Model	Oct 24, 2024	Multi-Task Learningspeech-recognition	CodeCode Available	1
PRESENT: Zero-Shot Text-to-Prosody Control	Aug 13, 2024	Prosody PredictionSpeech Synthesis	CodeCode Available	1
Generative Expressive Conversational Speech Synthesis	Jul 31, 2024	Speech Synthesis	CodeCode Available	1
VoxSim: A perceptual voice similarity dataset	Jul 26, 2024	BenchmarkingSpeaker Recognition	CodeCode Available	1
dMel: Speech Tokenization made Simple	Jul 22, 2024	DecoderLanguage Modeling	CodeCode Available	1
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings	Jul 19, 2024	Expressive Speech SynthesisSpeech Synthesis	CodeCode Available	1
Fine-Grained and Interpretable Neural Speech Editing	Jul 7, 2024	Data AugmentationSpeech Synthesis	CodeCode Available	1
Articulatory Phonetics Informed Controllable Expressive Speech Synthesis	Jun 15, 2024	Expressive Speech SynthesisSpeech Synthesis	CodeCode Available	1
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts	Apr 29, 2024	Contrastive LearningSpeech Synthesis	CodeCode Available	1
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks	Apr 6, 2024	Domain AdaptationSpeech Synthesis	CodeCode Available	1
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis	Apr 1, 2024	Speech Synthesistext-to-speech	CodeCode Available	1

Show:10 25 50

← PrevPage 2 of 25Next →

All datasets LibriTTS North American English LJSpeech Mandarin Chinese Blizzard Challenge 2013

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	PeriodWave-Turbo-L	PESQ	4.45	—	Unverified
2	BigVGAN-v2	PESQ	4.36	—	Unverified
3	EVA-GAN-big	PESQ	4.35	—	Unverified
4	PeriodWave + FreeU	PESQ	4.25	—	Unverified
5	RFWave	PESQ	4.23	—	Unverified
6	BigVSAN (w/ snakebeta)	PESQ	4.12	—	Unverified
7	BigVSAN	PESQ	4.12	—	Unverified
8	EVA-GAN-base	PESQ	4.03	—	Unverified
9	BigVGAN	PESQ	4.03	—	Unverified
10	Vocos	PESQ	3.7	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Tacotron 2	Mean Opinion Score	4.53	—	Unverified
2	WaveNet (Linguistic)	Mean Opinion Score	4.34	—	Unverified
3	WaveNet (L+F)	Mean Opinion Score	4.21	—	Unverified
4	Tacotron	Mean Opinion Score	4	—	Unverified
5	HMM-driven concatenative	Mean Opinion Score	3.86	—	Unverified
6	LSTM-RNN parametric	Mean Opinion Score	3.67	—	Unverified
7	means	Mean Opinion Score	0	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	BDDM vocoder	Mean Opinion Score	4.48	—	Unverified
2	DiffWave LARGE	Mean Opinion Score	4.44	—	Unverified
3	Neural HMM	Mean Opinion Score	3.24	—	Unverified
4	Neural HMM Ablation with 1 state per phone	Mean Opinion Score	2.68	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	WaveNet (L+F)	Mean Opinion Score	4.08	—	Unverified
2	LSTM-RNN parametric	Mean Opinion Score	3.79	—	Unverified
3	HMM-driven concatenative	Mean Opinion Score	3.47	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SampleRNN (2-tier)	NLL	1.39	—	Unverified
2	SampleRNN (3-tier)	NLL	1.39	—	Unverified