Speech Synthesis

Speech synthesis is the task of generating speech from some other modality like text, lip movements etc.

Please note that the leaderboards here are not really comparable between studies - as they use mean opinion score as a metric and collect different samples from Amazon Mechnical Turk.

( Image credit: WaveNet: A generative model for raw audio )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–100 of 1249 papers

Title	Date	Tasks	Status	Hype
A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models	Apr 22, 2025	cross-modal alignmentScript Generation	—Unverified	0
SOLIDO: A Robust Watermarking Method for Speech Synthesis via Low-Rank Adaptation	Apr 21, 2025	parameter-efficient fine-tuningSpeech Synthesis	—Unverified	0
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue	Apr 20, 2025	DiversitySpeech Synthesis	CodeCode Available	0
Collective Learning Mechanism based Optimal Transport Generative Adversarial Network for Non-parallel Voice Conversion	Apr 18, 2025	Generative Adversarial NetworkImage Generation	—Unverified	0
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis	Apr 14, 2025	Language ModelingLanguage Modelling	—Unverified	0
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis	Apr 14, 2025	RAGRetrieval-augmented Generation	—Unverified	0
SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis	Apr 14, 2025	Face SwappingSpeech Synthesis	CodeCode Available	1
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis	Apr 12, 2025	Speech Synthesis	—Unverified	0
Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis	Apr 10, 2025	Speech Synthesistext-to-speech	—Unverified	0
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow	Apr 10, 2025	Speech Synthesistext-to-speech	—Unverified	0
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models	Apr 3, 2025	Speech Synthesis	—Unverified	0
SpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM Development	Mar 31, 2025	Speech SynthesisVoice Cloning	CodeCode Available	0
SupertonicTTS: Towards Highly Scalable and Efficient Text-to-Speech System	Mar 29, 2025	Speech Synthesistext-to-speech	—Unverified	0
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech	Mar 21, 2025	Speech Synthesis	—Unverified	0
WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching	Mar 20, 2025	Speech Synthesis	CodeCode Available	2
MoonCast: High-Quality Zero-Shot Podcast Generation	Mar 18, 2025	Speech Synthesistext-to-speech	CodeCode Available	3
DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility	Mar 7, 2025	Speech Synthesis	CodeCode Available	0
Good practices for evaluation of synthesized speech	Mar 5, 2025	Speech Synthesis	—Unverified	0
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology	Mar 3, 2025	Speech SynthesisVoice Cloning	—Unverified	0
PodAgent: A Comprehensive Framework for Podcast Generation	Mar 1, 2025	Audio GenerationSpeech Synthesis	CodeCode Available	2
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models	Feb 27, 2025	DiversityLanguage Modeling	—Unverified	0
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis	Feb 26, 2025	Speech Synthesistext-to-speech	—Unverified	0
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM	Feb 24, 2025	Automatic Speech RecognitionLanguage Modeling	—Unverified	0
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions	Feb 18, 2025	Speech Synthesis	—Unverified	0
High-Fidelity Music Vocoder using Neural Audio Codecs	Feb 18, 2025	DecoderSpeech Synthesis	—Unverified	0
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing	Feb 17, 2025	Lip to Speech Synthesisspeech-recognition	—Unverified	0
A Survey on Bridging EEG Signals and Generative AI: From Image and Text to Beyond	Feb 17, 2025	Contrastive LearningEEG	—Unverified	0
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching	Feb 16, 2025	Language ModelingLanguage Modelling	—Unverified	0
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech	Feb 13, 2025	Adversarial AttackAdversarial Attack Detection	—Unverified	0
LoRP-TTS: Low-Rank Personalized Text-To-Speech	Feb 11, 2025	Speech Synthesistext-to-speech	—Unverified	0
Non-invasive electromyographic speech neuroprosthesis: a geometric perspective	Feb 9, 2025	Speech Synthesis	—Unverified	0
Gender Bias in Instruction-Guided Speech Synthesis Models	Feb 8, 2025	Expressive Speech SynthesisSpeech Synthesis	—Unverified	0
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis	Feb 6, 2025	Speech Synthesis	CodeCode Available	4
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet	Feb 4, 2025	Speech Synthesistext-to-speech	CodeCode Available	1
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis	Feb 3, 2025	QuantizationSpeech Synthesis	—Unverified	0
Compact Neural TTS Voices for Accessibility	Jan 28, 2025	Speech Synthesistext-to-speech	—Unverified	0
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation	Jan 24, 2025	Audio Deepfake DetectionDeepFake Detection	—Unverified	0
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement	Jan 23, 2025	Data AugmentationSpeech Enhancement	—Unverified	0
A Non-autoregressive Model for Joint STT and TTS	Jan 15, 2025	Automatic Speech Recognitionspeech-recognition	—Unverified	0
Speech Synthesis along Perceptual Voice Quality Dimensions	Jan 15, 2025	Expressive Speech SynthesisSpeech Synthesis	—Unverified	0
Exploring the encoding of linguistic representations in the Fully-Connected Layer of generative CNNs for Speech	Jan 13, 2025	Speech Synthesis	—Unverified	0
Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis	Jan 11, 2025	AttributeBenchmarking	CodeCode Available	1
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control	Jan 10, 2025	Speech Synthesistext-to-speech	—Unverified	0
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron	Jan 10, 2025	Speech Synthesistext-to-speech	—Unverified	0
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer	Jan 10, 2025	speech-recognitionSpeech Recognition	—Unverified	0
AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder	Jan 9, 2025	Pitch ClassificationPitch control	CodeCode Available	1
Probing Speaker-specific Features in Speaker Representations	Jan 9, 2025	Self-Supervised LearningSpeaker Verification	—Unverified	0
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis	Jan 9, 2025	Emotion RecognitionLanguage Modeling	—Unverified	0
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts	Jan 8, 2025	Speech Synthesis	—Unverified	0
OpenOmni: Large Language Models Pivot Zero-shot Omnimodal Alignment across Language with Real-time Self-Aware Emotional Speech Synthesis	Jan 8, 2025	DecoderEmotional Speech Synthesis	CodeCode Available	2

Show:10 25 50

← PrevPage 2 of 25Next →

All datasets LibriTTS North American English LJSpeech Mandarin Chinese Blizzard Challenge 2013

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	PeriodWave-Turbo-L	PESQ	4.45	—	Unverified
2	BigVGAN-v2	PESQ	4.36	—	Unverified
3	EVA-GAN-big	PESQ	4.35	—	Unverified
4	PeriodWave + FreeU	PESQ	4.25	—	Unverified
5	RFWave	PESQ	4.23	—	Unverified
6	BigVSAN (w/ snakebeta)	PESQ	4.12	—	Unverified
7	BigVSAN	PESQ	4.12	—	Unverified
8	EVA-GAN-base	PESQ	4.03	—	Unverified
9	BigVGAN	PESQ	4.03	—	Unverified
10	Vocos	PESQ	3.7	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Tacotron 2	Mean Opinion Score	4.53	—	Unverified
2	WaveNet (Linguistic)	Mean Opinion Score	4.34	—	Unverified
3	WaveNet (L+F)	Mean Opinion Score	4.21	—	Unverified
4	Tacotron	Mean Opinion Score	4	—	Unverified
5	HMM-driven concatenative	Mean Opinion Score	3.86	—	Unverified
6	LSTM-RNN parametric	Mean Opinion Score	3.67	—	Unverified
7	means	Mean Opinion Score	0	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	BDDM vocoder	Mean Opinion Score	4.48	—	Unverified
2	DiffWave LARGE	Mean Opinion Score	4.44	—	Unverified
3	Neural HMM	Mean Opinion Score	3.24	—	Unverified
4	Neural HMM Ablation with 1 state per phone	Mean Opinion Score	2.68	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	WaveNet (L+F)	Mean Opinion Score	4.08	—	Unverified
2	LSTM-RNN parametric	Mean Opinion Score	3.79	—	Unverified
3	HMM-driven concatenative	Mean Opinion Score	3.47	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SampleRNN (2-tier)	NLL	1.39	—	Unverified
2	SampleRNN (3-tier)	NLL	1.39	—	Unverified