Speech Synthesis

Speech synthesis is the task of generating speech from some other modality like text, lip movements etc.

Please note that the leaderboards here are not really comparable between studies - as they use mean opinion score as a metric and collect different samples from Amazon Mechnical Turk.

( Image credit: WaveNet: A generative model for raw audio )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 1226–1249 of 1249 papers

Title	Date	Tasks	Status
ULex: new data models and a mobile environment for corpus enrichment.	May 1, 2012	Speech Synthesis	—Unverified
Building Synthetic Voices in the META-NET Framework	May 1, 2012	Speech SynthesisVoice Conversion	—Unverified
Evaluating expressive speech synthesis from audiobook corpora for conversational phrases	May 1, 2012	ClusteringExpressive Speech Synthesis	—Unverified
Designing French Tale Corpora for Entertaining Text To Speech Synthesis	May 1, 2012	SentenceSpeech Synthesis	—Unverified
Building Text-To-Speech Voices in the Cloud	May 1, 2012	Speech RecognitionSpeech Synthesis	—Unverified
BUCEADOR, a multi-language search engine for digital libraries	May 1, 2012	Automatic Speech RecognitionAutomatic Speech Recognition (ASR)	—Unverified
Learning Sentiment Lexicons in Spanish	May 1, 2012	Opinion MiningQuestion Answering	—Unverified
Method for Collection of Acted Speech Using Various Situation Scripts	May 1, 2012	Speech Synthesis	—Unverified
Building a synchronous corpus of acoustic and 3D facial marker data for adaptive audio-visual speech synthesis	May 1, 2012	Audio-Visual Speech RecognitionSpeech Recognition	—Unverified
The Herme Database of Spontaneous Multimodal Human-Robot Dialogues	May 1, 2012	Gesture RecognitionSpeech Recognition	—Unverified
Constructive Interaction for Talking about Interesting Topics	May 1, 2012	ManagementSpeech Recognition	—Unverified
Comparing performance of different set-covering strategies for linguistic content optimization in speech corpora	May 1, 2012	DescriptiveSpeech Recognition	—Unverified
Versatile Speech Databases for High Quality Synthesis for Basque	May 1, 2012	Emotional Speech SynthesisSpeech Synthesis	—Unverified
Fast Labeling and Transcription with the Speechalyzer Toolkit	May 1, 2012	Audio ClassificationBenchmarking	—Unverified
Towards Fully Automatic Annotation of Audio Books for TTS	May 1, 2012	Speech RecognitionSpeech Synthesis	—Unverified
A Multilingual Natural Stress Emotion Database	May 1, 2012	Emotion RecognitionSpeech Synthesis	—Unverified
Statistical Evaluation of Pronunciation Encoding	May 1, 2012	Speech RecognitionSpeech Synthesis	—Unverified
A Galician Syntactic Corpus with Application to Intonation Modeling	May 1, 2012	Speech Synthesis	—Unverified
Open-Source Boundary-Annotated Corpus for Arabic Speech and Language Processing	May 1, 2012	ChunkingDescriptive	—Unverified
Texto4Science: a Quebec French Database of Annotated Short Text Messages	May 1, 2012	Speech SynthesisText-To-Speech Synthesis	—Unverified
CoALT: A Software for Comparing Automatic Labelling Tools	May 1, 2012	Speech RecognitionSpeech Synthesis	—Unverified
Practical Evaluation of Human and Synthesized Speech for Virtual Human Dialogue Systems	May 1, 2012	Speech Synthesis	—Unverified
Predicting Phrase Breaks in Classical and Modern Standard Arabic Text	May 1, 2012	ChunkingHuman Parsing	—Unverified
BAD: An Assistant tool for making verses in Basque	Apr 1, 2012	Speech SynthesisText-To-Speech Synthesis	—Unverified

Show:10 25 50

← PrevPage 50 of 50Next →

All datasets LibriTTS North American English LJSpeech Mandarin Chinese Blizzard Challenge 2013

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	PeriodWave-Turbo-L	PESQ	4.45	—	Unverified
2	BigVGAN-v2	PESQ	4.36	—	Unverified
3	EVA-GAN-big	PESQ	4.35	—	Unverified
4	PeriodWave + FreeU	PESQ	4.25	—	Unverified
5	RFWave	PESQ	4.23	—	Unverified
6	BigVSAN (w/ snakebeta)	PESQ	4.12	—	Unverified
7	BigVSAN	PESQ	4.12	—	Unverified
8	EVA-GAN-base	PESQ	4.03	—	Unverified
9	BigVGAN	PESQ	4.03	—	Unverified
10	Vocos	PESQ	3.7	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Tacotron 2	Mean Opinion Score	4.53	—	Unverified
2	WaveNet (Linguistic)	Mean Opinion Score	4.34	—	Unverified
3	WaveNet (L+F)	Mean Opinion Score	4.21	—	Unverified
4	Tacotron	Mean Opinion Score	4	—	Unverified
5	HMM-driven concatenative	Mean Opinion Score	3.86	—	Unverified
6	LSTM-RNN parametric	Mean Opinion Score	3.67	—	Unverified
7	means	Mean Opinion Score	0	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	BDDM vocoder	Mean Opinion Score	4.48	—	Unverified
2	DiffWave LARGE	Mean Opinion Score	4.44	—	Unverified
3	Neural HMM	Mean Opinion Score	3.24	—	Unverified
4	Neural HMM Ablation with 1 state per phone	Mean Opinion Score	2.68	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	WaveNet (L+F)	Mean Opinion Score	4.08	—	Unverified
2	LSTM-RNN parametric	Mean Opinion Score	3.79	—	Unverified
3	HMM-driven concatenative	Mean Opinion Score	3.47	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SampleRNN (2-tier)	NLL	1.39	—	Unverified
2	SampleRNN (3-tier)	NLL	1.39	—	Unverified