SOTAVerified

Speech Synthesis

Speech synthesis is the task of generating speech from some other modality like text, lip movements etc.

Please note that the leaderboards here are not really comparable between studies - as they use mean opinion score as a metric and collect different samples from Amazon Mechnical Turk.

( Image credit: WaveNet: A generative model for raw audio )

Papers

Showing 801–850 of 1249 papers

TitleStatusHype
Towards Natural Bilingual and Code-Switched Speech Synthesis Based on Mix of Monolingual Recordings and Cross-Lingual Voice Conversion—0
Towards Personalised Synthesised Voices for Individuals with Vocal Disabilities: Voice Banking and Reconstruction—0
Towards Robust Neural Vocoding for Speech Generation: A Survey—0
Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis—0
Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models—0
Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS—0
Towards Zero-Shot Text-To-Speech for Arabic Dialects—0
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator—0
Training Wake Word Detection with Synthesized Speech Data on Confusion Words—0
Trans-disciplinary spoken language processing studies for scientific understanding of second language learner's characteristics—0
Transferring neural speech waveform synthesizers to musical instrument sounds generation—0
Transformer-based Models of Text Normalization for Speech Applications—0
Transformer-Based Speech Synthesizer Attribution in an Open Set Scenario—0
Transformers in Speech Processing: A Survey—0
Transformer VQ-VAE for Unsupervised Unit Discovery and Speech Synthesis: ZeroSpeech 2020 Challenge—0
Transplantation of Conversational Speaking Style with Interjections in Sequence-to-Sequence Speech Synthesis—0
Triple M: A Practical Text-to-speech Synthesis System With Multi-guidance Attention And Multi-band Multi-time LPCNet—0
TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis using ranking support vector machine with variational autoencoder—0
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer—0
ULex: new data models and a mobile environment for corpus enrichment.—0
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis—0
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching—0
Uncovering Latent Style Factors for Expressive Speech Synthesis—0
UniFLG: Unified Facial Landmark Generator from Text or Speech—0
UniTTS: Residual Learning of Unified Embedding Space for Speech Style Control—0
Universal Neural Vocoding with Parallel WaveNet—0
Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks—0
Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis—0
Unsupervised Quantized Prosody Representation for Controllable Speech Synthesis—0
Unsupervised Style and Content Separation by Minimizing Mutual Information for Speech Synthesis—0
Unsupervised word-level prosody tagging for controllable speech synthesis—0
Using a machine learning model to assess the complexity of stress systems—0
Using a Pitch-Synchronous Residual Codebook for Hybrid HMM/Frame Selection Speech Synthesis—0
Using Audio Books for Training a Text-to-Speech System—0
Using GANs to Synthesise Minimum Training Data for Deepfake Generation—0
Using multiple reference audios and style embedding constraints for speech synthesis—0
Using NLG for speech synthesis of mathematical sentences—0
Using previous acoustic context to improve Text-to-Speech synthesis—0
Using VAEs and Normalizing Flows for One-shot Text-To-Speech Synthesis of Expressive Speech—0
Utterance-level Sequential Modeling For Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit—0
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE—0
UzbekTagger: The rule-based POS tagger for Uzbek language—0
VAKTA-SETU: A Speech-to-Speech Machine Translation Service in Select Indic Languages—0
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers—0
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment—0
VANI: Very-lightweight Accent-controllable TTS for Native and Non-native speakers with Identity Preservation—0
VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention—0
Variations prosodiques en synth\`ese par s\'election d'unit\'es: l'exemple des phrases interrogatives (Prosodic variations in unit-based speech synthesis: the example of interrogative sentences) [in French]—0
VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion—0
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders—0
Show:102550
← PrevPage 17 of 25Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1PeriodWave-Turbo-LPESQ4.45—Unverified
2BigVGAN-v2PESQ4.36—Unverified
3EVA-GAN-bigPESQ4.35—Unverified
4PeriodWave + FreeUPESQ4.25—Unverified
5RFWavePESQ4.23—Unverified
6BigVSAN (w/ snakebeta)PESQ4.12—Unverified
7BigVSANPESQ4.12—Unverified
8EVA-GAN-basePESQ4.03—Unverified
9BigVGANPESQ4.03—Unverified
10VocosPESQ3.7—Unverified
#ModelMetricClaimedVerifiedStatus
1Tacotron 2Mean Opinion Score4.53—Unverified
2WaveNet (Linguistic)Mean Opinion Score4.34—Unverified
3WaveNet (L+F)Mean Opinion Score4.21—Unverified
4TacotronMean Opinion Score4—Unverified
5HMM-driven concatenativeMean Opinion Score3.86—Unverified
6LSTM-RNN parametricMean Opinion Score3.67—Unverified
7meansMean Opinion Score0—Unverified
#ModelMetricClaimedVerifiedStatus
1BDDM vocoderMean Opinion Score4.48—Unverified
2DiffWave LARGEMean Opinion Score4.44—Unverified
3Neural HMMMean Opinion Score3.24—Unverified
4Neural HMM Ablation with 1 state per phoneMean Opinion Score2.68—Unverified
#ModelMetricClaimedVerifiedStatus
1WaveNet (L+F)Mean Opinion Score4.08—Unverified
2LSTM-RNN parametricMean Opinion Score3.79—Unverified
3HMM-driven concatenativeMean Opinion Score3.47—Unverified
#ModelMetricClaimedVerifiedStatus
1SampleRNN (2-tier)NLL1.39—Unverified
2SampleRNN (3-tier)NLL1.39—Unverified