SOTAVerified

Speech Synthesis

Speech synthesis is the task of generating speech from some other modality like text, lip movements etc.

Please note that the leaderboards here are not really comparable between studies - as they use mean opinion score as a metric and collect different samples from Amazon Mechnical Turk.

( Image credit: WaveNet: A generative model for raw audio )

Papers

Showing 751–800 of 1249 papers

TitleStatusHype
Prosody-TTS: An end-to-end speech synthesis system with prosody control—0
GANtron: Emotional Speech Synthesis with Generative Adversarial Networks—0
On the Interplay Between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis—0
Neural Speech Synthesis in German—0
Incorporating speaker embedding and post-filter network for improving speaker similarity of personalized speech synthesis system—0
Speech-MLP: a simple MLP architecture for speech processing—0
Conditioning Sequence-to-sequence Networks with Learned Activations—0
SynCLR: A Synthesis Framework for Contrastive Learning of out-of-domain Speech Representations—0
Guided-TTS:Text-to-Speech with Untranscribed Speech—0
FlowVocoder: A small Footprint Neural Vocoder based Normalizing flow for Speech Synthesis—0
Low-Latency Incremental Text-to-Speech Synthesis with Distilled Context Prediction Network—0
"Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World—0
On-device neural speech synthesis—0
fairseq S^2: A Scalable and Integrable Speech Synthesis ToolkitCode0
Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis—0
Physiological-Physical Feature Fusion for Automatic Voice Spoofing Detection—0
Neural Sequence-to-Sequence Speech Synthesis Using a Hidden Semi-Markov Model Based Structured Attention Mechanism—0
Full Attention Bidirectional Deep Learning Structure for Single Channel Speech Enhancement—0
Integrated Speech and Gesture SynthesisCode0
A Unified Transformer-based Framework for Duplex Text Normalization—0
Enhancing audio quality for expressive Neural Text-to-Speech—0
A Streamwise GAN Vocoder for Wideband Speech Coding at Very Low Bit Rate—0
Improved pronunciation prediction accuracy using morphology—0
Cross-speaker Style Transfer with Prosody Bottleneck in Neural Speech Synthesis—0
Exploring the Potential of Lexical Paraphrases for Mitigating Noise-Induced Comprehension Errors—0
Extending Text-to-Speech Synthesis with Articulatory Movement Prediction using Ultrasound Tongue ImagingCode0
Location, Location: Enhancing the Evaluation of Text-to-Speech Synthesis Using the Rapid Prosody Transcription Paradigm—0
Speech Synthesis from Text and Ultrasound Tongue Image-based Articulatory InputCode0
An Objective Evaluation Framework for Pathological Speech Synthesis—0
GANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech Synthesis—0
Preliminary study on using vector quantization latent spaces for TTS/VC systems with consistent performance—0
Distilling the Knowledge from Conditional Normalizing FlowsCode0
UniTTS: Residual Learning of Unified Embedding Space for Speech Style Control—0
Controllable Context-aware Conversational Speech Synthesis—0
Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis—0
Non-native English lexicon creation for bilingual speech synthesis—0
Advances in Speech Vocoding for Text-to-Speech with Continuous Parameters—0
EMOVIE: A Mandarin Emotion Speech Dataset with a Simple Emotional Text-to-Speech Model—0
A Flow-Based Neural Network for Time Domain Speech Enhancement—0
Ctrl-P: Temporal Control of Prosodic Variation for Speech Synthesis—0
Pathological voice adaptation with autoencoder-based voice conversion—0
Continuous Wavelet Vocoder-based Decomposition of Parametric Speech Waveform Synthesis—0
Sprachsynthese -- State-of-the-Art in englischer und deutscher Sprache—0
PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior—0
Learning to Efficiently Sample from Diffusion Probabilistic Models—0
Mathematical Vocoder Algorithm : Modified Spectral Inversion for Efficient Neural Speech Synthesis—0
Speaker verification-derived loss and data augmentation for DNN-based multispeaker speech synthesis—0
An objective evaluation of the effects of recording conditions and speaker characteristics in multi-speaker deep neural speech synthesis—0
NVC-Net: End-to-End Adversarial Voice ConversionCode0
Dual Script E2E framework for Multilingual and Code-Switching ASR—0
Show:102550
← PrevPage 16 of 25Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1PeriodWave-Turbo-LPESQ4.45—Unverified
2BigVGAN-v2PESQ4.36—Unverified
3EVA-GAN-bigPESQ4.35—Unverified
4PeriodWave + FreeUPESQ4.25—Unverified
5RFWavePESQ4.23—Unverified
6BigVSAN (w/ snakebeta)PESQ4.12—Unverified
7BigVSANPESQ4.12—Unverified
8EVA-GAN-basePESQ4.03—Unverified
9BigVGANPESQ4.03—Unverified
10VocosPESQ3.7—Unverified
#ModelMetricClaimedVerifiedStatus
1Tacotron 2Mean Opinion Score4.53—Unverified
2WaveNet (Linguistic)Mean Opinion Score4.34—Unverified
3WaveNet (L+F)Mean Opinion Score4.21—Unverified
4TacotronMean Opinion Score4—Unverified
5HMM-driven concatenativeMean Opinion Score3.86—Unverified
6LSTM-RNN parametricMean Opinion Score3.67—Unverified
7meansMean Opinion Score0—Unverified
#ModelMetricClaimedVerifiedStatus
1BDDM vocoderMean Opinion Score4.48—Unverified
2DiffWave LARGEMean Opinion Score4.44—Unverified
3Neural HMMMean Opinion Score3.24—Unverified
4Neural HMM Ablation with 1 state per phoneMean Opinion Score2.68—Unverified
#ModelMetricClaimedVerifiedStatus
1WaveNet (L+F)Mean Opinion Score4.08—Unverified
2LSTM-RNN parametricMean Opinion Score3.79—Unverified
3HMM-driven concatenativeMean Opinion Score3.47—Unverified
#ModelMetricClaimedVerifiedStatus
1SampleRNN (2-tier)NLL1.39—Unverified
2SampleRNN (3-tier)NLL1.39—Unverified