Ensemble prosody prediction for expressive speech synthesis

2023-04-03Unverified0· sign in to hype

Tian Huey Teh, Vivian Hu, Devang S Ram Mohan, Zack Hodari, Christopher G. R. Wallis, Tomás Gomez Ibarrondo, Alexandra Torresquintero, James Leoni, Mark Gales, Simon King

arXiv PDF

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Yet, in evaluations it is generally found that no single model is preferred for all input texts. This suggests an approach that has rarely been used before for Text-to-Speech: an ensemble of models. We apply ensemble learning to prosody prediction. We construct simple ensembles of prosody predictors by varying either model architecture or model parameter values. To automatically select amongst the models in the ensemble when performing Text-to-Speech, we propose a novel, and computationally trivial, variance-based criterion. We demonstrate that even a small ensemble of prosody predictors yields useful diversity, which, combined with the proposed selection criterion, outperforms any individual model from the ensemble.

Tasks

Diversity Ensemble Learning Expressive Speech Synthesis Prediction Prosody Prediction Speech Synthesis text-to-speech Text to Speech

Ensemble prosody prediction for expressive speech synthesis

Abstract

Tasks

Reproductions