ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech

2018-07-19ICLR 2019Code Available1· sign in to hype

Wei Ping, Kainan Peng, Jitong Chen

Code Available — Be the first to reproduce this paper.

Code

github.com/tiberiu44/TTS-Cube
pytorch★ 223
github.com/dhgrs/chainer-ClariNet
none★ 0
github.com/kensun0/Parallel-Wavenet
tf★ 0
github.com/ksw0306/ClariNet
pytorch★ 0
github.com/rickyHong/ClariNet-WaveNet-repl
pytorch★ 0

Abstract

In this work, we propose a new solution for parallel wave generation by WaveNet. In contrast to parallel WaveNet (van den Oord et al., 2018), we distill a Gaussian inverse autoregressive flow from the autoregressive WaveNet by minimizing a regularized KL divergence between their highly-peaked output distributions. Our method computes the KL divergence in closed-form, which simplifies the training algorithm and provides very efficient distillation. In addition, we introduce the first text-to-wave neural architecture for speech synthesis, which is fully convolutional and enables fast end-to-end training from scratch. It significantly outperforms the previous pipeline that connects a text-to-spectrogram model to a separately trained WaveNet (Ping et al., 2018). We also successfully distill a parallel waveform synthesizer conditioned on the hidden representation in this end-to-end model.

Tasks

Speech Synthesis text-to-speech Text to Speech

ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech

Code

Abstract

Tasks

Reproductions