Corpus Generation for Voice Command in Smart Home and the Effect of Speech Synthesis on End-to-End SLU

2020-05-01LREC 2020Unverified0· sign in to hype

Thierry Desot, Fran{\c{c}}ois Portet, Michel Vacher

Unverified — Be the first to reproduce this paper.

Abstract

Massive amounts of annotated data greatly contributed to the advance of the machine learning field. However such large data sets are often unavailable for novel tasks performed in realistic environments such as smart homes. In this domain, semantically annotated large voice command corpora for Spoken Language Understanding (SLU) are scarce, especially for non-English languages. We present the automatic generation process of a synthetic semantically-annotated corpus of French commands for smart-home to train pipeline and End-to-End (E2E) SLU models. SLU is typically performed through Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU) in a pipeline. Since errors at the ASR stage reduce the NLU performance, an alternative approach is End-to-End (E2E) SLU to jointly perform ASR and NLU. To that end, the artificial corpus was fed to a text-to-speech (TTS) system to generate synthetic speech data. All models were evaluated on voice commands acquired in a real smart home. We show that artificial data can be combined with real data within the same training set or used as a stand-alone training corpus. The synthetic speech quality was assessedby comparing it to real data using dynamic time warping (DTW).

Tasks

Automatic Speech Recognition Automatic Speech Recognition (ASR)Dynamic Time Warping Natural Language Understanding speech-recognition Speech Recognition Speech Synthesis Spoken Language Understanding text-to-speech Text to Speech

Corpus Generation for Voice Command in Smart Home and the Effect of Speech Synthesis on End-to-End SLU

Abstract

Tasks

Reproductions