Multi-blank Transducers for Speech Recognition

2022-11-04Code Available1· sign in to hype

Hainan Xu, Fei Jia, Somshubra Majumdar, Shinji Watanabe, Boris Ginsburg

Code Available — Be the first to reproduce this paper.

Code

github.com/NVIDIA/NeMo
OfficialIn paperpytorch★ 16,967
github.com/chimechallenge/C8DASR-Baseline-NeMo
pytorch★ 13
github.com/kehanlu/Nemo
pytorch★ 1
github.com/wd929/NeMo
pytorch★ 0

Abstract

This paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR). In standard RNN-T, the emission of a blank symbol consumes exactly one input frame; in our proposed method, we introduce additional blank symbols, which consume two or more input frames when emitted. We refer to the added symbols as big blanks, and the method multi-blank RNN-T. For training multi-blank RNN-Ts, we propose a novel logit under-normalization method in order to prioritize emissions of big blanks. With experiments on multiple languages and datasets, we show that multi-blank RNN-T methods could bring relative speedups of over +90%/+139% to model inference for English Librispeech and German Multilingual Librispeech datasets, respectively. The multi-blank RNN-T method also improves ASR accuracy consistently. We will release our implementation of the method in the NeMo (https://github.com/NVIDIA/NeMo) toolkit.

Tasks

Automatic Speech Recognition Automatic Speech Recognition (ASR)speech-recognition Speech Recognition

Multi-blank Transducers for Speech Recognition

Code

Abstract

Tasks

Reproductions