SOTAVerified

Towards Precise Lexicon Integration in Neural Machine Translation

2021-09-01RANLP 2021Code Available0· sign in to hype

Ogün Öz, Maria Sukhareva

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

Terminological consistency is an essential requirement for industrial translation. High-quality, hand-crafted terminologies contain entries in their nominal forms. Integrating such a terminology into machine translation is not a trivial task. The MT system must be able to disambiguate homographs on the source side and choose the correct wordform on the target side. In this work, we propose a simple but effective method for homograph disambiguation and a method of wordform selection by introducing multi-choice lexical constraints. We also propose a metric to measure the terminological consistency of the translation. Our results have a significant improvement over the current SOTA in terms of terminological consistency without any loss of the BLEU score. All the code used in this work will be published as open-source.

Tasks

Reproductions