SOTAVerified

NTT's Machine Translation Systems for WMT19 Robustness Task

2019-07-09WS 2019Unverified0· sign in to hype

Soichiro Murakami, Makoto Morishita, Tsutomu Hirao, Masaaki Nagata

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

This paper describes NTT's submission to the WMT19 robustness task. This task mainly focuses on translating noisy text (e.g., posts on Twitter), which presents different difficulties from typical translation tasks such as news. Our submission combined techniques including utilization of a synthetic corpus, domain adaptation, and a placeholder mechanism, which significantly improved over the previous baseline. Experimental results revealed the placeholder mechanism, which temporarily replaces the non-standard tokens including emojis and emoticons with special placeholder tokens during translation, improves translation accuracy even with noisy texts.

Tasks

Reproductions