SOTAVerified

Inducing Bilingual Lexica From Non-Parallel Data With Earth Mover's Distance Regularization

2016-12-01COLING 2016Unverified0· sign in to hype

Meng Zhang, Yang Liu, Huanbo Luan, Yiqun Liu, Maosong Sun

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

Being able to induce word translations from non-parallel data is often a prerequisite for cross-lingual processing in resource-scarce languages and domains. Previous endeavors typically simplify this task by imposing the one-to-one translation assumption, which is too strong to hold for natural languages. We remove this constraint by introducing the Earth Mover's Distance into the training of bilingual word embeddings. In this way, we take advantage of its capability to handle multiple alternative word translations in a natural form of regularization. Our approach shows significant and consistent improvements across four language pairs. We also demonstrate that our approach is particularly preferable in resource-scarce settings as it only requires a minimal seed lexicon.

Tasks

Reproductions