SOTAVerified

Don't Just Scratch the Surface: Enhancing Word Representations for Korean with Hanja

2019-08-25IJCNLP 2019Code Available0· sign in to hype

Kang Min Yoo, Taeuk Kim, Sang-goo Lee

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

We propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.e. Hanja). We employ cross-lingual transfer learning in training word representations by leveraging the fact that Hanja is closely related to Chinese. We evaluate the intrinsic quality of representations learned through our approach using the word analogy and similarity tests. In addition, we demonstrate their effectiveness on several downstream tasks, including a novel Korean news headline generation task.

Tasks

Reproductions