Utilizing constituent structure for compound analysis

2014-05-01LREC 2014Unverified0· sign in to hype

Krist{\'\i}n Bjarnad{\'o}ttir, J{\'o}n Da{\dh}ason

Unverified — Be the first to reproduce this paper.

Abstract

Compounding is extremely productive in Icelandic and multi-word compounds are common. The likelihood of finding previously unseen compounds in texts is thus very high, which makes out-of-vocabulary words a problem in the use of NLP tools. The tool de-scribed in this paper splits Icelandic compounds and shows their binary constituent structure. The probability of a constituent in an unknown (or unanalysed) compound forming a combined constituent with either of its neighbours is estimated, with the use of data on the constituent structure of over 240 thousand compounds from the Database of Modern Icelandic Inflection, and word frequencies from \'Islenskur or asj\'o ur, a corpus of approx. 550 million words. Thus, the structure of an unknown compound is derived by com-parison with compounds with partially the same constituents and similar structure in the training data. The granularity of the split re-turned by the decompounder is important in tasks such as semantic analysis or machine translation, where a flat (non-structured) se-quence of constituents is insufficient.

Tasks

Information Retrieval Machine Translation Part-Of-Speech Tagging Speech Recognition Translation

Utilizing constituent structure for compound analysis

Abstract

Tasks

Reproductions