SOTAVerified

An efficient language independent toolkit for complete morphological disambiguation

2014-05-01LREC 2014Unverified0· sign in to hype

L{\'a}szl{\'o} Laki, Gy{\"o}rgy Orosz

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

In this paper a Moses SMT toolkit-based language-independent complete morphological annotation tool is presented called HuLaPos2. Our system performs PoS tagging and lemmatization simultaneously. Amongst others, the algorithm used is able to handle phrases instead of unigrams, and can perform the tagging in a not strictly left-to-right order. With utilizing these gains, our system outperforms the HMM-based ones. In order to handle the unknown words, a suffix-tree based guesser was integrated into HuLaPos2. To demonstrate the performance of our system it was compared with several systems in different languages and PoS tag sets. In general, it can be concluded that the quality of HuLaPos2 is comparable with the state-of-the-art systems, and in the case of PoS tagging it outperformed many available systems.

Tasks

Reproductions