SOTAVerified

Quality Focused Approach to a Learner Corpus Development

2020-05-01LREC 2020Unverified0· sign in to hype

Roberts Dar{\c{g}}is, Ilze Auzi{\c{n}}a, Krist{\=\i}ne Lev{\=a}ne-Petrova, Inga Kaija

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

The paper presents quality focused approach to a learner corpus development. The methodology was developed with multiple design considerations put in place to make the annotation process easier and at the same time reduce the amount of mistakes that could be introduced due to inconsistent text correction or carelessness. The approach suggested in this paper consists of multiple parts: comparison of digitized texts by several annotators, text correction, automated morphological analysis, and manual review of annotations. The described approach is used to create Latvian Language Learner corpus (LaVA) which is part of a currently ongoing project Development of Learner corpus of Latvian: methods, tools and applications.

Tasks

Reproductions