SOTAVerified

LeSpell - A Multi-Lingual Benchmark Corpus of Spelling Errors to Develop Spellchecking Methods for Learner Language

2022-06-01LREC 2022Code Available0· sign in to hype

Marie Bexte, Ronja Laarmann-Quante, Andrea Horbach, Torsten Zesch

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

Spellchecking text written by language learners is especially challenging because errors made by learners differ both quantitatively and qualitatively from errors made by already proficient learners. We introduce LeSpell, a multi-lingual (English, German, Italian, and Czech) evaluation data set of spelling mistakes in context that we compiled from seven underlying learner corpora. Our experiments show that existing spellcheckers do not work well with learner data. Thus, we introduce a highly customizable spellchecking component for the DKPro architecture, which improves performance in many settings.

Reproductions