SOTAVerified

Lemmatization

Lemmatization is a process of determining a base or dictionary form (lemma) for a given surface form. Especially for languages with rich morphology it is important to be able to normalize words into their base forms to better support for example search engines and linguistic studies. Main difficulties in Lemmatization arise from encountering previously unseen words during inference time as well as disambiguating ambiguous surface forms which can be inflected variants of several different base forms depending on the context.

Source: Universal Lemmatizer: A Sequence to Sequence Model for Lemmatizing Universal Dependencies Treebanks

Papers

Showing 301–350 of 351 papers

TitleStatusHype
Analysing cross-lingual transfer in lemmatisation for Indian languages—0
An Analysis of Lemmatization on Topic Models of Morphologically Rich Language—0
Analyzing and Aligning German compound nouns—0
An annotated English child language database—0
An efficient language independent toolkit for complete morphological disambiguation—0
An ELECTRA Model for Latin Token Tagging Tasks—0
A Neural Lemmatizer for Bengali—0
An Evaluation of Lexicon-based Sentiment Analysis Techniques for the Plays of Gotthold Ephraim Lessing—0
An extended morphological analyzer of German handling verbal forms with separated separable particles (Un analyseur morphologique \'etendu de l'allemand traitant les formes verbales \`a particule s\'epar\'ee) [in French]—0
An Extensible Multilingual Open Source Lemmatizer—0
ANNLOR: A Na\" Notation-system for Lexical Outputs Ranking—0
An NLP Pipeline for Coptic—0
A Publicly Available Cross-Platform Lemmatizer for Bulgarian—0
Arabic Word-level Readability Visualization for Assisted Text Simplification—0
A Resource for Studying Chatino Verbal Morphology—0
A set of open source tools for Turkish natural language processing—0
A Simple Joint Model for Improved Contextual Neural Lemmatization—0
ASOBEK at SemEval-2016 Task 1: Sentence Representation with Character N-gram Embeddings for Semantic Textual Similarity—0
Attention-free encoder decoder for morphological processing—0
A unified lexical processing framework based on the Margin Infused Relaxed Algorithm. A case study on the Romanian Language—0
Authorship Attribution Based on Life-Like Network Automata—0
Automated Identification of Disaster News For Crisis Management Using Machine Learning—0
Automatically Acquired Lexical Knowledge Improves Japanese Joint Morphological and Dependency Analysis—0
Automatic Categorization of Tagalog Documents Using Support Vector Machines—0
Automatic Extraction of Synonyms for German Particle Verbs from Parallel Data with Distributional Similarity as a Re-Ranking Feature—0
Automatic Translation of English Text to Indian Sign Language Synthetic Animations—0
BabyFST - Towards a Finite-State Based Computational Model of Ancient Babylonian—0
Better Together: Modern Methods Plus Traditional Thinking in NP Alignment—0
Biaffine Dependency and Semantic Graph Parsing for EnhancedUniversal Dependencies—0
BioRo: The Biomedical Corpus for the Romanian Language—0
bleu2vec: the Painfully Familiar Metric on Continuous Vector Space Steroids—0
Breaking the Fake News Barrier: Deep Learning Approaches in Bangla Language—0
Build Fast and Accurate Lemmatization for Arabic—0
Building a Lemmatizer and a Spell-checker for Sorani Kurdish—0
Building a multilingual parallel corpus for human users—0
Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages—0
CBNU System for SIGMORPHON 2019 Shared Task 2: a Pipeline Model—0
CELI: An Experiment with Cross Language Textual Entailment—0
CEPLEXicon ― A Lexicon of Child European Portuguese—0
Character-level Supervision for Low-resource POS Tagging—0
Chimera -- Three Heads for English-to-Czech Translation—0
CNGL-CORE: Referential Translation Machines for Measuring Semantic Similarity—0
Comparison of Current Approaches to Lemmatization: A Case Study in Estonian—0
Compounds and distributional thesauri—0
Constraint 2021: Machine Learning Models for COVID-19 Fake News Detection Shared Task—0
Context Aware Lemmatization and Morphological Tagging Method in Turkish—0
Context based lemmatizer for Polish language—0
Context Sensitive Lemmatization Using Two Successive Bidirectional Gated Recurrent Networks—0
Context Sensitive Neural Lemmatization with Lematus—0
Coreference Resolution in FreeLing 4.0—0
Show:102550
← PrevPage 7 of 8Next →

No leaderboard results yet.