SOTAVerified

Lemmatization

Lemmatization is a process of determining a base or dictionary form (lemma) for a given surface form. Especially for languages with rich morphology it is important to be able to normalize words into their base forms to better support for example search engines and linguistic studies. Main difficulties in Lemmatization arise from encountering previously unseen words during inference time as well as disambiguating ambiguous surface forms which can be inflected variants of several different base forms depending on the context.

Source: Universal Lemmatizer: A Sequence to Sequence Model for Lemmatizing Universal Dependencies Treebanks

Papers

Showing 176–200 of 351 papers

TitleStatusHype
NeoN: A Tool for Automated Detection, Linguistic and LLM-Driven Analysis of Neologisms in Polish—0
NeoTag: a POS Tagger for Grammatical Neologism Detection—0
Neural Lemmatization of Multiword Expressions—0
Neural Polysynthetic Language Modelling—0
N-Gramas de Caractere como T\'ecnica de Normaliza \~ao Morfol\'ogica para L\' Portuguesa: Um Estudo em Categoriza \~ao de Textos (Character N-grams as a Morphological Normalization Technique for Portuguese Language: A Study in Text Categorization)—0
Non-Standard Words as Features for Text Categorization—0
NRC: A Machine Translation Approach to Cross-Lingual Word Sense Disambiguation (SemEval-2013 Task 10)—0
NRC Russian-English Machine Translation System for WMT 2016—0
On the Effectiveness of Dataset Embeddings in Mono-lingual,Multi-lingual and Zero-shot Conditions—0
On the Role of Morphological Information for Contextual Lemmatization—0
Open-Source Tools for Morphology, Lemmatization, POS Tagging and Named Entity Recognition—0
Optimizing a Distributional Semantic Model for the Prediction of German Particle Verb Compositionality—0
Overview of the EvaLatin 2020 Evaluation Campaign—0
Overview of the EvaLatin 2022 Evaluation Campaign—0
Oxford at SemEval-2017 Task 9: Neural AMR Parsing with Pointer-Augmented Attention—0
Parser combinators for Tigrinya and Oromo morphology—0
Persian Sentiment Analyzer: A Framework based on a Novel Feature Selection Method—0
Polish Coreference Corpus in Numbers—0
POS tagging, lemmatization and dependency parsing of West Frisian—0
Predicting the Compositionality of Nominal Compounds: Giving Word Embeddings a Hard Time—0
Probabilistic Lexical Generalization for French Dependency Parsing—0
Processing and Normalizing Hashtags—0
Producing Corpora of Medieval and Premodern Occitan—0
PurePos 2.0: a hybrid tool for morphological disambiguation—0
QLUT at SemEval-2017 Task 1: Semantic Textual Similarity Based on Word Embeddings—0
Show:102550
← PrevPage 8 of 15Next →

No leaderboard results yet.