SOTAVerified

Lexical Normalization

Lexical normalization is the task of translating/transforming a non standard text to a standard register.

Example:

new pix comming tomoroe
new pictures coming tomorrow

Datasets usually consists of tweets, since these naturally contain a fair amount of these phenomena.

For lexical normalization, only replacements on the word-level are annotated. Some corpora include annotation for 1-N and N-1 replacements. However, word insertion/deletion and reordering is not part of the task.

Papers

Showing 2647 of 47 papers

TitleStatusHype
An In-depth Analysis of the Effect of Lexical Normalization on the Dependency Parsing of Social Media0
Normalization of Indonesian-English Code-Mixed Twitter Data0
Lexical Normalization of User-Generated Medical Text0
MoNoise: A Multi-lingual and Easy-to-use Lexical Normalization ToolCode0
Adapting Sequence to Sequence models for Text Normalization in Social MediaCode0
Modeling Input Uncertainty in Neural Network Dependency ParsingCode0
Noise-Robust Morphological Disambiguation for Dialectal Arabic0
A Taxonomy for In-depth Evaluation of Normalization for User Generated Content0
Handling Normalization Issues for Part-of-Speech Tagging of Online Conversational Text0
MoNoise: Modeling Noise Using a Modular Normalization SystemCode0
The Denoised Web Treebank: Evaluating Dependency Parsing under Noisy Input Conditions0
NCSU-SAS-Ning: Candidate Generation and Feature Engineering for Supervised Lexical Normalization0
IHS\_RD: Lexical Normalization for English Tweets0
NCSU\_SAS\_SAM: Deep Encoding and Reconstruction for Normalization of Noisy Text0
Shared Tasks of the 2015 Workshop on Noisy User-generated Text: Twitter Lexical Normalization and Named Entity Recognition0
USZEGED: Correction Type-sensitive Normalization of English Tweets Using Efficiently Indexed n-gram Statistics0
Tweet Normalization with Syllables0
Accurate Word Segmentation and POS Tagging for Japanese Microblogs: Corpus Annotation and Joint Modeling with Lexical Normalization0
TweetNorm\_es: an annotated corpus for Spanish microtext normalization0
Towards Shared Datasets for Normalization Research0
A Large Corpus of Product Reviews in Portuguese: Tackling Out-Of-Vocabulary Words0
A Log-Linear Model for Unsupervised Text Normalization0
Show:102550
← PrevPage 2 of 2Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MoNoiseAccuracy87.63Unverified
2Syllable basedAccuracy86.08Unverified
3TextNormAccuracy83.94Unverified
4unLOLAccuracy82.06Unverified