SOTAVerified

Grammatical Error Correction

Grammatical Error Correction (GEC) is the task of correcting different kinds of errors in text such as spelling, punctuation, grammatical, and word choice errors.

GEC is typically formulated as a sentence correction task. A GEC system takes a potentially erroneous sentence as input and is expected to transform it to its corrected version. See the example given below:

| Input (Erroneous) | Output (Corrected) | | ------------------------- | ---------------------- | |She see Tom is catched by policeman in park at last night. | She saw Tom caught by a policeman in the park last night.|

Papers

Showing 1–50 of 415 papers

TitleStatusHype
End-to-End Spoken Grammatical Error Correction—0
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator—0
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction—0
gec-metrics: A Unified Library for Grammatical Error Correction EvaluationCode0
Exploring the Feasibility of Multilingual Grammatical Error Correction with a Single LLM up to 9B parameters: A Comparative Study of 17 ModelsCode0
Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring—0
Deep Learning Model Deployment in Multiple Cloud Providers: an Exploratory Study Using Low Computing Power Environments—0
Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case Study—0
Corrections Meet Explanations: A Unified Framework for Explainable Grammatical Error Correction—0
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human?Code1
Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error CorrectionCode0
Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing—0
Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction—0
DSGram: Dynamic Weighting Sub-Metrics for Grammatical Error Correction in the Era of Large Language ModelsCode0
Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error CorrectionCode0
LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning—0
Speak & Improve Challenge 2025: Tasks and Baseline Systems—0
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback—0
Enhancing Grammatical Error Detection using BERT with Cleaned Lang-8 DatasetCode0
Tibyan Corpus: Balanced and Comprehensive Error Coverage Corpus Using ChatGPT for Arabic Grammatical Error Correction—0
Efficient and Interpretable Grammatical Error Correction with Mixture of ExpertsCode0
A Simple Yet Effective Corpus Construction Framework for Indonesian Grammatical Error CorrectionCode0
The Write & Improve Corpus 2024: Error-annotated and CEFR-labelled essays by learners of English—0
Multi-head Sequence Tagging Model for Grammatical Error CorrectionCode0
Grammatical Error Correction for Low-Resource Languages: The Case of Zarma—0
LLM-based Code-Switched Text Generation for Grammatical Error Correction—0
Grammatical Error Feedback: An Implicit Evaluation Approach—0
CLEME2.0: Towards More Interpretable Evaluation by Disentangling Edits for Grammatical Error CorrectionCode1
EXCGEC: A Benchmark of Edit-wise Explainable Chinese Grammatical Error Correction—0
Improving Grammatical Error Correction via Contextual Data AugmentationCode0
ChatLang-8: An LLM-Based Synthetic Data Generation Framework for Grammatical Error Correction—0
Detection-Correction Structure via General Language Model for Grammatical Error CorrectionCode1
Organic Data-Driven Approach for Turkish Grammatical Error Correction and LLMsCode0
GPT-3.5 for Grammatical Error Correction—0
Comparative study of models trained on synthetic data for Ukrainian grammatical error correctionCode0
Spivavtor: An Instruction Tuned Ukrainian Text Editing Model—0
Pillars of Grammatical Error Correction: Comprehensive Inspection Of Contemporary Approaches In The Era of Large Language ModelsCode1
Grammatical Error Correction for Code-Switched Sentences by Learners of EnglishCode0
Ungrammatical-syntax-based In-context Example Selection for Grammatical Error Correction—0
Large Language Models Are State-of-the-Art Evaluator for Grammatical Error Correction—0
LM-Combiner: A Contextual Rewriting Model for Chinese Grammatical Error CorrectionCode0
To Err Is Human, but Llamas Can Learn It TooCode0
Revisiting Meta-evaluation for Grammatical Error CorrectionCode0
Neural Automated Writing Evaluation with Corrective Feedback—0
mEdIT: Multilingual Text Editing via Instruction TuningCode1
Likelihood-based Mitigation of Evaluation Bias in Large Language ModelsCode0
Evaluating Prompting Strategies for Grammatical Error Correction Based on Language Proficiency—0
Rethinking the Roles of Large Language Models in Chinese Grammatical Error Correction—0
Alirector: Alignment-Enhanced Chinese Grammatical Error CorrectorCode1
Prompting open-source and commercial language models for grammatical error correction of English learner text—0
Show:102550
← PrevPage 1 of 9Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Ensembles of best 7 models + GRECO + GTP-rerankF0.572.8—Unverified
2Majority-voting ensemble on best 7 modelsF0.571.8—Unverified
3GRECO (voting+ESC)F0.571.12—Unverified
4GEC-DI (LM+GED)F0.569.6—Unverified
5Unsupervised GEC + cLang8F0.569.6—Unverified
6ESCF0.569.51—Unverified
7T5F0.568.87—Unverified
8MoECEF0.567.79—Unverified
9SynGECF0.567.6—Unverified
10Sequence tagging + token-level transformations + two-stage fine-tuning (+BERT, RoBERTa, XLNet)F0.566.5—Unverified
#ModelMetricClaimedVerifiedStatus
1Majority-voting ensemble on best 7 modelsF0.581.4—Unverified
2GRECO (voting+ESC)F0.580.84—Unverified
3ESCF0.579.9—Unverified
4RedPenNetF0.577.6—Unverified
5clang_large_ft2-gectorF0.577.1—Unverified
6Unsupervised GEC + cLang8F0.576.5—Unverified
7DeBERTa + RoBERTa + XLNetF0.576.05—Unverified
8MoECEF0.574.07—Unverified
9Sequence tagging + token-level transformations + two-stage fine-tuning (+RoBERTa, XLNet)F0.573.7—Unverified
10BEA CombinationF0.573.2—Unverified
#ModelMetricClaimedVerifiedStatus
1Llama + 1M BT + goldF0.576.75—Unverified
2mT5-based multimodal MoEF0.576.3—Unverified
3gT5 xxlF0.575.96—Unverified
4TransformerF0.573.71—Unverified
5Transformer - synthetic pretrain onlyF0.551.41—Unverified
6Multilayer Convolutional Encoder-DecoderF0.543.35—Unverified
#ModelMetricClaimedVerifiedStatus
1VERNetGLEU62.1—Unverified
2Transformer + Pre-train with Pseudo Data + BERTGLEU62—Unverified
3SMT + BiGRUGLEU61.5—Unverified
4Copy-augmented Model (4 Ensemble +Denoising Autoencoder)GLEU61—Unverified
5TransformerGLEU59.9—Unverified
6CNN Seq2SeqGLEU57.47—Unverified
#ModelMetricClaimedVerifiedStatus
1Llama + 1M BT + goldF0.574.09—Unverified
2mBART-based model with synthetic dataF0.568.17—Unverified
3mT5 large + 10M synthF0.568.09—Unverified
4RedPenNetF0.567.71—Unverified
5ChatGPT (zero-shot)F0.527.4—Unverified
#ModelMetricClaimedVerifiedStatus
1GRECO (vote+ESC)F0.585.21—Unverified
2SMT + BiGRUF0.572.04—Unverified
3CNN Seq2SeqF0.570.14—Unverified
#ModelMetricClaimedVerifiedStatus
1CNN Seq2Seq + Quality EstimationF0.556.52—Unverified
2TransformerF0.555.8—Unverified
3+ BIFI with no criticF0.518.7—Unverified
#ModelMetricClaimedVerifiedStatus
1CNN Seq2Seq + Fluency Boost and inferenceGLEU62.37—Unverified
2CNN Seq2Seq + Fluency BoostF0.561.34—Unverified
3+ BIFI (ours)F0.542.4—Unverified
#ModelMetricClaimedVerifiedStatus
1TransformerGLEU59.9—Unverified
2CNN Seq2SeqGLEU57.47—Unverified
#ModelMetricClaimedVerifiedStatus
1Llama + 1M BT + goldF0.569.97—Unverified
#ModelMetricClaimedVerifiedStatus
1STG-Jointexact match34.1—Unverified
#ModelMetricClaimedVerifiedStatus
1GEC-DI (LM+GED)F0.548.61—Unverified
#ModelMetricClaimedVerifiedStatus
1RedPenNetF0.577.6—Unverified