Stronger Baselines for Grammatical Error Correction Using Pretrained Encoder-Decoder Model

2020-05-24Code Available1· sign in to hype

Satoru Katsumata, Mamoru Komachi

Code Available — Be the first to reproduce this paper.

Code

github.com/Katsumata420/generic-pretrained-GEC
OfficialIn paperpytorch★ 37
github.com/gotutiyan/GEC-BART
pytorch★ 4

Abstract

Studies on grammatical error correction (GEC) have reported the effectiveness of pretraining a Seq2Seq model with a large amount of pseudodata. However, this approach requires time-consuming pretraining for GEC because of the size of the pseudodata. In this study, we explore the utility of bidirectional and auto-regressive transformers (BART) as a generic pretrained encoder-decoder model for GEC. With the use of this generic pretrained model for GEC, the time-consuming pretraining can be eliminated. We find that monolingual and multilingual BART models achieve high performance in GEC, with one of the results being comparable to the current strong results in English GEC. Our implementations are publicly available at GitHub (https://github.com/Katsumata420/generic-pretrained-GEC).

Tasks

Decoder Grammatical Error Correction

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
CoNLL-2014 Shared Task	BART	F0.5	63	—	Unverified

Stronger Baselines for Grammatical Error Correction Using Pretrained Encoder-Decoder Model

Code

Abstract

Tasks

Benchmark Results

Reproductions