BLEURT: Learning Robust Metrics for Text Generation

2020-04-09ACL 2020Code Available1· sign in to hype

Thibault Sellam, Dipanjan Das, Ankur P. Parikh

Code Available — Be the first to reproduce this paper.

Code

github.com/google-research/bleurt
OfficialIn papertf★ 789
github.com/thu-coai/OpenMEVA
tf★ 51
github.com/ShiYaya/Awesome_Evaluation_Metrics_for_Text_Generation
none★ 6
github.com/sharanya-dasgupta001/hallushift
pytorch★ 4

Abstract

Text generation has made significant advances in the last few years. Yet, evaluation metrics have lagged behind, as the most popular choices (e.g., BLEU and ROUGE) may correlate poorly with human judgments. We propose BLEURT, a learned evaluation metric based on BERT that can model human judgments with a few thousand possibly biased training examples. A key aspect of our approach is a novel pre-training scheme that uses millions of synthetic examples to help the model generalize. BLEURT provides state-of-the-art results on the last three years of the WMT Metrics shared task and the WebNLG Competition dataset. In contrast to a vanilla BERT-based approach, it yields superior results even when the training data is scarce and out-of-distribution.

Tasks

Text Generation

BLEURT: Learning Robust Metrics for Text Generation

Code

Abstract

Tasks

Reproductions