Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

2020-04-23ACL 2020Code Available1· sign in to hype

Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith

Code Available — Be the first to reproduce this paper.

Code

github.com/allenai/dont-stop-pretraining
OfficialIn paperpytorch★ 540
github.com/shizhediao/black-box-prompt-learning
pytorch★ 56
github.com/shizhediao/t-dna
pytorch★ 19
github.com/SUSTechBruce/G-MAP
pytorch★ 9
github.com/bionlu-coling2024/biomed-ner-intent_detection
pytorch★ 8
github.com/sagjounkani/Dont-Stop-Pretraining-Use-Adapters-Instead
pytorch★ 1

Abstract

Language models pretrained on text from a wide variety of sources form the foundation of today's NLP. In light of the success of these broad-coverage models, we investigate whether it is still helpful to tailor a pretrained model to the domain of a target task. We present a study across four domains (biomedical and computer science publications, news, and reviews) and eight classification tasks, showing that a second phase of pretraining in-domain (domain-adaptive pretraining) leads to performance gains, under both high- and low-resource settings. Moreover, adapting to the task's unlabeled data (task-adaptive pretraining) improves performance even after domain-adaptive pretraining. Finally, we show that adapting to a task corpus augmented using simple data selection strategies is an effective alternative, especially when resources for domain-adaptive pretraining might be unavailable. Overall, we consistently find that multi-phase adaptive pretraining offers large gains in task performance.

Tasks

Citation Intent Classification

Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

Code

Abstract

Tasks

Reproductions