Fine-tuning BERT to classify COVID19 tweets containing symptoms

2021-06-01NAACL (SMM4H) 2021Unverified0· sign in to hype

Rajarshi Roychoudhury, Sudip Naskar

Unverified — Be the first to reproduce this paper.

Abstract

Twitter is a valuable source of patient-generated data that has been used in various population health studies. The first step in many of these studies is to identify and capture Twitter messages (tweets) containing medication mentions. Identifying personal mentions of COVID19 symptoms requires distinguishing personal mentions from other mentions such as symptoms reported by others and references to news articles or other sources. In this article, we describe our submission to Task 6 of the Social Media Mining for Health Applications (SMM4H) Shared Task 2021. This task challenged participants to classify tweets where the target classes are:(1) self-reports,(2) non-personal reports, and (3) literature/news mentions. Our system used a handcrafted preprocessing and word embeddings from BERT encoder model. We achieved an F1 score of 93%

Tasks

Articles Word Embeddings

Fine-tuning BERT to classify COVID19 tweets containing symptoms

Abstract

Tasks

Reproductions