A Twitter Corpus for Named Entity Recognition in Turkish

2022-06-01LREC 2022Code Available1· sign in to hype

Buse Çarık, Reyyan Yeniterzi

Code Available — Be the first to reproduce this paper.

Code

github.com/su-nlp/sunlp-twitter-ner-dataset
OfficialIn papernone★ 10

Abstract

This paper introduces a new Turkish Twitter Named Entity Recognition dataset. The dataset, which consists of 5000 tweets from a year-long period, was labeled by multiple annotators with a high agreement score. The dataset is also diverse in terms of the named entity types as it contains not only person, organization, and location but also time, money, product, and tv-show categories. Our initial experiments with pretrained language models (like BertTurk) over this dataset returned F1 scores of around 80%. We share this dataset publicly.

Tasks

named-entity-recognition Named Entity Recognition Named Entity Recognition (NER)

A Twitter Corpus for Named Entity Recognition in Turkish

Code

Abstract

Tasks

Reproductions