SOTAVerified

Building a Biomedical Full-Text Part-of-Speech Corpus Semi-Automatically

2022-06-01LREC (LAW) 2022Code Available0· sign in to hype

Nicholas Elder, Robert E. Mercer, Sudipta Singha Roy

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

This paper presents a method for semi-automatically building a corpus of full-text English-language biomedical articles annotated with part-of-speech tags. The outcomes are a semi-automatic procedure to create a large silver standard corpus of 5 million sentences drawn from a large corpus of full-text biomedical articles annotated for part-of-speech, and a robust, easy-to-use software tool that assists the investigation of differences in two tagged datasets. The method to build the corpus uses two part-of-speech taggers designed to tag biomedical abstracts followed by a human dispute settlement when the two taggers differ on the tagging of a token. The dispute resolution aspect is facilitated by the software tool which organizes and presents the disputed tags. The corpus and all of the software that has been implemented for this study are made publicly available.

Tasks

Reproductions