PatternRank: Leveraging Pretrained Language Models and Part of Speech for Unsupervised Keyphrase Extraction

2022-10-11Code Available2· sign in to hype

Tim Schopf, Simon Klimek, Florian Matthes

Code Available — Be the first to reproduce this paper.

Code

github.com/timschopf/keyphrasevectorizers
OfficialIn papertf★ 267

Abstract

Keyphrase extraction is the process of automatically selecting a small set of most relevant phrases from a given text. Supervised keyphrase extraction approaches need large amounts of labeled training data and perform poorly outside the domain of the training data. In this paper, we present PatternRank, which leverages pretrained language models and part-of-speech for unsupervised keyphrase extraction from single documents. Our experiments show PatternRank achieves higher precision, recall and F1-scores than previous state-of-the-art approaches. In addition, we present the KeyphraseVectorizers package, which allows easy modification of part-of-speech patterns for candidate keyphrase selection, and hence adaptation of our approach to any domain.

Tasks

Keyphrase Extraction

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
Inspec	PatternRank	F1@10	30.99	—	Unverified

PatternRank: Leveraging Pretrained Language Models and Part of Speech for Unsupervised Keyphrase Extraction

Code

Abstract

Tasks

Benchmark Results

Reproductions