A Domain Knowledge Enhanced Pre-Trained Language Model for Vertical Search: Case Study on Medicinal Products

2022-10-01COLING 2022Code Available0· sign in to hype

Kesong Liu, Jianhui Jiang, Feifei Lyu

Code Available — Be the first to reproduce this paper.

Code

github.com/liuks/ep_plm
OfficialIn papertf★ 5

Abstract

We present a biomedical knowledge enhanced pre-trained language model for medicinal product vertical search. Following ELECTRA’s replaced token detection (RTD) pre-training, we leverage biomedical entity masking (EM) strategy to learn better contextual word representations. Furthermore, we propose a novel pre-training task, product attribute prediction (PAP), to inject product knowledge into the pre-trained language model efficiently by leveraging medicinal product databases directly. By sharing the parameters of PAP’s transformer encoder with that of RTD’s main transformer, these two pre-training tasks are jointly learned. Experiments demonstrate the effectiveness of PAP task for pre-trained language model on medicinal product vertical search scenario, which includes query-title relevance, query intent classification, and named entity recognition in query.

Tasks

Attribute intent-classification Intent Classification Language Modeling Language Modelling named-entity-recognition Named Entity Recognition Named Entity Recognition (NER)

A Domain Knowledge Enhanced Pre-Trained Language Model for Vertical Search: Case Study on Medicinal Products

Code

Abstract

Tasks

Reproductions