SOTAVerified

Measuring Topic Coherence through Optimal Word Buckets

2017-04-01EACL 2017Unverified0· sign in to hype

Nitin Ramrakhiyani, Sachin Pawar, Swapnil Hingmire, Girish Palshikar

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

Measuring topic quality is essential for scoring the learned topics and their subsequent use in Information Retrieval and Text classification. To measure quality of Latent Dirichlet Allocation (LDA) based topics learned from text, we propose a novel approach based on grouping of topic words into buckets (TBuckets). A single large bucket signifies a single coherent theme, in turn indicating high topic coherence. TBuckets uses word embeddings of topic words and employs singular value decomposition (SVD) and Integer Linear Programming based optimization to create coherent word buckets. TBuckets outperforms the state-of-the-art techniques when evaluated using 3 publicly available datasets and on another one proposed in this paper.

Tasks

Reproductions