SOTAVerified

Topic Modeling for Maternal Health Using Reddit

2021-04-01EACL (Louhi) 2021Unverified0· sign in to hype

Shuang Gao, Shivani Pandya, Smisha Agarwal, João Sedoc

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

This paper applies topic modeling to understand maternal health topics, concerns, and questions expressed in online communities on social networking sites. We examine Latent Dirichlet Analysis (LDA) and two state-of-the-art methods: neural topic model with knowledge distillation (KD) and Embedded Topic Model (ETM) on maternal health texts collected from Reddit. The models are evaluated on topic quality and topic inference, using both auto-evaluation metrics and human assessment. We analyze a disconnect between automatic metrics and human evaluations. While LDA performs the best overall with the auto-evaluation metrics NPMI and Coherence, Neural Topic Model with Knowledge Distillation is favorable by expert evaluation. We also create a new partially expert annotated gold-standard maternal health topic

Tasks

Reproductions