Text Document Clustering: Wordnet vs. TF-IDF vs. Word Embeddings

2021-01-01EACL (GWC) 2021Unverified0· sign in to hype

Michał Marcińczuk, Mateusz Gniewkowski, Tomasz Walkowiak, Marcin Będkowski

Unverified — Be the first to reproduce this paper.

Abstract

In the paper, we deal with the problem of unsupervised text document clustering for the Polish language. Our goal is to compare the modern approaches based on language modeling (doc2vec and BERT) with the classical ones, i.e., TF-IDF and wordnet-based. The experiments are conducted on three datasets containing qualification descriptions. The experiments’ results showed that wordnet-based similarity measures could compete and even outperform modern embedding-based approaches.

Tasks

Clustering Language Modeling Language Modelling Word Embeddings

Text Document Clustering: Wordnet vs. TF-IDF vs. Word Embeddings

Abstract

Tasks

Reproductions