SOTAVerified

A Comparative Study of Text Preprocessing Approaches for Topic Detection of User Utterances

2016-05-01LREC 2016Unverified0· sign in to hype

Roman Sergienko, Muhammad Shan, Wolfgang Minker

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

The paper describes a comparative study of existing and novel text preprocessing and classification techniques for domain detection of user utterances. Two corpora are considered. The first one contains customer calls to a call centre for further call routing; the second one contains answers of call centre employees with different kinds of customer orientation behaviour. Seven different unsupervised and supervised term weighting methods were applied. The collective use of term weighting methods is proposed for classification effectiveness improvement. Four different dimensionality reduction methods were applied: stop-words filtering with stemming, feature selection based on term weights, feature transformation based on term clustering, and a novel feature transformation method based on terms belonging to classes. As classification algorithms we used k-NN and a SVM-based algorithm. The numerical experiments have shown that the simultaneous use of the novel proposed approaches (collectives of term weighting methods and the novel feature transformation method) allows reaching the high classification results with very small number of features.

Tasks

Reproductions