SOTAVerified

Shortest-Path Graph Kernels for Document Similarity

2017-09-01EMNLP 2017Unverified0· sign in to hype

Giannis Nikolentzos, Polykarpos Meladianos, Fran{\c{c}}ois Rousseau, Yannis Stavrakas, Michalis Vazirgiannis

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

In this paper, we present a novel document similarity measure based on the definition of a graph kernel between pairs of documents. The proposed measure takes into account both the terms contained in the documents and the relationships between them. By representing each document as a graph-of-words, we are able to model these relationships and then determine how similar two documents are by using a modified shortest-path graph kernel. We evaluate our approach on two tasks and compare it against several baseline approaches using various performance metrics such as DET curves and macro-average F1-score. Experimental results on a range of datasets showed that our proposed approach outperforms traditional techniques and is capable of measuring more accurately the similarity between two documents.

Tasks

Reproductions