SOTAVerified

Dynamic Classification in Web Archiving Collections

2020-05-01LREC 2020Unverified0· sign in to hype

Krutarth Patel, Cornelia Caragea, Mark Phillips

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

The Web archived data usually contains high-quality documents that are very useful for creating specialized collections of documents. To create such collections, there is a substantial need for automatic approaches that can distinguish the documents of interest for a collection out of the large collections (of millions in size) from Web Archiving institutions. However, the patterns of the documents of interest can differ substantially from one document to another, which makes the automatic classification task very challenging. In this paper, we explore dynamic fusion models to find, on the fly, the model or combination of models that performs best on a variety of document types. Our experimental results show that the approach that fuses different models outperforms individual models and other ensemble methods on three datasets.

Tasks

Reproductions