SOTAVerified

Don't Rule Out Simple Models Prematurely: A Large Scale Benchmark Comparing Linear and Non-linear Classifiers in OpenML

2018-10-05IDA 2018: Advances in Intelligent Data Analysis XVII 2018Code Available0· sign in to hype

Benjamin Strang, Peter van der Putten, Jan N. van Rijn, Frank Hutter

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

A basic step for each data-mining or machine learning task is to determine which model to choose based on the problem and the data at hand. In this paper we investigate when non-linear classifiers outperform linear classifiers by means of a large scale experiment. We benchmark linear and non-linear versions of three types of classifiers (support vector machines; neural networks; and decision trees), and analyze the results to determine on what type of datasets the non-linear version performs better. To the best of our knowledge, this work is the first principled and large scale attempt to support the common assumption that non-linear classifiers excel only when large amounts of data are available.

Reproductions