SOTAVerified

Sample selection from a given dataset to validate machine learning models

2021-04-27Unverified0· sign in to hype

Bertrand Iooss

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

The selection of a validation basis from a full dataset is often required in industrial use of supervised machine learning algorithm. This validation basis will serve to realize an independent evaluation of the machine learning model. To select this basis, we propose to adopt a "design of experiments" point of view, by using statistical criteria. We show that the "support points" concept, based on Maximum Mean Discrepancy criteria, is particularly relevant. An industrial test case from the company EDF illustrates the practical interest of the methodology.

Tasks

Reproductions