Distributionally Robust Losses for Latent Covariate Mixtures
John Duchi, Tatsunori Hashimoto, Hongseok Namkoong
Code Available — Be the first to reproduce this paper.
ReproduceCode
- github.com/hsnamkoong/marginal-droOfficialIn paperpytorch★ 5
Abstract
While modern large-scale datasets often consist of heterogeneous subpopulations -- for example, multiple demographic groups or multiple text corpora -- the standard practice of minimizing average loss fails to guarantee uniformly low losses across all subpopulations. We propose a convex procedure that controls the worst-case performance over all subpopulations of a given size. Our procedure comes with finite-sample (nonparametric) convergence guarantees on the worst-off subpopulation. Empirically, we observe on lexical similarity, wine quality, and recidivism prediction tasks that our worst-case procedure learns models that do well against unseen subpopulations.