SOTAVerified

PAC Mode Estimation using PPR Martingale Confidence Sequences

2021-09-10Code Available0· sign in to hype

Shubham Anand Jain, Rohan Shah, Sanit Gupta, Denil Mehta, Inderjeet Jayakumar Nair, Jian Vora, Sushil Khyalia, Sourav Das, Vinay J. Ribeiro, Shivaram Kalyanakrishnan

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

We consider the problem of correctly identifying the mode of a discrete distribution P with sufficiently high probability by observing a sequence of i.i.d. samples drawn from P. This problem reduces to the estimation of a single parameter when P has a support set of size K = 2. After noting that this special case is tackled very well by prior-posterior-ratio (PPR) martingale confidence sequences waudby-ramdas-ppr, we propose a generalisation to mode estimation, in which P may take K 2 values. To begin, we show that the "one-versus-one" principle to generalise from K = 2 to K 2 classes is more efficient than the "one-versus-rest" alternative. We then prove that our resulting stopping rule, denoted PPR-1v1, is asymptotically optimal (as the mistake probability is taken to 0). PPR-1v1 is parameter-free and computationally light, and incurs significantly fewer samples than competitors even in the non-asymptotic regime. We demonstrate its gains in two practical applications of sampling: election forecasting and verification of smart contracts in blockchains.

Reproductions