Planning in entropy-regularized Markov decision processes and games
2019-12-01NeurIPS 2019Code Available0· sign in to hype
Jean-bastien Grill, Omar Darwiche Domingues, Pierre Menard, Remi Munos, Michal Valko
Code Available — Be the first to reproduce this paper.
ReproduceCode
- github.com/omardrwch/smoothcruiser-checkOfficialnone★ 0
Abstract
We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model of the SmoothCruiser. SmoothCruiser makes use of the smoothness of the Bellman operator promoted by the regularization to achieve problem-independent sample complexity of order O(1/^4) for a desired accuracy , whereas for non-regularized settings there are no known algorithms with guaranteed polynomial sample complexity in the worst case.