SOTAVerified

Convergent Reinforcement Learning with Function Approximation: A Bilevel Optimization Perspective

2018-09-27Unverified0· sign in to hype

Zhuoran Yang, Zuyue Fu, Kaiqing Zhang, Zhaoran Wang

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

We study reinforcement learning algorithms with nonlinear function approximation in the online setting. By formulating both the problems of value function estimation and policy learning as bilevel optimization problems, we propose online Q-learning and actor-critic algorithms for these two problems respectively. Our algorithms are gradient-based methods and thus are computationally efficient. Moreover, by approximating the iterates using differential equations, we establish convergence guarantees for the proposed algorithms. Thorough numerical experiments are conducted to back up our theory.

Tasks

Reproductions