SOTAVerified

Distributional Reinforcement Learning

Value distribution is the distribution of the random return received by a reinforcement learning agent. it been used for a specific purpose such as implementing risk-aware behaviour.

We have random return Z whose expectation is the value Q. This random return is also described by a recursive equation, but one of a distributional nature

Papers

Showing 81–90 of 137 papers

TitleStatusHype
Safe Distributional Reinforcement Learning—0
Sample-based Distributional Policy Gradient—0
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss—0
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation—0
Statistical Efficiency of Distributional Temporal Difference Learning and Freedman's Inequality in Hilbert Spaces—0
Statistics and Samples in Distributional Reinforcement Learning—0
Stochastically Dominant Distributional Reinforcement Learning—0
The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement Learning—0
The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning—0
The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation—0
Show:102550
← PrevPage 9 of 14Next →

No leaderboard results yet.