SOTAVerified

Atari Games

The Atari 2600 Games task (and dataset) involves training an agent to achieve high game scores.

( Image credit: Playing Atari with Deep Reinforcement Learning )

Papers

Showing 351–400 of 625 papers

TitleStatusHype
Convex Regularization in Monte-Carlo Tree Search—0
Reinforcement Learning and its Connections with Neuroscience and Psychology—0
RL Unplugged: A Suite of Benchmarks for Offline Reinforcement LearningCode0
On Effective Parallelization of Monte Carlo Tree Search—0
A Decentralized Policy Gradient Approach to Multi-task Reinforcement Learning—0
The Value-Improvement Path: Towards Better Representations for Reinforcement Learning—0
Gradient Monitored Reinforcement Learning—0
A Metric Learning Approach to Anomaly Detection in Video GamesCode0
Local and Global Explanations of Agent Behavior: Integrating Strategy Summaries with Saliency MapsCode0
A Simple Imitation Learning Method via Contrastive Regularization—0
Accelerating Deep Neuroevolution on Distributed FPGAs for Reinforcement Learning Problems—0
Model Based Reinforcement Learning for Atari—0
Explain Your Move: Understanding Agent Actions Using Focused Feature SaliencyCode0
Episodic Reinforcement Learning with Associative Memory—0
Learning Dialog Policies from Weak Demonstrations—0
Self Punishment and Reward Backfill for Deep Q-LearningCode0
Using Generative Adversarial Nets on Atari Games for Feature Extraction in Deep Reinforcement Learning—0
Exploring Unknown States with Action BalanceCode0
Relevance-Guided Modeling of Object Dynamics for Reinforcement Learning—0
A Self-Tuning Actor-Critic Algorithm—0
On Catastrophic Interference in Atari 2600 GamesCode0
Efficiently Guiding Imitation Learning Agents with Human Gaze—0
ConQUR: Mitigating Delusional Bias in Deep Q-learningCode0
Disentangling Controllable Object through Video Prediction Improves Visual Reinforcement Learning—0
Value-driven Hindsight Modelling—0
Data Efficient Training for Reinforcement Learning with Adaptive Behavior Policy Sharing—0
Temporal-adaptive Hierarchical Reinforcement Learning—0
FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback—0
The Past and Present of Imitation Learning: A Citation Chain Study—0
On Bonus Based Exploration Methods In The Arcade Learning Environment—0
Deep Reinforcement Learning with Implicit Human Feedback—0
On the Role of Weight Sharing During Deep Option Learning—0
Speeding up reinforcement learning by combining attention and agency features—0
SLM Lab: A Comprehensive Benchmark and Modular Software Framework for Reproducible Deep Reinforcement LearningCode0
Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionCode0
Adapting Behaviour for Learning Progress—0
Deep Bayesian Reward Learning from Preferences—0
Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning—0
No-Regret Exploration in Goal-Oriented Reinforcement Learning—0
Adversary A3C for Robust Reinforcement Learning—0
Reconciling λ-Returns with Experience ReplayCode0
Maximum Entropy Monte-Carlo Planning—0
R2D2: Reliable and Repeatable Detector and DescriptorCode0
Propagating Uncertainty in Reinforcement Learning via Wasserstein BarycentersCode0
Contrastive Learning of Structured World ModelsCode0
GRIm-RePR: Prioritising Generating Important Features for Pseudo-Rehearsal—0
Memory-Efficient Episodic Control Reinforcement Learning with Dynamic Online k-meansCode0
Sample-Efficient Reinforcement Learning with Maximum Entropy Mellowmax Episodic ControlCode0
Efficient decorrelation of features using Gramian in Reinforcement Learning—0
Gamifying the Vehicle Routing Problem with Stochastic Requests—0
Show:102550
← PrevPage 8 of 13Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GDI-H3(200M frames)Score864—Unverified
2GDI-I3(200M frames)Score864—Unverified
3GDI-I3Score864—Unverified
4GDI-H3Score864—Unverified
5Bootstrapped DQNScore855—Unverified
6FQFScore854.2—Unverified
7R2D2Score837.7—Unverified
8Ape-XScore800.9—Unverified
9Agent57Score790.4—Unverified
10IMPALA (deep)Score787.34—Unverified
#ModelMetricClaimedVerifiedStatus
1GDI-H3(200M frames)Score34—Unverified
2GDI-I3Score34—Unverified
3GDI-H3Score34—Unverified
4TRPO-hashScore34—Unverified
5IQNScore34—Unverified
6NoisyNet-DuelingScore34—Unverified
7QR-DQN-1Score34—Unverified
8Go-ExploreScore34—Unverified
9Bootstrapped DQNScore33.9—Unverified
10C51 noopScore33.9—Unverified
#ModelMetricClaimedVerifiedStatus
1Agent57Score580,328.14—Unverified
2QR-DQN-1Score572,510—Unverified
3R2D2Score408,850—Unverified
4IMPALA (deep)Score351,200.12—Unverified
5Ape-XScore302,391.3—Unverified
6A2C + SILScore104,975.6—Unverified
7MuZero (Res2 Adam)Score94,906.25—Unverified
8DreamerV2Score94,688—Unverified
9MuZeroScore72,276—Unverified
10DNAScore52,398—Unverified
#ModelMetricClaimedVerifiedStatus
1GDI-H3Score1,000,000—Unverified
2GDI-H3(200M frames)Score1,000,000—Unverified
3Agent57Score999,997.63—Unverified
4R2D2Score999,996.7—Unverified
5MuZeroScore999,976.52—Unverified
6MuZero (Res2 Adam)Score999,659.18—Unverified
7GDI-I3Score943,910—Unverified
8Ape-XScore392,952.3—Unverified
9C51 noopScore266,434—Unverified
10Duel noopScore50,254.2—Unverified