SOTAVerified

Atari Games

The Atari 2600 Games task (and dataset) involves training an agent to achieve high game scores.

( Image credit: Playing Atari with Deep Reinforcement Learning )

Papers

Showing 551–600 of 625 papers

TitleStatusHype
Pre-training Neural Networks with Human Demonstrations for Deep Reinforcement Learning—0
Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks—0
Provable Representation Learning for Imitation with Contrastive Fourier Features—0
Reinforcement Learning and its Connections with Neuroscience and Psychology—0
Deep Reinforcement Learning via L-BFGS Optimization—0
Reward Prediction Error as an Exploration Objective in Deep RL—0
Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals—0
Regulatory Focus: Promotion and Prevention Inclinations in Policy Search—0
Reinforcement Learning and Video Games—0
Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards—0
Reinforcement Learning in Strategy-Based and Atari Games: A Review of Google DeepMinds Innovations—0
Reinforcement Learning with Attention that Works: A Self-Supervised Approach—0
Reinforcement Learning with Structured Hierarchical Grammar Representations of Actions—0
Return-Based Contrastive Representation Learning for Reinforcement Learning—0
Return-based Scaling: Yet Another Normalisation Trick for Deep RL—0
Reuse of Neural Modules for General Video Game Playing—0
Rewarding Episodic Visitation Discrepancy for Exploration in Reinforcement Learning—0
Reward Scale Robustness for Proximal Policy Optimization via DreamerV3 Tricks—0
Sample Efficient Deep Neuroevolution in Low Dimensional Latent Space—0
Sample-Efficient Reinforcement Learning through Transfer and Architectural Priors—0
Sample Efficient Reinforcement Learning In Continuous State Spaces: A Perspective Beyond Linearity—0
Selective Eye-gaze Augmentation To Enhance Imitation Learning In Atari Games—0
Self-Imitation Advantage Learning—0
Relevance-Guided Modeling of Object Dynamics for Reinforcement Learning—0
Self-Supervised Structured Representations for Deep Reinforcement Learning—0
A Self-Tuning Actor-Critic Algorithm—0
Shallow Updates for Deep Reinforcement Learning—0
Shared Learning : Enhancing Reinforcement in Q-Ensembles—0
Should I Run Offline Reinforcement Learning or Behavioral Cloning?—0
Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning—0
Simultaneously Updating All Persistence Values in Reinforcement Learning—0
Sketch-Based Linear Value Function Approximation—0
Slot Contrastive Networks: A Contrastive Approach for Representing Objects—0
Soft Decomposed Policy-Critic: Bridging the Gap for Effective Continuous Control with Discrete RL—0
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization—0
Speeding up reinforcement learning by combining attention and agency features—0
State Distribution-aware Sampling for Deep Q-learning—0
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL—0
Strategic Attentive Writer for Learning Macro-Actions—0
Strategy and Benchmark for Converting Deep Q-Networks to Event-Driven Spiking Neural Networks—0
Striving for Simplicity in Off-Policy Deep Reinforcement Learning—0
Successor-Predecessor Intrinsic Exploration—0
SwitchMT: An Adaptive Context Switching Methodology for Scalable Multi-Task Learning in Intelligent Autonomous Agents—0
Tactics of Adversarial Attack on Deep Reinforcement Learning Agents—0
CopyCAT: Taking Control of Neural Policies with Constant Attacks—0
Target Entropy Annealing for Discrete Soft Actor-Critic—0
Target Return Optimizer for Multi-Game Decision Transformer—0
Temporal-adaptive Hierarchical Reinforcement Learning—0
Terminal Prediction as an Auxiliary Task for Deep Reinforcement Learning—0
The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces—0
Show:102550
← PrevPage 12 of 13Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GDI-H3Score864—Unverified
2GDI-H3(200M frames)Score864—Unverified
3GDI-I3(200M frames)Score864—Unverified
4GDI-I3Score864—Unverified
5Bootstrapped DQNScore855—Unverified
6FQFScore854.2—Unverified
7R2D2Score837.7—Unverified
8Ape-XScore800.9—Unverified
9Agent57Score790.4—Unverified
10IMPALA (deep)Score787.34—Unverified
#ModelMetricClaimedVerifiedStatus
1GDI-H3Score34—Unverified
2GDI-H3(200M frames)Score34—Unverified
3GDI-I3Score34—Unverified
4Go-ExploreScore34—Unverified
5QR-DQN-1Score34—Unverified
6NoisyNet-DuelingScore34—Unverified
7IQNScore34—Unverified
8TRPO-hashScore34—Unverified
9Bootstrapped DQNScore33.9—Unverified
10C51 noopScore33.9—Unverified
#ModelMetricClaimedVerifiedStatus
1Agent57Score580,328.14—Unverified
2QR-DQN-1Score572,510—Unverified
3R2D2Score408,850—Unverified
4IMPALA (deep)Score351,200.12—Unverified
5Ape-XScore302,391.3—Unverified
6A2C + SILScore104,975.6—Unverified
7MuZero (Res2 Adam)Score94,906.25—Unverified
8DreamerV2Score94,688—Unverified
9MuZeroScore72,276—Unverified
10DNAScore52,398—Unverified
#ModelMetricClaimedVerifiedStatus
1GDI-H3(200M frames)Score1,000,000—Unverified
2GDI-H3Score1,000,000—Unverified
3Agent57Score999,997.63—Unverified
4R2D2Score999,996.7—Unverified
5MuZeroScore999,976.52—Unverified
6MuZero (Res2 Adam)Score999,659.18—Unverified
7GDI-I3Score943,910—Unverified
8Ape-XScore392,952.3—Unverified
9C51 noopScore266,434—Unverified
10Duel noopScore50,254.2—Unverified