SOTAVerified

Atari Games

The Atari 2600 Games task (and dataset) involves training an agent to achieve high game scores.

( Image credit: Playing Atari with Deep Reinforcement Learning )

Papers

Showing 301–350 of 625 papers

TitleStatusHype
Combining Off and On-Policy Training in Model-Based Reinforcement Learning—0
Return-Based Contrastive Representation Learning for Reinforcement Learning—0
Q-Value Weighted Regression: Reinforcement Learning with Limited DataCode0
Measuring Progress in Deep Reinforcement Learning Sample Efficiency—0
Learning State Representations from Random Deep Action-conditional PredictionsCode0
An advantage actor-critic algorithm for robotic motion planning in dense and dynamic scenarios—0
Revisiting Prioritized Experience Replay: A Value PerspectiveCode0
Shielding Atari Games with Bounded PrescienceCode0
Benchmarking Perturbation-based Saliency Maps for Explaining Atari AgentsCode0
Hierarchical Width-Based Planning and LearningCode0
Deep Reinforcement Learning with Quantum-inspired Experience Replay—0
Compute- and Memory-Efficient Reinforcement Learning with Latent Experience Replay—0
Unsupervised Task Clustering for Multi-Task Reinforcement LearningCode0
Unsupervised Active Pre-Training for Reinforcement Learning—0
Adaptive N-step Bootstrapping with Off-policy Data—0
Learning Efficient Planning-based Rewards for Imitation Learning—0
Deep Q-Learning with Low Switching Cost—0
Optimistic Exploration with Backward Bootstrapped Bonus for Deep Reinforcement Learning—0
Towards Understanding Asynchronous Advantage Actor-critic: Convergence and Linear Speedup—0
Self-Imitation Advantage Learning—0
Evaluating Agents without RewardsCode0
Learning to Stop: Dynamic Simulation Monte-Carlo Tree Search—0
Selective Eye-gaze Augmentation To Enhance Imitation Learning In Atari Games—0
Non-Crossing Quantile Regression for Distributional Reinforcement Learning—0
Optimizing the Neural Architecture of Reinforcement Learning AgentsCode0
Predictive PER: Balancing Priority and Diversity towards Stable Deep Reinforcement Learning—0
Leveraging the Variance of Return Sequences for Exploration Policy—0
Perturbation-based exploration methods in deep reinforcement learning—0
Machine versus Human Attention in Deep Reinforcement Learning Tasks—0
Learning to Represent Action Values as a Hypergraph on the Action Vertices—0
VisualHints: A Visual-Lingual Environment for Multimodal Reinforcement Learning—0
Explanation Augmented Feedback in Human-in-the-Loop Reinforcement Learning—0
Smaller World Models for Reinforcement Learning—0
Student-Initiated Action Advising via Advice NoveltyCode0
Strategy and Benchmark for Converting Deep Q-Networks to Event-Driven Spiking Neural Networks—0
Lucid Dreaming for Experience Replay: Refreshing Past States with the Current PolicyCode0
Population-Guided Imitation Learning—0
Multiplayer Support for the Arcade Learning Environment—0
Reconstructing Actions To Explain Deep Reinforcement Learning—0
Evolutionary Reinforcement Learning via Cooperative Coevolutionary Negatively Correlated Search—0
Ranking Policy DecisionsCode0
Model-Free Episodic Control with State Aggregation—0
Adversary Agnostic Robust Deep Reinforcement Learning—0
Noisy Agents: Self-supervised Exploration by Predicting Auditory Events—0
Slot Contrastive Networks: A Contrastive Approach for Representing Objects—0
Analysis of Q-learning with Adaptation and Momentum Restart for Gradient Descent—0
DinerDash Gym: A Benchmark for Policy Learning in High-Dimensional Action SpaceCode0
Learning Abstract Models for Strategic Exploration and Fast Reward Transfer—0
Attention or memory? Neurointerpretable agents in space and time—0
Double Prioritized State Recycled Experience Replay—0
Show:102550
← PrevPage 7 of 13Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GDI-I3(200M frames)Score864—Unverified
2GDI-H3Score864—Unverified
3GDI-I3Score864—Unverified
4GDI-H3(200M frames)Score864—Unverified
5Bootstrapped DQNScore855—Unverified
6FQFScore854.2—Unverified
7R2D2Score837.7—Unverified
8Ape-XScore800.9—Unverified
9Agent57Score790.4—Unverified
10IMPALA (deep)Score787.34—Unverified
#ModelMetricClaimedVerifiedStatus
1GDI-H3(200M frames)Score34—Unverified
2Go-ExploreScore34—Unverified
3GDI-H3Score34—Unverified
4GDI-I3Score34—Unverified
5QR-DQN-1Score34—Unverified
6NoisyNet-DuelingScore34—Unverified
7IQNScore34—Unverified
8TRPO-hashScore34—Unverified
9ASL DDQNScore33.9—Unverified
10C51 noopScore33.9—Unverified
#ModelMetricClaimedVerifiedStatus
1Agent57Score580,328.14—Unverified
2QR-DQN-1Score572,510—Unverified
3R2D2Score408,850—Unverified
4IMPALA (deep)Score351,200.12—Unverified
5Ape-XScore302,391.3—Unverified
6A2C + SILScore104,975.6—Unverified
7MuZero (Res2 Adam)Score94,906.25—Unverified
8DreamerV2Score94,688—Unverified
9MuZeroScore72,276—Unverified
10DNAScore52,398—Unverified
#ModelMetricClaimedVerifiedStatus
1GDI-H3(200M frames)Score1,000,000—Unverified
2GDI-H3Score1,000,000—Unverified
3Agent57Score999,997.63—Unverified
4R2D2Score999,996.7—Unverified
5MuZeroScore999,976.52—Unverified
6MuZero (Res2 Adam)Score999,659.18—Unverified
7GDI-I3Score943,910—Unverified
8Ape-XScore392,952.3—Unverified
9C51 noopScore266,434—Unverified
10Duel noopScore50,254.2—Unverified