SOTAVerified

Benchmarking Deep Reinforcement Learning for Continuous Control

2016-04-22Code Available2· sign in to hype

Yan Duan, Xi Chen, Rein Houthooft, John Schulman, Pieter Abbeel

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

Recently, researchers have made significant progress combining the advances in deep learning for learning feature representations with reinforcement learning. Some notable examples include training agents to play Atari games based on raw pixel data and to acquire advanced manipulation skills using raw sensory inputs. However, it has been difficult to quantify progress in the domain of continuous control due to the lack of a commonly adopted benchmark. In this work, we present a benchmark suite of continuous control tasks, including classic tasks like cart-pole swing-up, tasks with very high state and action dimensionality such as 3D humanoid locomotion, tasks with partial observations, and tasks with hierarchical structure. We report novel findings based on the systematic evaluation of a range of implemented reinforcement learning algorithms. Both the benchmark and reference implementations are released at https://github.com/rllab/rllab in order to facilitate experimental reproducibility and to encourage adoption by other researchers.

Tasks

Benchmark Results

DatasetModelMetricClaimedVerifiedStatus
2D WalkerTRPOScore1,353.8—Unverified
AcrobotTRPOScore-326—Unverified
Acrobot (limited sensors)TRPOScore-83.3—Unverified
Acrobot (noisy observations)TRPOScore-149.6—Unverified
Acrobot (system identifications)TRPOScore-170.9—Unverified
AntTRPOScore730.2—Unverified
Ant + GatheringTRPOScore-0.4—Unverified
Ant + MazeTRPOScore0—Unverified
Cart-Pole BalancingTRPOScore4,869.8—Unverified
Cart-Pole Balancing (limited sensors)TRPOScore960.2—Unverified
Cart-Pole Balancing (noisy observations)TRPOScore606.2—Unverified
Cart-Pole Balancing (system identifications)TRPOScore980.3—Unverified
Double Inverted PendulumTRPOScore4,412.4—Unverified
Full HumanoidTRPOScore287—Unverified
Half-CheetahTRPOScore1,914—Unverified
HopperTRPOScore1,183.3—Unverified
Inverted PendulumTRPOScore247.2—Unverified
Inverted Pendulum (limited sensors)TRPOScore4.5—Unverified
Inverted Pendulum (noisy observations)TRPOScore10.4—Unverified
Inverted Pendulum (system identifications)TRPOScore14.1—Unverified
Mountain CarTRPOScore-61.7—Unverified
Mountain Car (limited sensors)TRPOScore-64.2—Unverified
Mountain Car (noisy observations)TRPOScore-60.2—Unverified
Mountain Car (system identifications)TRPOScore-61.6—Unverified
Simple HumanoidTRPOScore269.7—Unverified
SwimmerTRPOScore96—Unverified
Swimmer + GatheringTRPOScore0—Unverified
Swimmer + MazeTRPOScore0—Unverified

Reproductions