Pith. sign in

REVIEW 1 cited by

Evolved Policy Gradients

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.04821 v2 pith:55EG7L3F submitted 2018-02-13 cs.LG cs.AI

classification cs.LGcs.AI
keywords losslearningpolicyagentalgorithmsevolvedgradientmetalearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a metalearning approach for learning gradient-based reinforcement learning (RL) algorithms. The idea is to evolve a differentiable loss function, such that an agent, which optimizes its policy to minimize this loss, will achieve high rewards. The loss is parametrized via temporal convolutions over the agent's experience. Because this loss is highly flexible in its ability to take into account the agent's history, it enables fast task learning. Empirical results show that our evolved policy gradient algorithm (EPG) achieves faster learning on several randomized environments compared to an off-the-shelf policy gradient method. We also demonstrate that EPG's learned loss can generalize to out-of-distribution test time tasks, and exhibits qualitatively different behavior from other popular metalearning algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Resetting only Adam's timestep counter at objective changes, not its momentum, improves RL performance on Atari and Craftax.

Pith tools