REVIEW 2 cited by
Revisiting Design Choices in Proximal Policy Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Proximal Policy Optimization (PPO) is a popular deep policy gradient algorithm. In standard implementations, PPO regularizes policy updates with clipped probability ratios, and parameterizes policies with either continuous Gaussian distributions or discrete Softmax distributions. These design choices are widely accepted, and motivated by empirical performance comparisons on MuJoCo and Atari benchmarks. We revisit these practices outside the regime of current benchmarks, and expose three failure modes of standard PPO. We explain why standard design choices are problematic in these cases, and show that alternative choices of surrogate objectives and policy parameterizations can prevent the failure modes. We hope that our work serves as a reminder that many algorithmic design choices in reinforcement learning are tied to specific simulation environments. We should not implicitly accept these choices as a standard part of a more general algorithm.
Forward citations
Cited by 2 Pith papers
-
CORE: Constraint-Aware One-Step Reinforcement Learning for Simulation-Guided Neural Network Accelerator Design
A critic-free one-step RL method with a scaling-graph decoder improves sample efficiency for simulation-guided DNN accelerator co-design, reportedly beating GA and HASCO by over an order of magnitude.
-
Inverse Design in Distributed Circuits Using Single-Step Reinforcement Learning
DCIDA uses single-step reinforcement learning with a deterministic action-to-layout mapping to generate distributed resonator circuits that match target transfer functions more closely than prior methods on a surrogat...
Discussion (0). Continue with ORCID to comment.