Pith. sign in

REVIEW 2 cited by

Generalization, Mayhems and Limits in Recurrent Proximal Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.11104 v1 pith:FIJ7BXNS submitted 2022-05-23 cs.LG

classification cs.LG
keywords recurrentagentagentschallengeenvironmentsgeneralizationmayhemmemory
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

At first sight it may seem straightforward to use recurrent layers in Deep Reinforcement Learning algorithms to enable agents to make use of memory in the setting of partially observable environments. Starting from widely used Proximal Policy Optimization (PPO), we highlight vital details that one must get right when adding recurrence to achieve a correct and efficient implementation, namely: properly shaping the neural net's forward pass, arranging the training data, correspondingly selecting hidden states for sequence beginnings and masking paddings for loss computation. We further explore the limitations of recurrent PPO by benchmarking the contributed novel environments Mortar Mayhem and Searing Spotlights that challenge the agent's memory beyond solely capacity and distraction tasks. Remarkably, we can demonstrate a transition to strong generalization in Mortar Mayhem when scaling the number of training seeds, while the agent does not succeed on Searing Spotlights, which seems to be a tough challenge for memory-based agents.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AlphaZeroBeta: Deep Reinforcement Learning for Market-Neutral Portfolios

    q-fin.PM 2026-07 conditional novelty 5.0 of 10

    A deep RL policy with a composite reward and hard dollar-neutral projection beat convex baselines on Sharpe with near-zero benchmark correlation in seven-equity-index walk-forward backtests.

  2. Towards Safe and Robust Autonomous Vehicle Platooning: A Self-Organizing Cooperative Control Framework

    cs.RO 2024-08 unverdicted novelty 3.0 of 10

    TriCoD is a cooperative decision-making framework using twin-world deduction and adaptive switching between DRL and model-driven methods to enable safe, dynamic AV platooning in hybrid traffic.

Pith tools