Pith. sign in

REVIEW 3 cited by

A Closer Look at Deep Policy Gradients

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.02553 v4 pith:HWMCJR4P submitted 2018-11-06 cs.LG cs.NEcs.ROstat.ML

classification cs.LGcs.NEcs.ROstat.ML
keywords gradientbehaviordeepframeworkmethodspolicytruevalue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study how the behavior of deep policy gradient algorithms reflects the conceptual framework motivating their development. To this end, we propose a fine-grained analysis of state-of-the-art methods based on key elements of this framework: gradient estimation, value prediction, and optimization landscapes. Our results show that the behavior of deep policy gradient algorithms often deviates from what their motivating framework would predict: the surrogate objective does not match the true reward landscape, learned value estimators fail to fit the true value function, and gradient estimates poorly correlate with the "true" gradient. The mismatch between predicted and empirical behavior we uncover highlights our poor understanding of current methods, and indicates the need to move beyond current benchmark-centric evaluation methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator

    math.OC 2026-02 conditional novelty 7.0 of 10

    REINFORCE's gradient-estimator noise-to-signal ratio is exactly computable for linear and polynomial systems, and it typically blows up as policies approach deterministic optima.

  2. Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments

    cs.LG 2026-03 conditional novelty 5.0 of 10

    PPO plateaus can be avoided by increasing the number of parallel environments, which reduces both the outer-loop step size and update noise; scaling to 1M environments sustained improvement to 1T transitions.

  3. OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization

    cs.AI 2026-02 conditional novelty 5.0 of 10

    HARPO reweights per-sample and per-task advantages in GRPO to balance multitask learning, yielding a 7B model with the best average rank across 10 social-behavior tasks and top zero-shot scores on two held-out benchmarks.

Pith tools