REVIEW 3 cited by
A Closer Look at Deep Policy Gradients
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study how the behavior of deep policy gradient algorithms reflects the conceptual framework motivating their development. To this end, we propose a fine-grained analysis of state-of-the-art methods based on key elements of this framework: gradient estimation, value prediction, and optimization landscapes. Our results show that the behavior of deep policy gradient algorithms often deviates from what their motivating framework would predict: the surrogate objective does not match the true reward landscape, learned value estimators fail to fit the true value function, and gradient estimates poorly correlate with the "true" gradient. The mismatch between predicted and empirical behavior we uncover highlights our poor understanding of current methods, and indicates the need to move beyond current benchmark-centric evaluation methods.
Forward citations
Cited by 3 Pith papers
-
Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
REINFORCE's gradient-estimator noise-to-signal ratio is exactly computable for linear and polynomial systems, and it typically blows up as policies approach deterministic optima.
-
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
PPO plateaus can be avoided by increasing the number of parallel environments, which reduces both the outer-loop step size and update noise; scaling to 1M environments sustained improvement to 1T transitions.
-
OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization
HARPO reweights per-sample and per-task advantages in GRPO to balance multitask learning, yielding a 7B model with the best average rank across 10 social-behavior tasks and top zero-shot scores on two held-out benchmarks.
Discussion (0). Continue with ORCID to comment.