REVIEW 5 cited by
Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent analyses of certain gradient descent optimization methods have shown that performance can degrade in some settings - such as with stochasticity or implicit momentum. In deep reinforcement learning (Deep RL), such optimization methods are often used for training neural networks via the temporal difference error or policy gradient. As an agent improves over time, the optimization target changes and thus the loss landscape (and local optima) change. Due to the failure modes of those methods, the ideal choice of optimizer for Deep RL remains unclear. As such, we provide an empirical analysis of the effects that a wide range of gradient descent optimizers and their hyperparameters have on policy gradient methods, a subset of Deep RL algorithms, for benchmark continuous control tasks. We find that adaptive optimizers have a narrow window of effective learning rates, diverging in other cases, and that the effectiveness of momentum varies depending on the properties of the environment. Our analysis suggests that there is significant interplay between the dynamics of the environment and Deep RL algorithm properties which aren't necessarily accounted for by traditional adaptive gradient methods. We provide suggestions for optimal settings of current methods and further lines of research based on our findings.
Forward citations
Cited by 5 Pith papers
-
Reweighting Adversarial Networks for Unbinned Unfolding
RANs generalize moment unfolding to full phase-space unbinned unfolding via detector-level Wasserstein critics without requiring support overlap or multiple iterations.
-
Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps
Resetting only Adam's timestep counter at objective changes, not its momentum, improves RL performance on Atari and Craftax.
-
Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
EXSEARCH trains LLMs for agentic search by treating search trajectories as latent variables and optimizing a weighted likelihood via expectation-maximization, yielding gains on NQ, HotpotQA, MuSiQue, and 2WikiQA.
-
Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($\Delta$)
SARSA(Delta) decomposes action-value estimates across discount factors, but the claimed gains are not consistently supported by the paper's own tables.
-
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition
A proposed Q-learning port of TD(Delta) decomposes action values by discount factor, but the core Bellman equation for the delta components is derived incorrectly and the claimed experiments are missing.
Discussion (0). Continue with ORCID to comment.