Establishes global linear convergence of entropy-regularized policy gradient in continuous MDPs with log-linear softmax policies under Q-realizability by bounding non-uniform PL constants in two feature regimes.
Rethinking the global convergence of softmax policy gradient with linear function approximation
2 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2roles
other 1polarities
unclear 1representative citing papers
Certain errors in proxy rewards for policy gradient methods can be benign or beneficial by preventing policies from stalling on outputs with mediocre ground truth rewards, enabling improved RLHF metrics and reward design insights.
citing papers explorer
-
Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs
Establishes global linear convergence of entropy-regularized policy gradient in continuous MDPs with log-linear softmax policies under Q-realizability by bounding non-uniform PL constants in two feature regimes.
-
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
Certain errors in proxy rewards for policy gradient methods can be benign or beneficial by preventing policies from stalling on outputs with mediocre ground truth rewards, enabling improved RLHF metrics and reward design insights.