Replacing the importance sampling ratio with a stop-gradient self-anchored ratio for positive advantages yields unclipped, REINFORCE-equivalent gradients that improve exploration without training instability.
If πθ < π lower, the gradient is nullified to prevent the probability from dropping too low (over-penalization)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma
Replacing the importance sampling ratio with a stop-gradient self-anchored ratio for positive advantages yields unclipped, REINFORCE-equivalent gradients that improve exploration without training instability.