Pith. sign in

Policy gradient methods for reinforcement learning with function approximation.Advances in neural information processing systems, 12, 1999

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

method 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

method 1

polarities

use method 1

representative citing papers

Logit Dynamics in Softmax Policy Gradient Methods

cs.LG · 2025-06-15 · conditional · novelty 2.0

The L2 norm of the softmax policy gradient logit update is η|A| sqrt(1 - 2P_c + C(P)), where P_c is the chosen action's probability and C(P) is the collision probability.

citing papers explorer

Showing 1 of 1 citing paper.

  • Logit Dynamics in Softmax Policy Gradient Methods cs.LG · 2025-06-15 · conditional · none · ref 2

    The L2 norm of the softmax policy gradient logit update is η|A| sqrt(1 - 2P_c + C(P)), where P_c is the chosen action's probability and C(P) is the collision probability.