Pith. sign in

For PQN, we use the CleanRL implementation (Huang et al., 2022)

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

StaQ it! Growing neural networks for Policy Mirror Descent

cs.LG · 2025-06-16 · conditional · novelty 7.0

StaQ, a finite-memory Policy Mirror Descent algorithm, converges to the optimal entropy-regularized policy with a sufficiently large window of past Q-functions and performs competitively with baselines.

citing papers explorer

Showing 1 of 1 citing paper.

  • StaQ it! Growing neural networks for Policy Mirror Descent cs.LG · 2025-06-16 · conditional · none · ref 15

    StaQ, a finite-memory Policy Mirror Descent algorithm, converges to the optimal entropy-regularized policy with a sufficiently large window of past Q-functions and performs competitively with baselines.