PPO-BR adapts PPO's clipping threshold with entropy and reward signals, claiming faster convergence and lower variance, but the proof is incomplete and the experiments are not verifiable.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization
PPO-BR adapts PPO's clipping threshold with entropy and reward signals, claiming faster convergence and lower variance, but the proof is incomplete and the experiments are not verifiable.