GenPO++ achieves exact Jacobian-free likelihood ratio computation for generative flow policies by embedding history states as auxiliary memory in a high-order reversible ODE solver.
Maximum entropy reinforcement learning via energy-based normalizing flow
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2representative citing papers
Frictional Q-Learning extends batch-constrained Q-learning with a contrastive autoencoder trained against orthonormal 'friction' actions, reporting wins on Humanoid and Walker2D but losses to TD3 on Ant and DDPG on HalfCheetah.
citing papers explorer
-
GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios
GenPO++ achieves exact Jacobian-free likelihood ratio computation for generative flow policies by embedding history states as auxiliary memory in a high-order reversible ODE solver.
-
Frictional Q-Learning
Frictional Q-Learning extends batch-constrained Q-learning with a contrastive autoencoder trained against orthonormal 'friction' actions, reporting wins on Humanoid and Walker2D but losses to TD3 on Ant and DDPG on HalfCheetah.