Pith. sign in

Maximum entropy reinforcement learning via energy-based normalizing flow

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

fields

cs.LG 2

years

2026 1 2025 1

representative citing papers

Frictional Q-Learning

cs.LG · 2025-09-24 · reject · novelty 5.0

Frictional Q-Learning extends batch-constrained Q-learning with a contrastive autoencoder trained against orthonormal 'friction' actions, reporting wins on Humanoid and Walker2D but losses to TD3 on Ant and DDPG on HalfCheetah.

citing papers explorer

Showing 2 of 2 citing papers.

  • GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios cs.LG · 2026-06-05 · unverdicted · none · ref 5

    GenPO++ achieves exact Jacobian-free likelihood ratio computation for generative flow policies by embedding history states as auxiliary memory in a high-order reversible ODE solver.

  • Frictional Q-Learning cs.LG · 2025-09-24 · reject · none · ref 4

    Frictional Q-Learning extends batch-constrained Q-learning with a contrastive autoencoder trained against orthonormal 'friction' actions, reporting wins on Humanoid and Walker2D but losses to TD3 on Ant and DDPG on HalfCheetah.