Boosting Trust Region Policy Optimization by Normalizing Flows Policy

· 2018 · cs.AI · arXiv 1809.10326

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

We propose to improve trust region policy search with normalizing flows policy. We illustrate that when the trust region is constructed by KL divergence constraints, normalizing flows policy generates samples far from the 'center' of the previous policy iterate, which potentially enables better exploration and helps avoid bad local optima. Through extensive comparisons, we show that the normalizing flows policy significantly improves upon baseline architectures especially on high-dimensional tasks with complex dynamics.

representative citing papers

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios

cs.LG · 2026-06-05 · unverdicted · novelty 6.0

GenPO++ achieves exact Jacobian-free likelihood ratio computation for generative flow policies by embedding history states as auxiliary memory in a high-order reversible ODE solver.

citing papers explorer

Showing 1 of 1 citing paper.

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios cs.LG · 2026-06-05 · unverdicted · none · ref 42 · internal anchor
GenPO++ achieves exact Jacobian-free likelihood ratio computation for generative flow policies by embedding history states as auxiliary memory in a high-order reversible ODE solver.

Boosting Trust Region Policy Optimization by Normalizing Flows Policy

fields

years

verdicts

representative citing papers

citing papers explorer