Boosting Trust Region Policy Optimization by Normalizing Flows Policy

Shipra Agrawal; Yunhao Tang

arxiv: 1809.10326 · v3 · pith:UYWPOZ5Vnew · submitted 2018-09-27 · 💻 cs.AI · cs.LG· stat.ML

Boosting Trust Region Policy Optimization by Normalizing Flows Policy

Yunhao Tang , Shipra Agrawal This is my paper

classification 💻 cs.AI cs.LGstat.ML

keywords policyflowsnormalizingregiontrustarchitecturesavoidbaseline

0 comments

read the original abstract

We propose to improve trust region policy search with normalizing flows policy. We illustrate that when the trust region is constructed by KL divergence constraints, normalizing flows policy generates samples far from the 'center' of the previous policy iterate, which potentially enables better exploration and helps avoid bad local optima. Through extensive comparisons, we show that the normalizing flows policy significantly improves upon baseline architectures especially on high-dimensional tasks with complex dynamics.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios
cs.LG 2026-06 unverdicted novelty 6.0

GenPO++ achieves exact Jacobian-free likelihood ratio computation for generative flow policies by embedding history states as auxiliary memory in a high-order reversible ODE solver.