Pith. sign in

Let offline rl flow: Training conservative agents in the latent space of normalizing flows

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

years

2026 3

representative citing papers

Generative OOD-regularized Model-based Policy Optimization

cs.LG · 2026-05-23 · unverdicted · novelty 6.0

GORMPO uses generative models for density-based regularization in model-based offline RL, outperforming baselines by 17% on a medical dataset while providing theoretical guarantees under mild assumptions.

citing papers explorer

Showing 3 of 3 citing papers.