GORMPO uses generative models for density-based regularization in model-based offline RL, outperforming baselines by 17% on a medical dataset while providing theoretical guarantees under mild assumptions.
Let offline rl flow: Training conservative agents in the latent space of normalizing flows
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
SERNF fine-tunes dexterous manipulation policies on real hardware by pairing normalizing-flow policies with action-chunked critics and conservative off-policy RL.
Develops Bishop-style constructive apparatus for geometric sets, integration, extremum theorems, selectors, differential inclusions, Markov chains, and densities in systems and control.
citing papers explorer
-
Generative OOD-regularized Model-based Policy Optimization
GORMPO uses generative models for density-based regularization in model-based offline RL, outperforming baselines by 17% on a medical dataset while providing theoretical guarantees under mild assumptions.
-
SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows
SERNF fine-tunes dexterous manipulation policies on real hardware by pairing normalizing-flow policies with action-chunked critics and conservative off-policy RL.
-
Some Essential Constructive Foundations for Systems and Control
Develops Bishop-style constructive apparatus for geometric sets, integration, extremum theorems, selectors, differential inclusions, Markov chains, and densities in systems and control.