FQL trains a one-step policy to maximize Q-values while distilling a flow-matching behavioral cloning policy, outperforming many offline RL baselines.
In tables, we denote values at or above 95% of the best performance in bold, following OGBench (Park et al., 2025)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Flow Q-Learning
FQL trains a one-step policy to maximize Q-values while distilling a flow-matching behavioral cloning policy, outperforming many offline RL baselines.