Hyperparameter Selection for Offline Reinforcement Learning

Alexander Novikov; Andrea Michi; Caglar Gulcehre; Cosmin Paduraru; Konrad Zolna; Nando de Freitas; Tom Le Paine; Ziyu Wang

arxiv: 2007.09055 · v1 · pith:RAIUAFB4new · submitted 2020-07-17 · 💻 cs.LG · cs.AI· stat.ML

Hyperparameter Selection for Offline Reinforcement Learning

Tom Le Paine , Cosmin Paduraru , Andrea Michi , Caglar Gulcehre , Konrad Zolna , Alexander Novikov , Ziyu Wang , Nando de Freitas This is my paper

classification 💻 cs.LG cs.AIstat.ML

keywords offlinehyperparameterselectionpoliciesbestchoicesdatafactors

0 comments

read the original abstract

Offline reinforcement learning (RL purely from logged data) is an important avenue for deploying RL techniques in real-world scenarios. However, existing hyperparameter selection methods for offline RL break the offline assumption by evaluating policies corresponding to each hyperparameter setting in the environment. This online execution is often infeasible and hence undermines the main aim of offline RL. Therefore, in this work, we focus on \textit{offline hyperparameter selection}, i.e. methods for choosing the best policy from a set of many policies trained using different hyperparameters, given only logged data. Through large-scale empirical evaluation we show that: 1) offline RL algorithms are not robust to hyperparameter choices, 2) factors such as the offline RL algorithm and method for estimating Q values can have a big impact on hyperparameter selection, and 3) when we control those factors carefully, we can reliably rank policies across hyperparameter choices, and therefore choose policies which are close to the best policy in the set. Overall, our results present an optimistic view that offline hyperparameter selection is within reach, even in challenging tasks with pixel observations, high dimensional action spaces, and long horizon.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

$\text{DT}^2$: Decision-Targeted Digital Twins
cs.LG 2026-06 unverdicted novelty 7.0

DT² trains digital twins to preserve pairwise policy rankings from fitted Q-evaluation on offline data rather than minimizing one-step transition errors, improving policy ranking and reducing decision regret.
Sample-efficient inductive matrix completion with noise and inexact side-information
stat.ML 2026-05 unverdicted novelty 7.0

Nonconvex projected gradient descent for noisy inductive matrix completion achieves linear convergence and order-optimal error at sample complexity scaling with side-information dimension a instead of ambient dimension n.
Sample-efficient inductive matrix completion with noise and inexact side-information
stat.ML 2026-05 unverdicted novelty 7.0

A projected gradient descent algorithm for noisy inductive matrix completion achieves linear convergence and stable recovery at sample complexity governed by side-information dimension, extending to inexact side-infor...
Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents
cs.LG 2026-06 unverdicted novelty 6.0

ADWM learns a latent diffusion world model with per-transition independent denoising and policy-conditioned guidance to enable accurate offline evaluation of LLM agent policies.
Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning
cs.LG 2026-05 unverdicted novelty 6.0

An adaptive UCB-based policy selection and fine-tuning strategy improves performance over standard O2O-RL baselines under interaction budgets.
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
cs.RO 2021-08 accept novelty 6.0

A comprehensive benchmark study of offline imitation learning methods on multi-stage robot manipulation tasks identifies key sensitivities to algorithm design, data quality, and stopping criteria while releasing all d...
Some Essential Constructive Foundations for Systems and Control
eess.SY 2026-06 unverdicted novelty 5.0

Develops Bishop-style constructive apparatus for geometric sets, integration, extremum theorems, selectors, differential inclusions, Markov chains, and densities in systems and control.