← back to paper
arxiv: 2606.00680 · 2 revisions
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief