Pith. sign in

REVIEW 1 cited by

Deterministic Uncertainty Propagation for Improved Model-Based Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04088 v3 pith:VVYNVDQG submitted 2024-06-06 cs.LG

classification cs.LG
keywords approachesdeterministicbellmancarlomodel-basedmombomonteoffline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current approaches to model-based offline reinforcement learning often incorporate uncertainty-based reward penalization to address the distributional shift problem. These approaches, commonly known as pessimistic value iteration, use Monte Carlo sampling to estimate the Bellman target to perform temporal difference-based policy evaluation. We find out that the randomness caused by this sampling step significantly delays convergence. We present a theoretical result demonstrating the strong dependency of suboptimality on the number of Monte Carlo samples taken per Bellman target calculation. Our main contribution is a deterministic approximation to the Bellman target that uses progressive moment matching, a method developed originally for deterministic variational inference. The resulting algorithm, which we call Moment Matching Offline Model-Based Policy Optimization (MOMBO), propagates the uncertainty of the next state through a nonlinear Q-network in a deterministic fashion by approximating the distributions of hidden layer activations by a normal distribution. We show that it is possible to provide tighter guarantees for the suboptimality of MOMBO than the existing Monte Carlo sampling approaches. We also observe MOMBO to converge faster than these approaches in a large set of benchmark tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Offline Diffusion Policy for Multi-User Delay-Constrained Scheduling

    cs.AI 2025-01 conditional novelty 4.0 of 10

    SOCD trains a diffusion-based scheduling policy offline, selects actions via a critic score, and tunes a Lagrange multiplier from the offline dataset to satisfy resource constraints.

Pith tools