Pith. sign in

REVIEW 2 cited by

Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19400 v4 pith:YPUQFPPL submitted 2024-10-25 cs.LG cs.AI

classification cs.LGcs.AI
keywords offlinescasstatecorrectionactionissueperformancestates
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In offline reinforcement learning (RL), addressing the out-of-distribution (OOD) action issue has been a focus, but we argue that there exists an OOD state issue that also impairs performance yet has been underexplored. Such an issue describes the scenario when the agent encounters states out of the offline dataset during the test phase, leading to uncontrolled behavior and performance degradation. To this end, we propose SCAS, a simple yet effective approach that unifies OOD state correction and OOD action suppression in offline RL. Technically, SCAS achieves value-aware OOD state correction, capable of correcting the agent from OOD states to high-value in-distribution states. Theoretical and empirical results show that SCAS also exhibits the effect of suppressing OOD actions. On standard offline RL benchmarks, SCAS achieves excellent performance without additional hyperparameter tuning. Moreover, benefiting from its OOD state correction feature, SCAS demonstrates enhanced robustness against environmental perturbations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized Environments

    cs.LG 2025-04 conditional novelty 6.0 of 10

    PDTS swaps the UCB rule in robust active task sampling for posterior sampling with diversity regularization, improving worst-case (CVaR) adaptation in meta-RL and domain randomization.

  2. Variational OOD State Correction for Offline Reinforcement Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DASP adds a variational density-aware term to offline RL policy optimization and reports improved average scores on MuJoCo and AntMaze benchmarks.

Pith tools