Pith. sign in

REVIEW 2 cited by

Distributionally Robust Offline Reinforcement Learning with Linear Function Approximation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.06620 v3 pith:BWORDNSR submitted 2022-09-14 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords environmentapproximationdatasetdistributionallyfunctionlinearpolicyrobust
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Among the reasons hindering reinforcement learning (RL) applications to real-world problems, two factors are critical: limited data and the mismatch between the testing environment (real environment in which the policy is deployed) and the training environment (e.g., a simulator). This paper attempts to address these issues simultaneously with distributionally robust offline RL, where we learn a distributionally robust policy using historical data obtained from the source environment by optimizing against a worst-case perturbation thereof. In particular, we move beyond tabular settings and consider linear function approximation. More specifically, we consider two settings, one where the dataset is well-explored and the other where the dataset has sufficient coverage of the optimal policy. We propose two algorithms~-- one for each of the two settings~-- that achieve error bounds $\tilde{O}(d^{1/2}/N^{1/2})$ and $\tilde{O}(d^{3/2}/N^{1/2})$ respectively, where $d$ is the dimension in the linear function approximation and $N$ is the number of trajectories in the dataset. To the best of our knowledge, they provide the first non-asymptotic results of the sample complexity in this setting. Diverse experiments are conducted to demonstrate our theoretical findings, showing the superiority of our algorithm against the non-robust one.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hybrid Cross-domain Robust Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    HYDRO combines a small offline robust RL dataset with a mismatched online simulator, filtering simulator samples by uncertainty and gap to the worst-case model to improve robust policy performance.

  2. Linear Mixture Distributionally Robust Markov Decision Processes

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Introduces linear mixture distributionally robust MDPs and proves offline suboptimality bounds of order 1/sqrt(K) for TV, KL, and chi-squared uncertainty sets.

Pith tools