Pith. sign in

REVIEW 2 cited by

Upper and Lower Bounds for Distributionally Robust Off-Dynamics Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20521 v1 pith:MC6Q3DHR submitted 2024-09-30 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords mathcalalgorithmlearningpolicyrobustdecisiondistributionallydrmdps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study off-dynamics Reinforcement Learning (RL), where the policy training and deployment environments are different. To deal with this environmental perturbation, we focus on learning policies robust to uncertainties in transition dynamics under the framework of distributionally robust Markov decision processes (DRMDPs), where the nominal and perturbed dynamics are linear Markov Decision Processes. We propose a novel algorithm We-DRIVE-U that enjoys an average suboptimality $\widetilde{\mathcal{O}}\big({d H \cdot \min \{1/{\rho}, H\}/\sqrt{K} }\big)$, where $K$ is the number of episodes, $H$ is the horizon length, $d$ is the feature dimension and $\rho$ is the uncertainty level. This result improves the state-of-the-art by $\mathcal{O}(dH/\min\{1/\rho,H\})$. We also construct a novel hard instance and derive the first information-theoretic lower bound in this setting, which indicates our algorithm is near-optimal up to $\mathcal{O}(\sqrt{H})$ for any uncertainty level $\rho\in(0,1]$. Our algorithm also enjoys a 'rare-switching' design, and thus only requires $\mathcal{O}(dH\log(1+H^2K))$ policy switches and $\mathcal{O}(d^2H\log(1+H^2K))$ calls for oracle to solve dual optimization problems, which significantly improves the computational efficiency of existing algorithms for DRMDPs, whose policy switch and oracle complexities are both $\mathcal{O}(K)$.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linear Mixture Distributionally Robust Markov Decision Processes

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Introduces linear mixture distributionally robust MDPs and proves offline suboptimality bounds of order 1/sqrt(K) for TV, KL, and chi-squared uncertainty sets.

  2. A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming

    cs.CR 2025-05 reject novelty 5.0 of 10

    A reward-driven, PPO-finetuned LLM pipeline claims to generate diverse, evasive webshell payloads with higher escape rates than prompt-engineering baselines.

Pith tools