Pith. sign in

REVIEW 1 cited by

Cross-Domain Policy Adaptation by Capturing Representation Mismatch

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15369 v1 pith:OPEOVFHI submitted 2024-05-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords domainrepresentationdynamicslearningmismatchsourcetargetonly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It is vital to learn effective policies that can be transferred to different domains with dynamics discrepancies in reinforcement learning (RL). In this paper, we consider dynamics adaptation settings where there exists dynamics mismatch between the source domain and the target domain, and one can get access to sufficient source domain data, while can only have limited interactions with the target domain. Existing methods address this problem by learning domain classifiers, performing data filtering from a value discrepancy perspective, etc. Instead, we tackle this challenge from a decoupled representation learning perspective. We perform representation learning only in the target domain and measure the representation deviations on the transitions from the source domain, which we show can be a signal of dynamics mismatch. We also show that representation deviation upper bounds performance difference of a given policy in the source domain and target domain, which motivates us to adopt representation deviation as a reward penalty. The produced representations are not involved in either policy or value function, but only serve as a reward penalizer. We conduct extensive experiments on environments with kinematic and morphology mismatch, and the results show that our method exhibits strong performance on many tasks. Our code is publicly available at https://github.com/dmksjfl/PAR.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    DADiff estimates cross-domain dynamics mismatch from diffusion-model latent-state trajectories and uses it for reward modification or data selection in policy adaptation.

Pith tools