Pith. sign in

REVIEW 1 cited by

Contrastive Representation for Data Filtering in Cross-Domain Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.06192 v1 pith:QMRN7BFR submitted 2024-05-10 cs.LG cs.AI

classification cs.LGcs.AI
keywords datadomaindomainsdynamicscontrastiveperformancetargetcross-domain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cross-domain offline reinforcement learning leverages source domain data with diverse transition dynamics to alleviate the data requirement for the target domain. However, simply merging the data of two domains leads to performance degradation due to the dynamics mismatch. Existing methods address this problem by measuring the dynamics gap via domain classifiers while relying on the assumptions of the transferability of paired domains. In this paper, we propose a novel representation-based approach to measure the domain gap, where the representation is learned through a contrastive objective by sampling transitions from different domains. We show that such an objective recovers the mutual-information gap of transition functions in two domains without suffering from the unbounded issue of the dynamics gap in handling significantly different domains. Based on the representations, we introduce a data filtering algorithm that selectively shares transitions from the source domain according to the contrastive score functions. Empirical results on various tasks demonstrate that our method achieves superior performance, using only 10% of the target data to achieve 89.2% of the performance on 100% target dataset with state-of-the-art methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    DADiff estimates cross-domain dynamics mismatch from diffusion-model latent-state trajectories and uses it for reward modification or data selection in policy adaptation.

Pith tools