Pith. sign in

REVIEW 1 cited by

QMP: Q-switch Mixture of Policies for Multi-Task Behavior Sharing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.00671 v3 pith:MMXD4PUQ submitted 2023-02-01 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords taskpoliciessharingbehaviorsmtrldatatasksbehavior
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-task reinforcement learning (MTRL) aims to learn several tasks simultaneously for better sample efficiency than learning them separately. Traditional methods achieve this by sharing parameters or relabeled data between tasks. In this work, we introduce a new framework for sharing behavioral policies across tasks, which can be used in addition to existing MTRL methods. The key idea is to improve each task's off-policy data collection by employing behaviors from other task policies. Selectively sharing helpful behaviors acquired in one task to collect training data for another task can lead to higher-quality trajectories, leading to more sample-efficient MTRL. Thus, we introduce a simple and principled framework called Q-switch mixture of policies (QMP) that selectively shares behavior between different task policies by using the task's Q-function to evaluate and select useful shareable behaviors. We theoretically analyze how QMP improves the sample efficiency of the underlying RL algorithm. Our experiments show that QMP's behavioral policy sharing provides complementary gains over many popular MTRL algorithms and outperforms alternative ways to share behaviors in various manipulation, locomotion, and navigation environments. Videos are available at https://qmp-mtrl.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A per-task guide policy selects other tasks' control policies to generate training trajectories, boosting multi-task RL performance across five baselines.

Pith tools