Pith. sign in

REVIEW 2 cited by

Active Fine-Tuning of Multi-Task Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05026 v3 pith:LEYCETFZ submitted 2024-10-07 cs.LG cs.RO

classification cs.LGcs.RO
keywords multi-tasktaskspoliciesactiveadaptationcollectingdemonstrateddemonstrations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained generalist policies are rapidly gaining relevance in robot learning due to their promise of fast adaptation to novel, in-domain tasks. This adaptation often relies on collecting new demonstrations for a specific task of interest and applying imitation learning algorithms, such as behavioral cloning. However, as soon as several tasks need to be learned, we must decide which tasks should be demonstrated and how often? We study this multi-task problem and explore an interactive framework in which the agent adaptively selects the tasks to be demonstrated. We propose AMF (Active Multi-task Fine-tuning), an algorithm to maximize multi-task policy performance under a limited demonstration budget by collecting demonstrations yielding the largest information gain on the expert policy. We derive performance guarantees for AMF under regularity assumptions and demonstrate its empirical effectiveness to efficiently fine-tune neural policies in complex and high-dimensional environments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Arnold: a generalist muscle transformer policy

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A single transformer policy with a compositional sensorimotor vocabulary achieves expert or super-expert performance on 14 musculoskeletal control tasks spanning four embodiments.

  2. Epistemically-guided forward-backward exploration

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Choosing exploration policies by the ensemble disagreement of forward-backward value estimates improves zero-shot RL sample efficiency on DeepMind Control Suite tasks.

Pith tools