Pith. sign in

REVIEW 4 cited by

Language Reward Modulation for Pretraining Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.12270 v1 pith:ZSHAAIW2 submitted 2023-08-23 cs.LG cs.AI

classification cs.LGcs.AI
keywords pretraininglearninglrfsrewardrewardslampreinforcementtextbf
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Using learned reward functions (LRFs) as a means to solve sparse-reward reinforcement learning (RL) tasks has yielded some steady progress in task-complexity through the years. In this work, we question whether today's LRFs are best-suited as a direct replacement for task rewards. Instead, we propose leveraging the capabilities of LRFs as a pretraining signal for RL. Concretely, we propose $\textbf{LA}$nguage Reward $\textbf{M}$odulated $\textbf{P}$retraining (LAMP) which leverages the zero-shot capabilities of Vision-Language Models (VLMs) as a $\textit{pretraining}$ utility for RL as opposed to a downstream task reward. LAMP uses a frozen, pretrained VLM to scalably generate noisy, albeit shaped exploration rewards by computing the contrastive alignment between a highly diverse collection of language instructions and the image observations of an agent in its pretraining environment. LAMP optimizes these rewards in conjunction with standard novelty-seeking exploration rewards with reinforcement learning to acquire a language-conditioned, pretrained policy. Our VLM pretraining approach, which is a departure from previous attempts to use LRFs, can warmstart sample-efficient learning on robot manipulation tasks in RLBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following

    cs.AI 2025-09 conditional novelty 6.0 of 10

    ExRAP couples LLM planning with a temporal knowledge-graph memory and information-based exploration, improving success and efficiency for continual embodied instruction following.

  3. World Model Implanting for Test-time Adaptation of Embodied Agents

    cs.AI 2025-09 conditional novelty 6.0 of 10

    WorMI adapts an LLM-based embodied policy to unseen domains by retrieving and compositionally implanting domain-specific world models at test time.

  4. Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    PrefVLM combines VLM-generated trajectory preferences with selective human feedback and inverse-dynamics VLM adaptation, matching PEBBLE on five Meta-World tasks with up to 2x fewer human labels.

Pith tools