Pith. sign in

REVIEW 4 cited by

Vision-Language Models Provide Promptable Representations for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02651 v3 pith:NL73225P submitted 2024-02-05 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords approachembeddingsknowledgepoliciesrepresentationstrainedvlmsbehaviors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans can quickly learn new behaviors by leveraging background world knowledge. In contrast, agents trained with reinforcement learning (RL) typically learn behaviors from scratch. We thus propose a novel approach that uses the vast amounts of general and indexable world knowledge encoded in vision-language models (VLMs) pre-trained on Internet-scale data for embodied RL. We initialize policies with VLMs by using them as promptable representations: embeddings that encode semantic features of visual observations based on the VLM's internal knowledge and reasoning capabilities, as elicited through prompts that provide task context and auxiliary information. We evaluate our approach on visually-complex, long horizon RL tasks in Minecraft and robot navigation in Habitat. We find that our policies trained on embeddings from off-the-shelf, general-purpose VLMs outperform equivalent policies trained on generic, non-promptable image embeddings. We also find our approach outperforms instruction-following methods and performs comparably to domain-specific embeddings. Finally, we show that our approach can use chain-of-thought prompting to produce representations of common-sense semantic reasoning, improving policy performance in novel scenes by 1.5 times.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Octo: An Open-Source Generalist Robot Policy

    cs.RO 2024-05 unverdicted novelty 6.0 of 10

    Octo is an open-source transformer-based generalist robot policy pretrained on 800k trajectories that serves as an effective initialization for finetuning across diverse robotic platforms.

  2. Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    SALSA aligns social features and adds future-risk signals in VLA models to cut near-collisions by 86.4% and raise social accuracy from 53% to 93% on SCAND and real robots.

  3. Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning

    cs.AI 2025-09 conditional novelty 5.0 of 10

    A curriculum-trained DRL agent with a VLM query action and layered rewards is claimed to improve semantic exploration and object discovery in AI2-THOR.

  4. Interpret Policies in Deep Reinforcement Learning using SILVER with RL-Guided Labeling: A Model-level Approach to High-dimensional and Multi-action Environments

    cs.LG 2025-10 reject novelty 4.0 of 10

    SILVER with RL-guided labeling: SHAP plus clustering plus policy-query labels plus decision trees or regression to interpret multi-action Atari policies.

Pith tools