Pith. sign in

REVIEW 2 cited by

Distilling Internet-Scale Vision-Language Models into Embodied Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12507 v2 pith:32UKFNZM submitted 2023-01-29 cs.AI

classification cs.AI
keywords agentslanguageembodiedmodelsagentgroundinternet-scaleobjects
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction-following agents must ground language into their observation and action spaces. Learning to ground language is challenging, typically requiring domain-specific engineering or large quantities of human interaction data. To address this challenge, we propose using pretrained vision-language models (VLMs) to supervise embodied agents. We combine ideas from model distillation and hindsight experience replay (HER), using a VLM to retroactively generate language describing the agent's behavior. Simple prompting allows us to control the supervision signal, teaching an agent to interact with novel objects based on their names (e.g., planes) or their features (e.g., colors) in a 3D rendered environment. Fewshot prompting lets us teach abstract category membership, including pre-existing categories (food vs toys) and ad-hoc ones (arbitrary preferences over objects). Our work outlines a new and effective way to use internet-scale VLMs, repurposing the generic language grounding acquired by such models to teach task-relevant groundings to embodied agents.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Consistent Model-based Adaptation for Visual Reinforcement Learning

    cs.CV 2025-02 conditional novelty 6.0 of 10

    SCMA trains a policy-agnostic observation denoiser, using a pre-trained world model as a clean-distribution reference, and shows improved visual RL performance under distractions.

  2. AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito

    cs.AI 2026-01 conditional novelty 4.0 of 10

    An AI agent combining GraphRAG, static Fortran analysis, and LLM code generation is reported to translate legacy Fortran finite-difference code into Devito, with Grade-A results claimed on roughly three-quarters of 13...

Pith tools