Pith. sign in

REVIEW 1 cited by

GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.04191 v2 pith:NNLAQP4I submitted 2025-04-05 cs.CV cs.RO

classification cs.CVcs.RO
keywords learningphysicalgroveopen-vocabularyrewardskillwhileacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning open-vocabulary physical skills for simulated agents presents a significant challenge in artificial intelligence. Current reinforcement learning approaches face critical limitations: manually designed rewards lack scalability across diverse tasks, while demonstration-based methods struggle to generalize beyond their training distribution. We introduce GROVE, a generalized reward framework that enables open-vocabulary physical skill learning without manual engineering or task-specific demonstrations. Our key insight is that Large Language Models(LLMs) and Vision Language Models(VLMs) provide complementary guidance -- LLMs generate precise physical constraints capturing task requirements, while VLMs evaluate motion semantics and naturalness. Through an iterative design process, VLM-based feedback continuously refines LLM-generated constraints, creating a self-improving reward system. To bridge the domain gap between simulation and natural images, we develop Pose2CLIP, a lightweight mapper that efficiently projects agent poses directly into semantic feature space without computationally expensive rendering. Extensive experiments across diverse embodiments and learning paradigms demonstrate GROVE's effectiveness, achieving 22.2% higher motion naturalness and 25.7% better task completion scores while training 8.4x faster than previous methods. These results establish a new foundation for scalable physical skill acquisition in simulated environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboMoRe: LLM-based Robot Co-design via Joint Optimization of Morphology and Reward

    cs.RO 2025-05 reject novelty 6.0 of 10

    An LLM-driven framework that jointly proposes robot morphologies and reward functions, using diversity reflection and alternating refinement, claims large efficiency gains over baselines across eight locomotion tasks.

Pith tools