Pith. sign in

REVIEW 1 cited by

Learning Reward for Physical Skills using Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.14092 v1 pith:Q7NR3W66 submitted 2023-10-21 cs.RO cs.AI

classification cs.ROcs.AI
keywords rewardfunctionslearningphysicalskillsfeedbackllmsenvironment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning reward functions for physical skills are challenging due to the vast spectrum of skills, the high-dimensionality of state and action space, and nuanced sensory feedback. The complexity of these tasks makes acquiring expert demonstration data both costly and time-consuming. Large Language Models (LLMs) contain valuable task-related knowledge that can aid in learning these reward functions. However, the direct application of LLMs for proposing reward functions has its limitations such as numerical instability and inability to incorporate the environment feedback. We aim to extract task knowledge from LLMs using environment feedback to create efficient reward functions for physical skills. Our approach consists of two components. We first use the LLM to propose features and parameterization of the reward function. Next, we update the parameters of this proposed reward function through an iterative self-alignment process. In particular, this process minimizes the ranking inconsistency between the LLM and our learned reward functions based on the new observations. We validated our method by testing it on three simulated physical skill learning tasks, demonstrating effective support for our design choices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. "Stack It Up!": 3D Stable Structure Generation from 2D Hand-drawn Sketch

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    StackItUp converts 2D hand-drawn sketches into stable 3D block arrangements using a symbolic relation graph and diffusion-based block pose generation.

Pith tools