Pith. sign in

REVIEW 2 cited by

Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.06687 v2 pith:227NWNLE submitted 2023-09-13 cs.RO cs.AI

classification cs.ROcs.AI
keywords rewardfunctionframeworklanguageroboticautomateddeepdesign
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although Deep Reinforcement Learning (DRL) has achieved notable success in numerous robotic applications, designing a high-performing reward function remains a challenging task that often requires substantial manual input. Recently, Large Language Models (LLMs) have been extensively adopted to address tasks demanding in-depth common-sense knowledge, such as reasoning and planning. Recognizing that reward function design is also inherently linked to such knowledge, LLM offers a promising potential in this context. Motivated by this, we propose in this work a novel LLM framework with a self-refinement mechanism for automated reward function design. The framework commences with the LLM formulating an initial reward function based on natural language inputs. Then, the performance of the reward function is assessed, and the results are presented back to the LLM for guiding its self-refinement process. We examine the performance of our proposed framework through a variety of continuous robotic control tasks across three diverse robotic systems. The results indicate that our LLM-designed reward functions are able to rival or even surpass manually designed reward functions, highlighting the efficacy and applicability of our approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Uncertainty-aware Reward Design Process

    cs.LG 2025-07 conditional novelty 6.0 of 10

    URDP couples LLM-based reward component design with uncertainty-weighted Bayesian optimization, reporting better reward quality and efficiency than Eureka and Text2Reward on three benchmarks.

  2. Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One

    cs.LG 2025-05 reject novelty 6.0 of 10

    An LLM classifies game states into situations and selects the RL agent with the best historical average reward for each situation, outperforming static ensemble baselines on Atari.

Pith tools