Pith. sign in

REVIEW 1 cited by

Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.10861 v1 pith:HTKJAXRJ submitted 2025-05-16 cs.LG

classification cs.LG
keywords learningpolicypurealgorithmalgorithmsbaselineenvironmentsloro
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We investigate the usage of Large Language Model (LLM) in collecting high-quality data to warm-start Reinforcement Learning (RL) algorithms for learning in some classical Markov Decision Process (MDP) environments. In this work, we focus on using LLM to generate an off-policy dataset that sufficiently covers state-actions visited by optimal policies, then later using an RL algorithm to explore the environment and improve the policy suggested by the LLM. Our algorithm, LORO, can both converge to an optimal policy and have a high sample efficiency thanks to the LLM's good starting policy. On multiple OpenAI Gym environments, such as CartPole and Pendulum, we empirically demonstrate that LORO outperforms baseline algorithms such as pure LLM-based policies, pure RL, and a naive combination of the two, achieving up to $4 \times$ the cumulative rewards of the pure RL baseline.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProDVI: Programmatic Dynamics Priors for Value Network Initialization

    cs.LG 2026-08 conditional novelty 6.0 of 10

    LLM-generated dynamics programs, used only to pretrain a value network's state-action encoder, improve sample efficiency of model-free RL on continuous control tasks.

Pith tools