REVIEW 2 cited by
Using Large Language Models to Automate and Expedite Reinforcement Learning with Reward Machine
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present LARL-RM (Large language model-generated Automaton for Reinforcement Learning with Reward Machine) algorithm in order to encode high-level knowledge into reinforcement learning using automaton to expedite the reinforcement learning. Our method uses Large Language Models (LLM) to obtain high-level domain-specific knowledge using prompt engineering instead of providing the reinforcement learning algorithm directly with the high-level knowledge which requires an expert to encode the automaton. We use chain-of-thought and few-shot methods for prompt engineering and demonstrate that our method works using these approaches. Additionally, LARL-RM allows for fully closed-loop reinforcement learning without the need for an expert to guide and supervise the learning since LARL-RM can use the LLM directly to generate the required high-level knowledge for the task at hand. We also show the theoretical guarantee of our algorithm to converge to an optimal policy. We demonstrate that LARL-RM speeds up the convergence by 30% by implementing our method in two case studies.
Forward citations
Cited by 2 Pith papers
-
A Domain Adaptation of Large Language Models for Classifying Mechanical Assembly Components
Fine-tuning GPT-3.5 Turbo on 681 OSDR parts yields about 89% accuracy on held-out OSDR data and produces function labels for ABC parts, although the ABC labels are never checked against ground truth.
-
Leveraging LLM for Automated Ontology Extraction and Knowledge Graph Generation
OntoKGen automates ontology extraction and knowledge graph generation from technical documents using LLMs with user-guided iterative prompting, demonstrated on a semiconductor equipment reliability case study.
Discussion (0). Continue with ORCID to comment.