Pith. sign in

REVIEW 1 cited by

Language Model-In-The-Loop: Data Optimal Approach to Learn-To-Recommend Actions in Text Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07687 v1 pith:QFANL3AJ submitted 2023-11-13 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords gamesllmstextannotatedduringhumanlanguagelearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have demonstrated superior performance in language understanding benchmarks. CALM, a popular approach, leverages linguistic priors of LLMs -- GPT-2 -- for action candidate recommendations to improve the performance in text games in Jericho without environment-provided actions. However, CALM adapts GPT-2 with annotated human gameplays and keeps the LLM fixed during the learning of the text based games. In this work, we explore and evaluate updating LLM used for candidate recommendation during the learning of the text based game as well to mitigate the reliance on the human annotated gameplays, which are costly to acquire. We observe that by updating the LLM during learning using carefully selected in-game transitions, we can reduce the dependency on using human annotated game plays for fine-tuning the LLMs. We conducted further analysis to study the transferability of the updated LLMs and observed that transferring in-game trained models to other games did not result in a consistent transfer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Solving the Content Gap in Roblox Game Recommendations: LLM-Based Profile Generation and Reranking

    cs.IR 2025-02 reject novelty 5.0 of 10

    LLM-generated game profiles from in-game text plus a personalized LLM reranker improve NDCG Engagement at rank 10 by 4.9% on average, with mixed and sometimes negative results at other cutoffs.

Pith tools