REVIEW 5 cited by
ICPL: Few-shot In-context Preference Learning via LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Preference-based reinforcement learning is an effective way to handle tasks where rewards are hard to specify but can be exceedingly inefficient as preference learning is often tabula rasa. We demonstrate that Large Language Models (LLMs) have native preference-learning capabilities that allow them to achieve sample-efficient preference learning, addressing this challenge. We propose In-Context Preference Learning (ICPL), which uses in-context learning capabilities of LLMs to reduce human query inefficiency. ICPL uses the task description and basic environment code to create sets of reward functions which are iteratively refined by placing human feedback over videos of the resultant policies into the context of an LLM and then requesting better rewards. We first demonstrate ICPL's effectiveness through a synthetic preference study, providing quantitative evidence that it significantly outperforms baseline preference-based methods with much higher performance and orders of magnitude greater efficiency. We observe that these improvements are not solely coming from LLM grounding in the task but that the quality of the rewards improves over time, indicating preference learning capabilities. Additionally, we perform a series of real human preference-learning trials and observe that ICPL extends beyond synthetic settings and can work effectively with humans-in-the-loop.
Forward citations
Cited by 5 Pith papers
-
Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices
LLM-derived willingness-to-pay for hotel attributes deviates systematically from human benchmarks; cheap-preference examples pull models closer, while expensive or business-persona prompts push them further away.
-
Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning
A hybrid method uses LLM-generated candidate gaits and then refines them with a few human preference rankings, achieving quadruped behaviors aligned with user intent in as few as four queries.
-
Instant Preference Alignment for Text-to-Image Diffusion Models
An MLLM-driven, training-free pipeline extracts preference keywords from a reference image and modulates diffusion cross-attention at global and regional levels for instant, multi-round preference-aligned image generation.
-
Activation Reward Models for Few-Shot Model Alignment
Mean attention-head activations from a few labeled examples, injected into selected heads, turn a frozen vision-language model into a few-shot reward model that beats prompting and scoring baselines and a new reward-h...
-
Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation
A review and position paper proposing a six-dimension success framework and risk diagnostics for evaluating LLM-based music recommendation systems.
Discussion (0). Continue with ORCID to comment.