Pith. sign in

REVIEW 5 cited by

ICPL: Few-shot In-context Preference Learning via LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.17233 v3 pith:JUZ3ULDX submitted 2024-10-22 cs.AI cs.LG

classification cs.AIcs.LG
keywords learningpreferenceicplcapabilitieshumanin-contextllmsrewards
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Preference-based reinforcement learning is an effective way to handle tasks where rewards are hard to specify but can be exceedingly inefficient as preference learning is often tabula rasa. We demonstrate that Large Language Models (LLMs) have native preference-learning capabilities that allow them to achieve sample-efficient preference learning, addressing this challenge. We propose In-Context Preference Learning (ICPL), which uses in-context learning capabilities of LLMs to reduce human query inefficiency. ICPL uses the task description and basic environment code to create sets of reward functions which are iteratively refined by placing human feedback over videos of the resultant policies into the context of an LLM and then requesting better rewards. We first demonstrate ICPL's effectiveness through a synthetic preference study, providing quantitative evidence that it significantly outperforms baseline preference-based methods with much higher performance and orders of magnitude greater efficiency. We observe that these improvements are not solely coming from LLM grounding in the task but that the quality of the rewards improves over time, indicating preference learning capabilities. Additionally, we perform a series of real human preference-learning trials and observe that ICPL extends beyond synthetic settings and can work effectively with humans-in-the-loop.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices

    cs.AI 2026-02 conditional novelty 6.0 of 10

    LLM-derived willingness-to-pay for hotel attributes deviates systematically from human benchmarks; cheap-preference examples pull models closer, while expensive or business-persona prompts push them further away.

  2. Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A hybrid method uses LLM-generated candidate gaits and then refines them with a few human preference rankings, achieving quadruped behaviors aligned with user intent in as few as four queries.

  3. Instant Preference Alignment for Text-to-Image Diffusion Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    An MLLM-driven, training-free pipeline extracts preference keywords from a reference image and modulates diffusion cross-attention at global and regional levels for instant, multi-round preference-aligned image generation.

  4. Activation Reward Models for Few-Shot Model Alignment

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Mean attention-head activations from a few labeled examples, injected into selected heads, turn a frozen vision-language model into a few-shot reward model that beats prompting and scoring baselines and a new reward-h...

  5. Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation

    cs.IR 2025-11 conditional novelty 4.0 of 10

    A review and position paper proposing a six-dimension success framework and risk diagnostics for evaluating LLM-based music recommendation systems.

Pith tools