Pith. sign in

REVIEW 3 cited by

The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07629 v1 pith:P26LQLJG submitted 2023-10-11 cs.CL cs.CY

classification cs.CLcs.CY
keywords feedbackhumanlanguagelearningmodelspreferencesvaluesbetter
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs). However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective human preferences and values. In this paper, we survey existing approaches for learning from human feedback, drawing on 95 papers primarily from the ACL and arXiv repositories.First, we summarise the past, pre-LLM trends for integrating human feedback into language models. Second, we give an overview of present techniques and practices, as well as the motivations for using feedback; conceptual frameworks for defining values and preferences; and how feedback is collected and from whom. Finally, we encourage a better future of feedback learning in LLMs by raising five unresolved conceptual and practical challenges.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pairwise Calibrated Rewards for Pluralistic Alignment

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A small ensemble of reward functions can be trained to match pairwise human preference frequencies, offering a practical route to pluralistic AI alignment.

  2. J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge

    cs.LG 2025-05 conditional novelty 6.0 of 10

    J1-7B, a judge LLM trained with supervised fine-tuning and reinforcement learning, improves when forced to reflect with 'wait' tokens, and the scaling ability emerges during the RL phase.

  3. Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment

    cs.CL 2025-04 conditional novelty 6.0 of 10

    Persona-judge applies speculative decoding between two preference-prompted copies of the same LLM to achieve training-free personalized alignment.

Pith tools