REVIEW 3 cited by
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs). However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective human preferences and values. In this paper, we survey existing approaches for learning from human feedback, drawing on 95 papers primarily from the ACL and arXiv repositories.First, we summarise the past, pre-LLM trends for integrating human feedback into language models. Second, we give an overview of present techniques and practices, as well as the motivations for using feedback; conceptual frameworks for defining values and preferences; and how feedback is collected and from whom. Finally, we encourage a better future of feedback learning in LLMs by raising five unresolved conceptual and practical challenges.
Forward citations
Cited by 3 Pith papers
-
Pairwise Calibrated Rewards for Pluralistic Alignment
A small ensemble of reward functions can be trained to match pairwise human preference frequencies, offering a practical route to pluralistic AI alignment.
-
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
J1-7B, a judge LLM trained with supervised fine-tuning and reinforcement learning, improves when forced to reflect with 'wait' tokens, and the scaling ability emerges during the RL phase.
-
Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment
Persona-judge applies speculative decoding between two preference-prompted copies of the same LLM to achieve training-free personalized alignment.
Discussion (0). Continue with ORCID to comment.