REVIEW 15 cited by
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) such as ChatGPT have seen widespread adoption due to their strong instruction-following abilities. Developing these LLMs involves a complex yet poorly understood workflow requiring training with human feedback. Replicating and understanding this instruction-following requires tackling three major challenges: the high cost of data collection, the lack of trustworthy evaluation, and the absence of reference method implementations. We address these challenges with AlpacaFarm, a simulator that enables research and development for learning from feedback at a low cost. First, we design LLM prompts to simulate human feedback that are 50x cheaper than crowdworkers and display high agreement with humans. Second, we propose an automatic evaluation and validate it against human instructions obtained on real-world interactions. Third, we contribute reference implementations for several methods (PPO, DPO, best-of-n, expert iteration, and more) that learn from pairwise feedback. Finally, as an end-to-end validation of AlpacaFarm, we train and evaluate eleven models on 10k pairs of real human feedback and show that rankings of models trained in AlpacaFarm match rankings of models trained on human data. As a demonstration of the research possible in AlpacaFarm, we find that methods that use a reward model can substantially improve over supervised fine-tuning and that our reference PPO implementation leads to a +10% improvement in win-rate against Davinci003. We release all components of AlpacaFarm at https://github.com/tatsu-lab/alpaca_farm.
Forward citations
Cited by 15 Pith papers
-
UQ: Assessing Language Models on Unsolved Questions
Unsolved Stack Exchange questions can serve as a dynamic benchmark: the best tested model passes validator screening on only 15% of 500 questions.
-
ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling
Per-step retrieval of solved exemplars injected into the reasoning trace improves test-time scaling accuracy, with up to 13.4 absolute points gained on AIME 2025.
-
Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration
CALM uses bilevel optimization to tune per-vocabulary temperature-like logit adjustments during LLM fine-tuning, and reports improved out-of-domain calibration for aligned language models.
-
Multi-Turn On-Policy Distillation with Prefix Replay
ReOPD offline-distills multi-turn agentic LLMs via teacher-prefix replay plus step-decay sampling, matching online OPD accuracy at ≥4× speed with zero tool calls.
-
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
HC-RLHF returns an aligned language model only after a held-out safety test certifies, with probability at least 1-delta, that expected harm (as judged by a learned cost model) is below a chosen threshold.
-
SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
SGDPO modifies DPO with a subsequence-based pilot term and reports up to 9.19% relative MT-Bench gains, though the claimed gradient mechanism is not fully derived for the implemented loss.
-
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
Increasing energy loss in an LLM's final layer during RLHF is linked to reward hacking, and penalizing that loss (EPPO) reduces hacking and improves RLHF quality.
-
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
AutoConverter converts open-ended VQA questions into multiple-choice format via multi-agent GPT-4o, and VMCBench applies it to 20 datasets to evaluate 33 vision-language models.
-
A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls
A rubric-plus-question-answering LLM framework for literary translation evaluation beats traditional MT metrics but still trails human agreement, especially on Korean honorifics.
-
RecoWorld: Building Simulated Environments for Agentic Recommender Systems
A design proposal, not a tested system: a dual-view simulation loop in which an LLM-simulated user issues reflective instructions when about to disengage, and an instruction-following recommender adapts to maximize si...
-
An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models
A training pipeline that scores a model's own responses for semantic, factual, and safety uncertainty, builds preference pairs from those scores, and trains in three difficulty stages improves reported alignment score...
-
Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
Prompt injection can make LLM agents leak personal data they observed while executing tasks, with measured attack success rates around 15-20 percent and password leakage much rarer.
-
DeepCRCEval: Revisiting the Evaluation of Code Review Comment Generation
An empirical study that rebuilds code review comment evaluation around a nine-criterion rubric and claims a training-free LLM baseline beats existing generators.
-
ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
ReSURE reduces the harm of noisy dialogue data during fine-tuning by grouping samples by dialogue depth and softly down-weighting high-loss examples, improving multi-turn benchmarks modestly.
-
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
Aya Expanse 8B and 32B report state-of-the-art multilingual win-rates on a new 23-language translated Arena-Hard benchmark, with the 32B beating Llama 3.1 70B by 54.0%.
Discussion (0). Continue with ORCID to comment.