Pith. sign in

REVIEW 5 cited by

A Critical Evaluation of AI Feedback for Aligning Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12366 v1 pith:7PUGVZHR submitted 2024-02-19 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords rlaiffeedbackmodelmodelscriticstepteacherevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reinforcement learning with AI feedback (RLAIF) is a popular paradigm for improving the instruction-following abilities of powerful pre-trained language models. RLAIF first performs supervised fine-tuning (SFT) using demonstrations from a teacher model and then further fine-tunes the model with reinforcement learning (RL), using feedback from a critic model. While recent popular open-source models have demonstrated substantial improvements in performance from the RL step, in this paper we question whether the complexity of this RL step is truly warranted for AI feedback. We show that the improvements of the RL step are virtually entirely due to the widespread practice of using a weaker teacher model (e.g. GPT-3.5) for SFT data collection than the critic (e.g., GPT-4) used for AI feedback generation. Specifically, we show that simple supervised fine-tuning with GPT-4 as the teacher outperforms existing RLAIF pipelines. More generally, we find that the gains from RLAIF vary substantially across base model families, test-time evaluation protocols, and critic models. Finally, we provide a mechanistic explanation for when SFT may outperform the full two-step RLAIF pipeline as well as suggestions for making RLAIF maximally useful in practice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Toward a Theory of Value in AI Alignment

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A systematic annotation of 94 AI alignment papers shows the field largely equates human values with measurable preferences, rarely defines values, and is increasingly removing humans from alignment evaluation.

  2. Graph-Enhanced Policy Optimization in LLM Agent Training

    cs.AI 2025-10 conditional novelty 6.0 of 10

    GEPO adds graph-centrality-based intrinsic rewards, dynamic discounts, and two-level advantage shaping to group-based RL, improving LLM agent success on ALFWorld, WebShop, and a private Workbench benchmark.

  3. Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Adding transcription and answer-selection auxiliary tasks to a fixed spoken-QA dataset improves MLLM performance with less data, validated on three datasets and a new ASK-QA benchmark.

  4. Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Adding a special refuse token to a fine-tuned LLM lets a developer tune refusal rates at inference time by thresholding the token's probability, with per-category control.

  5. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools