REVIEW 8 cited by
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large Language Models (LLMs) have achieved significant advancements, however, the common learning paradigm treats LLMs as passive information repositories, neglecting their potential for active learning and alignment. Some approaches train LLMs using their own generated synthetic data, exploring the possibility of active alignment. However, there is still a huge gap between these one-time alignment methods and the continuous automatic alignment of humans. In this paper, we introduce \textbf{I-SHEEP}, an \textbf{I}terative \textbf{S}elf-En\textbf{H}anc\textbf{E}m\textbf{E}nt \textbf{P}aradigm.This human-like paradigm enables LLMs to \textbf{continuously self-align from scratch with nothing}. Compared to the one-time alignment method Dromedary \cite{sun2023principledriven}, which refers to the first iteration in this paper, I-SHEEP can significantly enhance capacities on both Qwen and Llama models. I-SHEEP achieves a maximum relative improvement of 78.2\% in the Alpaca Eval, 24.0\% in the MT Bench, and an absolute increase of 8.88\% in the IFEval accuracy over subsequent iterations in Qwen-1.5 72B model. Additionally, I-SHEEP surpasses the base model in various standard benchmark generation tasks, achieving an average improvement of 24.77\% in code generation tasks, 12.04\% in TrivialQA, and 20.29\% in SQuAD. We also provide new insights based on the experiment results. Our codes, datasets, and models are available at \textbf{https://anonymous.4open.science/r/I-SHEEP}.
Forward citations
Cited by 8 Pith papers
-
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
SyncLoop jointly evolves multimodal training data and model capability through alternating SFT and RL, selecting error-prone samples to improve geometry reasoning.
-
LLMs for Customized Marketing Content Generation and Evaluation at Scale
MarketingFM generates e-commerce ad copy with RAG and an LLM; AutoEval uses LLM-as-a-Judge plus rule checks and self-refines its prompts, with online tests showing significant clicks and impressions lifts but no signi...
-
Aligning Instruction Tuning with Pre-training
AITP selects pre-training texts that are underrepresented relative to instruction-tuning data, rewrites them into instruction-response pairs, and improves average SFT performance on three open LLMs.
-
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
A controlled, multi-family study shows the relative generation-verification gap grows with pretraining flops for stable verification methods, and iterative self-improvement saturates quickly.
-
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
Preference Tree Optimization uses look-ahead simulations scored by an AI oracle to generate DPO preference data, and the resulting Motivational Interviewing agent scores higher on that same oracle than the base model.
-
Enterprise Large Language Model Evaluation Benchmark
A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...
-
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
Iterative self-training on a model's own correct outputs, with simple length and voting filters, lets transformers generalize to far longer arithmetic and path-finding problems than they saw in training.
-
Curiosity-Driven Reinforcement Learning from Human Feedback
Adding a prediction-error curiosity reward, masked to non-top-k tokens, improves output diversity of RLHF-trained LLMs while keeping reward-model judged quality roughly unchanged.
Discussion (0). Continue with ORCID to comment.