Pith. sign in

REVIEW 2 cited by

Reformatted Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12219 v2 pith:ZMNMWVZS submitted 2024-02-19 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords dataalignmentrealignabilityhumanllmsqualityapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The quality of finetuning data is crucial for aligning large language models (LLMs) with human values. Current methods to improve data quality are either labor-intensive or prone to factual errors caused by LLM hallucinations. This paper explores elevating the quality of existing instruction data to better align with human values, introducing a simple and effective approach named ReAlign, which reformats the responses of instruction data into a format that better aligns with pre-established criteria and the collated evidence. This approach minimizes human annotation, hallucination, and the difficulty in scaling, remaining orthogonal to existing alignment techniques. Experimentally, ReAlign significantly boosts the general alignment ability, math reasoning, factuality, and readability of the LLMs. Encouragingly, without introducing any additional data or advanced training techniques, and merely by reformatting the response, LLaMA-2-13B's mathematical reasoning ability on GSM8K can be improved from 46.77% to 56.63% in accuracy. Additionally, a mere 5% of ReAlign data yields a 67% boost in general alignment ability measured by the Alpaca dataset. This work highlights the need for further research into the science and mechanistic interpretability of LLMs. We have made the associated code and data publicly accessible to support future studies at https://github.com/GAIR-NLP/ReAlign.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Toward Cybersecurity-Expert Small Language Models

    cs.CL 2025-10 conditional novelty 5.0 of 10

    A family of 4B–20B cybersecurity models fine-tuned on an enriched, expert-steered reasoning dataset matches or beats larger frontier models on core CTI benchmarks.

  2. Alignment at Pre-training! Towards Native Alignment for Arabic LLMs

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Rewriting Arabic pre-training data with aligned LLM workers improves safety, helpfulness, and Arabic benchmark scores in the released LLaMA3-Tamed models.

Pith tools