Pith. sign in

REVIEW 6 cited by

Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.00402 v1 pith:JJDK2KZT submitted 2024-05-01 cs.CL

classification cs.CL
keywords modelslanguagereasoningabilitiesinstruction-tuningllmsself-refinedemonstrations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The alignments of reasoning abilities between smaller and larger Language Models are largely conducted via Supervised Fine-Tuning (SFT) using demonstrations generated from robust Large Language Models (LLMs). Although these approaches deliver more performant models, they do not show sufficiently strong generalization ability as the training only relies on the provided demonstrations. In this paper, we propose the Self-refine Instruction-tuning method that elicits Smaller Language Models to self-refine their abilities. Our approach is based on a two-stage process, where reasoning abilities are first transferred between LLMs and Small Language Models (SLMs) via Instruction-tuning on demonstrations provided by LLMs, and then the instructed models Self-refine their abilities through preference optimization strategies. In particular, the second phase operates refinement heuristics based on the Direct Preference Optimization algorithm, where the SLMs are elicited to deliver a series of reasoning paths by automatically sampling the generated responses and providing rewards using ground truths from the LLMs. Results obtained on commonsense and math reasoning tasks show that this approach significantly outperforms Instruction-tuning in both in-domain and out-domain scenarios, aligning the reasoning abilities of Smaller and Larger Language Models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aligning VLM Assistants with Personalized Situated Cognition

    cs.AI 2025-06 conditional novelty 6.0 of 10

    The authors present PCogAlignBench, an 18k-sample benchmark of visual scenes with role-based users, and PCogAlign, a framework using a cognition-aware reward model to produce responses aligned with each user's roles.

  2. ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ARB provides 1,356 Arabic multimodal questions with 5,119 human-reviewed reasoning steps and shows leading models score much higher on reasoning fluency than on correct answers.

  3. Future-KL Regularized GRPO: Process-Level Credit Assignment from $f$-Divergence Regularization

    cs.LG 2026-01 unverdicted novelty 5.0 of 10

    Abstract claims FRPO adds a future-KL correction to GRPO for better math reasoning; the provided manuscript body is a different PRL paper and does not support those claims.

  4. Self-Training Large Language Models with Confident Reasoning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    CORE-PO self-trains LLMs to prefer reasoning paths with high self-estimated confidence, improving answer and reasoning accuracy on several benchmarks.

  5. Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.

  6. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools