Pith. sign in

REVIEW 1 cited by

Rethinking Data Selection for Supervised Fine-Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06094 v1 pith:MUPEZDW7 submitted 2024-02-08 cs.CL

classification cs.CL
keywords dataselectiondiversityessentialfinetuninginstancesqualityresponses
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although supervised finetuning (SFT) has emerged as an essential technique to align large language models with humans, it is considered superficial, with style learning being its nature. At the same time, recent works indicate the importance of data selection for SFT, showing that finetuning with high-quality and diverse subsets of the original dataset leads to superior downstream performance. In this work, we rethink the intuition behind data selection for SFT. Considering SFT is superficial, we propose that essential demonstrations for SFT should focus on reflecting human-like interactions instead of data quality or diversity. However, it is not straightforward to directly assess to what extent a demonstration reflects human styles. Towards an initial attempt in this direction, we find selecting instances with long responses is surprisingly more effective for SFT than utilizing full datasets or instances selected based on quality and diversity. We hypothesize that such a simple heuristic implicitly mimics a crucial aspect of human-style conversation: detailed responses are usually more helpful.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Boosting LLM via Learning from Data Iteratively and Selectively

    cs.CL 2024-12 conditional novelty 6.0 of 10

    IterIT iteratively re-scores instruction samples during fine-tuning and greedily selects a small, diverse, high-complexity subset each epoch.

Pith tools