Pith. sign in

REVIEW 2 cited by

Instruction Tuned Models are Quick Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.05539 v1 pith:Q4W5BFER submitted 2023-05-17 cs.CL

classification cs.CL
keywords instructiondatalearningdownstreammodelstaskssotatraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction tuning of language models has demonstrated the ability to enhance model generalization to unseen tasks via in-context learning using a few examples. However, typical supervised learning still requires a plethora of downstream training data for finetuning. Often in real-world situations, there is a scarcity of data available for finetuning, falling somewhere between few shot inference and fully supervised finetuning. In this work, we demonstrate the sample efficiency of instruction tuned models over various tasks by estimating the minimal downstream training data required by them to perform transfer learning and match the performance of state-of-the-art (SOTA) supervised models. We conduct experiments on 119 tasks from Super Natural Instructions (SuperNI) in both the single task learning (STL) and multi task learning (MTL) settings. Our findings reveal that, in the STL setting, instruction tuned models equipped with 25% of the downstream train data surpass the SOTA performance on the downstream tasks. In the MTL setting, an instruction tuned model trained on only 6% of downstream training data achieve SOTA, while using 100% of the training data results in a 3.69% points improvement (ROUGE-L 74.68) over the previous SOTA. We conduct an analysis on T5 vs Tk-Instruct by developing several baselines to demonstrate that instruction tuning aids in increasing both sample efficiency and transfer learning. Additionally, we observe a consistent ~4% performance increase in both settings when pre-finetuning is performed with instructions. Finally, we conduct a categorical study and find that contrary to previous results, tasks in the question rewriting and title generation categories suffer from instruction tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adapting Biomedical Abstracts into Plain language using Large Language Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A one-shot GPT-4 prompt driven by distilled PLABA annotation guidelines ranked first on simplicity and third on accuracy in the plain-language biomedical abstract adaptation task.

  2. Steps are all you need: Rethinking STEM Education with Prompt Engineering

    cs.CL 2024-12 reject novelty 4.0 of 10

    A new prompting recipe that mixes few-shot chain-of-thought examples with model-generated analogies is claimed to improve STEM question-answering on Mixtral 8x7B, alongside a new 928-question dataset.

Pith tools