Pith. sign in

REVIEW 4 cited by

Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11270 v2 pith:ESV6FARR submitted 2023-06-20 cs.CL cs.LG

classification cs.CLcs.LG
keywords instructioninstructionsmodelsperformanceinstruction-tunedlanguagephrasingsconsistently
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction fine-tuning has recently emerged as a promising approach for improving the zero-shot capabilities of Large Language Models (LLMs) on new tasks. This technique has shown particular strength in improving the performance of modestly sized LLMs, sometimes inducing performance competitive with much larger model variants. In this paper we ask two questions: (1) How sensitive are instruction-tuned models to the particular phrasings of instructions, and, (2) How can we make them more robust to such natural language variation? To answer the former, we collect a set of 319 instructions manually written by NLP practitioners for over 80 unique tasks included in widely used benchmarks, and we evaluate the variance and average performance of these instructions as compared to instruction phrasings observed during instruction fine-tuning. We find that using novel (unobserved) but appropriate instruction phrasings consistently degrades model performance, sometimes substantially so. Further, such natural instructions yield a wide variance in downstream performance, despite their semantic equivalence. Put another way, instruction-tuned models are not especially robust to instruction re-phrasings. We propose a simple method to mitigate this issue by introducing ``soft prompt'' embedding parameters and optimizing these to maximize the similarity between representations of semantically equivalent instructions. We show that this method consistently improves the robustness of instruction-tuned models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise

    cs.IR 2026-07 accept novelty 6.5 of 10

    The Power-of-Noise effect in RAG is reproducible only under the original constrained setup and disappears or weakens once instruction templates, longer outputs, and modern LLMs are used.

  2. Beyond Prompt Content: Enhancing LLM Performance via Content-Format Integrated Prompt Optimization

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Jointly optimizing prompt wording and formatting with UCT-guided format search and LLM-generated formats improves accuracy over content-only prompt optimizers on several benchmarks and models.

  3. PlotTwist: A Creative Plot Generation Framework with Small Language Models

    cs.CL 2026-03 reject novelty 5.0 of 10

    PlotTwist aligns a 3B-active-parameter model with DPO to generate movie plots that its own Qwen-3-32B agentic evaluator scores above GPT-4.1, Claude Sonnet 4, and Gemini 2.0 Flash.

  4. Investigating the Robustness of Retrieval-Augmented Generation at the Query Level

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Retrieval-augmented generation performance drops noticeably under minor query perturbations, with end-to-end results often tracking retriever behavior.

Pith tools