REVIEW 3 cited by
The language of prompting: What linguistic properties make a prompt successful?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The latest generation of LLMs can be prompted to achieve impressive zero-shot or few-shot performance in many NLP tasks. However, since performance is highly sensitive to the choice of prompts, considerable effort has been devoted to crowd-sourcing prompts or designing methods for prompt optimisation. Yet, we still lack a systematic understanding of how linguistic properties of prompts correlate with task performance. In this work, we investigate how LLMs of different sizes, pre-trained and instruction-tuned, perform on prompts that are semantically equivalent, but vary in linguistic structure. We investigate both grammatical properties such as mood, tense, aspect and modality, as well as lexico-semantic variation through the use of synonyms. Our findings contradict the common assumption that LLMs achieve optimal performance on lower perplexity prompts that reflect language use in pretraining or instruction-tuning data. Prompts transfer poorly between datasets or models, and performance cannot generally be explained by perplexity, word frequency, ambiguity or prompt length. Based on our results, we put forward a proposal for a more robust and comprehensive evaluation standard for prompting research.
Forward citations
Cited by 3 Pith papers
-
A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
Relevance labels produced by LLMs depend on the prompt and the judge model; LLM-written prompts are more consistent than expert-written prompts across three assessment paradigms, and graded 0-to-3 judgments are unstab...
-
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
Multimodal language models trained with a medium number of instruction templates (5,000 for 7B, 100 for 13B) outperform both fewer and many more templates, with gains up to 10 points on small benchmark samples.
-
Trusting CHATGPT: how minor tweaks in the prompts lead to major differences in sentiment classification
Minor prompt rewording produces statistically significant shifts in GPT-4o mini's Spanish sentiment labels, yet overall agreement between prompts stays between 92% and 98%.
Discussion (0). Continue with ORCID to comment.