REVIEW 2 cited by
Prompting is not a substitute for probability measurements in large language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Prompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs). While other methods directly read out models' probability distributions over strings, prompting requires models to access this internal information by processing linguistic input, thereby implicitly testing a new type of emergent ability: metalinguistic judgment. In this study, we compare metalinguistic prompting and direct probability measurements as ways of measuring models' linguistic knowledge. Broadly, we find that LLMs' metalinguistic judgments are inferior to quantities directly derived from representations. Furthermore, consistency gets worse as the prompt query diverges from direct measurements of next-word probabilities. Our findings suggest that negative results relying on metalinguistic prompts cannot be taken as conclusive evidence that an LLM lacks a particular linguistic generalization. Our results also highlight the value that is lost with the move to closed APIs where access to probability distributions is limited.
Forward citations
Cited by 2 Pith papers
-
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
Across heavy NP shift, dative alternation and multiple PP shift, LLM ordering preferences correlate with human judgments, but particle movement preferences do not.
-
Human Psychometric Questionnaires Mischaracterize LLM Behavior
Standard psychometric questionnaires like the Big Five and PVQ produce different and more consistent results than ecologically valid questions drawn from real user conversations, suggesting the former may mischaracter...
Discussion (0). Continue with ORCID to comment.