REVIEW 7 cited by
Large Language Models Show Human-like Social Desirability Biases in Survey Responses
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
As Large Language Models (LLMs) become widely used to model and simulate human behavior, understanding their biases becomes critical. We developed an experimental framework using Big Five personality surveys and uncovered a previously undetected social desirability bias in a wide range of LLMs. By systematically varying the number of questions LLMs were exposed to, we demonstrate their ability to infer when they are being evaluated. When personality evaluation is inferred, LLMs skew their scores towards the desirable ends of trait dimensions (i.e., increased extraversion, decreased neuroticism, etc). This bias exists in all tested models, including GPT-4/3.5, Claude 3, Llama 3, and PaLM-2. Bias levels appear to increase in more recent models, with GPT-4's survey responses changing by 1.20 (human) standard deviations and Llama 3's by 0.98 standard deviations-very large effects. This bias is robust to randomization of question order and paraphrasing. Reverse-coding all the questions decreases bias levels but does not eliminate them, suggesting that this effect cannot be attributed to acquiescence bias. Our findings reveal an emergent social desirability bias and suggest constraints on profiling LLMs with psychometric tests and on using LLMs as proxies for human participants.
Forward citations
Cited by 7 Pith papers
-
Low Stage and High Order Explicit Runge--Kutta Methods via $Q$- and $D$-Conditions: Several Construction Details
A Q/D-space reformulation of Butcher simplifying assumptions yields sufficient order conditions and a recursive linear-system construction for explicit Runge-Kutta methods of even order p with s(p)=(p²-2p+8)/4 stages.
-
Bullying the Machine: How Personas Increase LLM Vulnerability
Persona-conditioned LLMs are more likely to produce unsafe outputs when pressured with bullying tactics, particularly when agreeableness or conscientiousness is weakened.
-
Debunking with Dialogue? Exploring AI-Generated Counterspeech to Challenge Conspiracy Theories
Zero-shot LLM-generated counterspeech to conspiracy theories is shallow, repetitive, and about 10% confabulated, with no model reaching practical effectiveness.
-
Departures from Standard Disk Predictions in Intensive Ground-Based Monitoring of Three AGN
Based only on the abstract, the paper reports that Mrk 509 inter-band continuum lags scale as wavelength to the 2.17 power, steeper than the thin-disk prediction, but the appended full text is a different article.
-
From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology
This Perspective paper proposes that LLM research in psychology must combine psychometric validity and causal inference standards, mapping evidence requirements to the type of claim being made.
-
Be.FM: Open Foundation Models for Human Behavior
Be.FM fine-tunes Llama models on behavioral data and claims improved behavior prediction, but its headline evaluation is compromised by testing on the same data it trained on.
-
A Survey on Training-free Alignment of Large Language Models
A survey that catalogs and categorizes training-free LLM alignment methods into pre-decoding, in-decoding, and post-decoding, with a limited experimental comparison on one model.
Discussion (0). Continue with ORCID to comment.