Pith. sign in

REVIEW 6 cited by

Do LLMs exhibit human-like response biases? A case study in survey design

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.04076 v5 pith:IFAOVHHN submitted 2023-11-07 cs.CL

classification cs.CL
keywords llmsbiaseshumansresponsechangesdesignhumanhuman-like
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As large language models (LLMs) become more capable, there is growing excitement about the possibility of using LLMs as proxies for humans in real-world tasks where subjective labels are desired, such as in surveys and opinion polling. One widely-cited barrier to the adoption of LLMs as proxies for humans in subjective tasks is their sensitivity to prompt wording - but interestingly, humans also display sensitivities to instruction changes in the form of response biases. We investigate the extent to which LLMs reflect human response biases, if at all. We look to survey design, where human response biases caused by changes in the wordings of "prompts" have been extensively explored in social psychology literature. Drawing from these works, we design a dataset and framework to evaluate whether LLMs exhibit human-like response biases in survey questionnaires. Our comprehensive evaluation of nine models shows that popular open and commercial LLMs generally fail to reflect human-like behavior, particularly in models that have undergone RLHF. Furthermore, even if a model shows a significant change in the same direction as humans, we find that they are sensitive to perturbations that do not elicit significant changes in humans. These results highlight the pitfalls of using LLMs as human proxies, and underscore the need for finer-grained characterizations of model behavior. Our code, dataset, and collected samples are available at https://github.com/lindiatjuatja/BiasMonkey

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

    cs.CL 2026-07 conditional novelty 8.0 of 10

    Across 45 LLMs, the 'right?' tag effect flips from sycophantic to resistant over four years of releases, while the 'maybe?' tag raises agreement in every model — anti-sycophancy training is grammar-keyed and one-sided.

  2. People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Across 10 LLMs and the European Social Survey, AI alignment favors wealthier, more educated, less religious, and more politically interested groups, with country of residence explaining as much variance as all sociode...

  3. RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A detector that projects a text's hidden neural activations onto a direction learned from human versus AI writing differences reports higher accuracy than baselines, including on unseen generators and attacked texts.

  4. Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs

    cs.AI 2025-05 conditional novelty 6.0 of 10

    LLM confidence is less sensitive to task difficulty than human confidence and bends to persona stereotypes, and separating confidence prompts from answer prompts (AFCE) improves calibration on hard tasks.

  5. Exploring LLMs for Automated Generation and Adaptation of Questionnaires

    cs.HC 2025-01 conditional novelty 5.0 of 10

    LLM-generated survey questions were rated as clear and specific, while LLM-based pretesting improved some adapted questions but often made original questions wordier and less clear.

  6. Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods

    cs.AI 2025-05 accept novelty 4.0 of 10

    LLMs extend, rather than replace, classical social science methods, with a proposed three-tier bias framework for LLM-augmented surveys.

Pith tools