Pith. sign in

REVIEW 3 major objections 5 minor 84 references

Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that large language models extend rather than replace surveys, experiments, and content analysis, and maps where they help and where they fail.

desk verdict A solid, useful methodological review of LLMs in social science methods; the three-tier bias framework is the main fresh contribution, while the 'interpretive pluralism' claim needs more empirical grounding than the paper provides. read the letter →

arxiv 2505.19402 v1 pith:BIDASA5C submitted 2025-05-26 cs.AI cs.CY

classification cs.AIcs.CY
keywords largelanguagemodelscontentanalysissurveyresearchexperimentaldesignsimulatedrespondentsbiastaxonomyLasswell's5Wscomputationalsocialscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Rather than treating large language models as replacements for classical methods, this paper argues that LLMs extend content analysis, surveys, and experiments by adding new affordances: automated and interpretation-sensitive text coding, simulated respondents and opinion trajectories, and personalized, interactive, and counterfactual experimental stimuli. The review synthesizes recent interdisciplinary evidence to show where LLMs work—structured, low-ambiguity annotation, aggregate opinion prediction, and main-effect replication—and where they break down, including contested normative texts, subgroup and interaction effects, and counterfactual reasoning. A sympathetic reader would care because the claim positions generative AI as an augmentation layer within existing research logics rather than a rupture, giving working social scientists a concrete map of when an LLM-backed result can be trusted and when human validation is still load-bearing.

What carries the argument

The organizing machinery is Lasswell's question, 'Who says what, in which channel, to whom, with what effect?', used as an analytic grid: each component maps to a class of LLM affordances, with message studies linked to interpretive variation, audience studies to trajectory simulation, and effect studies to counterfactual experimentation. A second load-bearing mechanism is the three-tier bias taxonomy—representational, procedural, and interactional—under which the paper groups seven sources of bias in LLM-backed survey simulation, from persona construction and training data to prompt wording, generation sampling, evaluation metrics, and compound unknown effects. A third is triangulation: LLM outputs are treated as one layer among human coding, classical classifiers, and experimental validation, with prompt design, model selection, and output format treated as experimental variables rather than background preprocessing choices.

What would settle it

Take a preregistered set of 50 published experiments spanning media and survey contexts, run the same persona-based prompts on a next-generation LLM using the 133-effect replication protocol, and compute replication rates separately for main and interaction effects. If interaction-effect replication rises to near main-effect levels (for example, above 60 percent), then the paper's empirically grounded boundary—that LLMs reliably model main effects but not conditional relationships—collapses. A second check would test whether the seven bias sources persist unchanged across models and tasks; if a new model removes whole bias categories, the taxonomy's exhaustiveness claim would fail.

Watch

Extended reading notes

Core claim

On the paper's own terms, LLMs do not alter the core logic of social science; they recalibrate which parts of an established method are automated, simulated, or generated. In content analysis, they shift the goal from intercoder agreement to dialectic intersubjectivity, meaning controlled divergence of interpretations across simulated perspectives rather than convergence on a single reading. In surveys, they allow researchers to build synthetic respondent pools and trace audience trajectories, provided fidelity claims are filtered through a three-tier bias taxonomy separating representational, procedural, and interactional bias. In experiments, they generate stimuli, replicate known effects, and run multi-agent simulations; the reviewed evidence suggests that main effects replicate at high rates while interaction effects do not, with one large replication study reporting 76% of main effects and only 27% of interaction effects reproduced. The paper asserts that classical designs remain the compass that validates, anchors, and interprets these new capabilities.

Load-bearing premise

The paper's advice assumes that its seven bias sources are stable and exhaustive across LLMs and tasks, and that current empirical patterns—such as LLMs replicating main effects but struggling with interaction effects—remain representative of future models.

Editorial extensions

If this is right

  • Researchers should treat LLM annotation output as one interpretive layer among several, cross-checking it against human coders or classical classifiers rather than accepting it as ground truth.
  • When using LLMs to simulate survey respondents, fidelity at the national average is not enough; subgroup variation and within-group diversity must be validated before claims about public opinion are made.
  • Experimental designs can now personalize stimuli, run interactive dialogues, and simulate counterfactual message conditions, but the reviewed evidence indicates that interaction effects and conditional relationships may not replicate reliably in silico.
  • Methods sections and preregistration should include prompt templates, model version, sampling parameters, and output format as core design choices, not as implementation details.
  • LLM-based simulations are best used for hypothesis exploration, piloting, and theory testing with human validation, not as standalone replacements for human samples.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 76%-versus-27% main/interaction effect gap holds for future models, LLM-simulated experiments could be expected to reproduce direct persuasive effects but silently underestimate conditional moderation, making them poor candidates for detecting interaction hypotheses without supplementing with human experiments.
  • The three-tier bias taxonomy could double as a reporting checklist for synthetic-data studies; a preregistration that states persona construction, model algorithm, training provenance, prompt wording, sampling strategy, evaluation metrics, and expected compound effects would make the paper's implicit methodological standard explicit.
  • The Lasswell framing implies a reflexive turn the authors approach but do not fully state: LLMs themselves are now speakers with messages, channels, audiences, and effects, so the same 5W logic applies to studying AI-mediated communication as a cultural form. Treating this as a target of research rather than only as a methodological tool is a testable extension.
  • A direct test of the paper's boundary pattern could compare next-generation models against the 133-effect replication protocol; if interaction-effect replication climbs toward main-effect rates, the field's cautionary stance should be recalibrated upward.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper argues that large language models (LLMs) should be integrated into, rather than used to replace, classical quantitative social science methods. It reviews applications in three domains—content analysis, survey research, and experimental studies—and then re-reads Lasswell's 5W framework to organize the claimed affordances: interpretive variation in message study, audience trajectory modeling, and counterfactual experimentation. The paper is an integrative literature review with several structured comparison tables (Tables 1–5), a proposed three-tier bias framework for LLM-based surveys (§3.4), and explicit recommendations for documenting prompts, triangulating with human coding, and developing standardized protocols.

Significance. The paper is a timely, readable, and unusually comprehensive synthesis of a fast-moving literature. Its tables (Tables 3–5) give the reader a concrete map of the current evidence base, and it consistently qualifies claims with limitations from the cited studies (e.g., interaction-effect failures, cultural bias, prompt sensitivity). If the central 'augment rather than replace' thesis holds, the paper provides a useful conceptual bridge between computational and traditional methodological traditions. However, the paper's most distinctive conceptual contributions—dialectic intersubjectivity as a route to interpretive pluralism, and the three-tier bias taxonomy—are not empirically validated in this manuscript; they are advanced as conclusions from the literature review, and the review's own cited evidence creates unresolved tensions that need to be addressed.

major comments (3)
  1. [§2 and §5.1] The 'interpretive variation' claim—that LLM personas can surface the pluralism of meaning by simulating multiple audience perspectives—is load-bearing for the paper's contribution to content analysis and for its extensions in §5.2–5.3 (audience trajectories, counterfactual messages). Yet the paper itself cites evidence of high prompt sensitivity (Zhao et al. [13], Lu et al. [12]) and of LLMs reflecting dominant cultural norms (Li et al. [14]). These are in tension: divergence induced by persona prompts may reflect arbitrary prompt perturbations or activation of internal stereotypes rather than the distribution of interpretations across real human subgroups. The paper should either provide or point to direct validation comparing persona-conditioned LLM interpretations against the interpretations of actual human interpretive communities (e.g., coding distributions by ideology, culture, or expertise), or explicitly reframe this as an untested hypothesis and adjust the strength of the claims in §5.1 accordingly.
  2. [§4.2 and Table 4] The replication evidence summarized in Table 4 is mixed in exactly the places where the paper draws a positive recommendation. Yeykelis et al. [62] report 27% replication of interaction effects, and Hewitt et al. [59] find accuracy declines for underrepresented populations; Chen et al. [63] report skewed prediction for gender-, ethnicity-, and norm-related interventions. Yet §4.2 concludes that LLM simulation is 'particularly useful for piloting hypotheses, exploring counterfactuals, and evaluating designs across diverse or underrepresented populations.' That final clause is directly contradicted by the cited findings. The paper should reconcile the recommendation with the evidence—either by narrowing the claimed usefulness, or by explaining why the piloting use case is robust despite the replication gaps.
  3. [§3.3–3.4 and Table 2] The seven bias sources and the three-tier framework (representational, procedural, interactional) are presented as a diagnostic structure for locating where bias enters LLM-based survey simulation. As stated, however, the framework is a taxonomy rather than a diagnostic: the paper provides no procedure for attributing an observed bias to a specific tier, and §3.3.7 itself concedes that interaction effects among sources make isolation hard. That is a reasonable caveat, but the paper should state more explicitly that the framework's purpose is conceptual organization, not measurement, and should avoid implying in Table 2 or §3.4 that it enables researchers to 'analyze where bias enters the research pipeline' in an operational sense.
minor comments (5)
  1. [Abstract] There are several typographical artifacts from LaTeX ligatures, e.g., 'tra nsforming' in the abstract and 'reflect' throughout; these should be cleaned in production.
  2. [§3 opening] 'Since 1940s' should read 'Since the 1940s' for grammatical correctness.
  3. [References] Several references are incomplete or lack venue information: [24], [26], [59], and [64] list arXiv IDs or preprints without full publication status, while [61] is an early arXiv version whose final presentation could be updated.
  4. [Table 4] The row for Yeykelis et al. [62] lists '19,000+ personas' in the text but the table does not show the persona count; adding a column or matching the text would improve readability.
  5. [§4.1.5] The phrase 'marks the emergence of a new experimental paradigm' is stronger than the heterogeneous evidence in Table 3 justifies; a more hedged formulation (e.g., 'signals a shift toward') would better match the paper's otherwise careful tone.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a literature-grounded review whose self-citations are empirical and non-load-bearing.

full rationale

The manuscript is an integrative review/position paper: it makes no formal derivations and fits no parameters, so the main circularity mechanisms (self-definitional equations, fitted inputs renamed predictions) do not apply. Load-bearing empirical assertions are attributed to external studies (e.g., [3], [4], [5], [6], [62], [63]) or to the authors' own empirical preprints ([7], [35], [38], [22]). These self-citations are used as evidence, not as axioms: [7] reports an LLM persona-simulation study of interpretive variation, [35] reports a 43-model political-bias benchmark, and [38] reports shortcut-feature experiments in survey prediction. None of these citations is asserted as a theorem that forbids alternatives, and the paper repeatedly hedges its central 'interpretive pluralism' claim by acknowledging prompt sensitivity, cultural bias, and subgroup fidelity problems (Sections 3.3 and 4.2). The 'seven bias sources' and 'three-tier framework' are taxonomic syntheses of cited work, not results derived from their own definitions. The skeptic's concern about persona-conditioned divergence lacking validation against real human subgroups is a validity/robustness critique, not a circularity reduction: the paper does not define 'pluralism' as whatever the prompt produces, and it explicitly calls for careful validation. Consequently no step reduces by construction to its input, and no claim is forced by a self-citation chain.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

The paper's conceptual contributions rest on assumptions about the stability and exhaustiveness of the bias taxonomy, the validity of human ground truth, and the normative priority of classical research logics. No free parameters or invented physical entities are involved.

assumptions (4)
  • domain assumption LLMs behave consistently enough across tasks and versions that observed annotation and simulation performance generalizes.
    The paper's recommendations presume the cited performance patterns are stable; see Sections 2.2 and 3.3.
  • ad hoc to paper The seven bias sources in Section 3.3 are exhaustive and distinct.
    The three-tier framework groups these seven, but the paper offers no proof that other bias sources do not exist.
  • domain assumption Classical research logics (validity, reliability, design control) remain the appropriate normative standard for LLM-augmented research.
    The entire 'compass' argument depends on this normative premise, stated in Sections 1 and 6.
  • domain assumption Human annotation and survey responses are a valid ground truth for evaluating LLMs.
    Used in Sections 2.1 and 3.3 to judge LLM performance.
invented entities (2)
  • Three-tier bias framework (representational, procedural, interactional)
    purpose: A taxonomy for diagnosing where bias enters LLM-augmented survey research.
    This is a conceptual organizer proposed in Section 3.4; it has no falsifiable handle beyond the paper's own structure.
  • Dialectic intersubjectivity
    purpose: A stance for content analysis that treats divergent LLM interpretations as meaningful variation rather than error.
    Introduced via the authors' prior work [7], Section 2; it is a methodological value, not an empirically testable entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods." pith.science (2026). https://pith.science/paper/BIDASA5C

@misc{pith2026250519402,
  author       = {Pith},
  title        = {Pith review of: Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BIDASA5C}},
  note         = {Machine review of arXiv:2505.19402}
}
read the original abstract

This paper examines how large language models (LLMs) are transforming core quantitative methods in communication research in particular, and in the social sciences more broadly-namely, content analysis, survey research, and experimental studies. Rather than replacing classical approaches, LLMs introduce new possibilities for coding and interpreting text, simulating dynamic respondents, and generating personalized and interactive stimuli. Drawing on recent interdisciplinary work, the paper highlights both the potential and limitations of LLMs as research tools, including issues of validity, bias, and interpretability. To situate these developments theoretically, the paper revisits Lasswell's foundational framework -- "Who says what, in which channel, to whom, with what effect?" -- and demonstrates how LLMs reconfigure message studies, audience analysis, and effects research by enabling interpretive variation, audience trajectory modeling, and counterfactual experimentation. Revisiting the metaphor of the methodological compass, the paper argues that classical research logics remain essential as the field integrates LLMs and generative AI. By treating LLMs not only as technical instruments but also as epistemic and cultural tools, the paper calls for thoughtful, rigorous, and imaginative use of LLMs in future communication and social science research.

Figures

Figures reproduced from arXiv: 2505.19402 by the authors.

Figure 1
Figure 1. Co-evolution of Media Technologies and Research M [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

84 extracted references · 51 canonical work pages

  1. [7]

    Embracing Dialectic Intersubjectivity: Coordination of D ifferent Perspectives in Content Analysis with LLM Persona Simulation, February 2025

    Taewoo Kang, Kjerstin Thorson, Tai-Quan Peng, Dan Hiaes hutter-Rice, Sanguk Lee, and Stuart Soroka. Embracing Dialectic Intersubjectivity: Coordination of D ifferent Perspectives in Content Analysis with LLM Persona Simulation, February 2025. arXiv:2502.00903 [ cs]. 19 Recalibrating the Compass

  2. [35]

    Beyond Partisan Leaning: A Comparative Analysis of Political Bias in Large Language Models, May 2025

    Tai-Quan Peng, Kaiqi Yang, Sanguk Lee, Hang Li, Yucheng Chu, Yuping Lin, and Hui Liu. Beyond Partisan Leaning: A Comparative Analysis of Political Bias in Large Language Models, May 2025. arXiv:2412.16746 [cs] version: 4

  3. [38]

    Kaiqi Yang, Hang Li, Hongzhi Wen, Tai-Quan Peng, Jilian g Tang, and Hui Liu. Are Large Language Models (LLMs) Good Social Predictors? In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics : EMNLP 2024 , pages 2718–2730, Miami, Florida, USA, November 2024. Association for Comput ational Linguistics

  4. [13]

    Zhao, Eric Wallace, Shi Feng, Dan Klein, and Same er Singh

    Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Same er Singh. Calibrate Before Use: Improving Few-Shot Performance of Language Models, June 2021. arXiv: 2102.09690 [cs]

  5. [12]

    Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt O rder Sensitivity, March 2022

    Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt O rder Sensitivity, March 2022. arXiv:2104.08786 [cs]

  6. [14]

    CultureLLM: Incorporat- ing Cultural Differences into Large Language Models, Decemb er 2024

    Cheng Li, Mengzhou Chen, Jindong Wang, Sunayana Sitara m, and Xing Xie. CultureLLM: Incorporat- ing Cultural Differences into Large Language Models, Decemb er 2024. arXiv:2402.10946 [cs]

  7. [62]

    Cummings, and Byr on Reeves

    Leo Yeykelis, Kaavya Pichai, James J. Cummings, and Byr on Reeves. Using Large Language Models to Create AI Personas for Replication, Generalization and P rediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings, Apri l 2025. arXiv:2408.16073 [cs]

  8. [59]

    Predicting Results of Social Science ExperimentsUsing Large Language Models

    Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Ro bb Willer. Predicting Results of Social Science ExperimentsUsing Large Language Models. August 2024

  9. [63]

    Predicting Field Experiments with Large Language Models

    Yaoyu Chen, Yuheng Hu, and Yingda Lu. Predicting Field E xperiments with Large Language Models, April 2025. arXiv:2504.01167 [cs] version: 1

Show all 84 references
  1. [1]

    Lazer, A

    D. Lazer, A. Pentland, L. Adamic, S. Aral, A. L. Barabasi, D. Brewer, N. Christakis, N. Contractor, J. Fowler, M. Gutmann, T. Jebara, G. King, M. Macy, D. Roy, and M. Van Alstyne. Computational Social Science. Science, 323(5915):721–723, February 2009

  2. [2]

    When Communicat ion Meets Computation: Opportuni- ties, Challenges, and Pitfalls in Computational Communica tion Science

    Wouter van Atteveldt and Tai-Quan Peng. When Communicat ion Meets Computation: Opportuni- ties, Challenges, and Pitfalls in Computational Communica tion Science. Communication Methods and Measures, 12(2-3):81–92, 2018

  3. [3]

    Chat GPT outperforms crowd workers for text- annotation tasks

    Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. Chat GPT outperforms crowd workers for text- annotation tasks. Proceedings of the National Academy of Sciences , 120(30):e2305016120, July 2023. Publisher: Proceedings of the National Academy of Sciences

  4. [4]

    “HOT” ChatGPT: The Promise of ChatGPT in Detecting and Discriminating Hateful, Offensive , and Toxic Comments on Social Media

    Lingyao Li, Lizhou Fan, Shubham Atreja, and Libby Hemphi ll. “HOT” ChatGPT: The Promise of ChatGPT in Detecting and Discriminating Hateful, Offensive , and Toxic Comments on Social Media. ACM Trans. Web, 18(2):30:1–30:36, March 2024

  5. [5]

    Do AIs Know What the Most Important Issue is? Using Language M odels to Code Open-Text Social Survey Responses At Scale, August 2023

    Jonathan Mellon, Jack Bailey, Ralph Scott, James Breckw oldt, Marta Miori, and Phillip Schmedeman. Do AIs Know What the Most Important Issue is? Using Language M odels to Code Open-Text Social Survey Responses At Scale, August 2023

  6. [6]

    DeVerna, Harry Yaojun Yan, Kai-Cheng Yang, an d Filippo Menczer

    Matthew R. DeVerna, Harry Yaojun Yan, Kai-Cheng Yang, an d Filippo Menczer. Fact-checking in- formation from large language models can decrease headline discernment. Proceedings of the National Academy of Sciences , 121(50):e2322823121, December 2024. Publisher: Proceed ings o...

  7. [8]

    Goodby e human annotators? Content analysis of social policy debates using ChatGPT

    Erwin Gielens, Jakub Sowula, and Philip Leifeld. Goodby e human annotators? Content analysis of social policy debates using ChatGPT. Journal of Social Policy , pages 1–20, January 2025

  8. [9]

    The Altern ative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators wi th LLMs, January 2025

    Nitay Calderon, Roi Reichart, and Rotem Dror. The Altern ative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators wi th LLMs, January 2025. arXiv:2501.10970 [cs]

  9. [10]

    Wu, William Thorne, Ambrose Robinson, Ni kolaos Aletras, Carolina Scarton, Kalina Bontcheva, and Xingyi Song

    Yida Mu, Ben P. Wu, William Thorne, Ambrose Robinson, Ni kolaos Aletras, Carolina Scarton, Kalina Bontcheva, and Xingyi Song. Navigating Prompt Complexity f or Zero-Shot Classification: A Study of Large Language Models in Computational Social Science, Sep tember 2023. arXiv:230...

  10. [11]

    Large Language Models Cannot Self-Correct R easoning Yet, October 2023

    Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. Large Language Models Cannot Self-Correct R easoning Yet, October 2023. arXiv:2310.01798 [cs]

  11. [15]

    Robertson, Steve Rathje, Jochen Hartm ann, Saif M

    Stefan Feuerriegel, Abdurahman Maarouf, Dominik Bär, Dominique Geissler, Jonas Schweisthal, Nicolas Pröllochs, Claire E. Robertson, Steve Rathje, Jochen Hartm ann, Saif M. Mohammad, Oded Netzer, Alexandra A. Siegel, Barbara Plank, and Jay J. Van Bavel. Usi ng natural language ...

  12. [16]

    Human-in-the-loop or AI-in-the-loop? Automate or Collabo rate?, December 2024

    Sriraam Natarajan, Saurabh Mathur, Sahil Sidheekh, Wo lfgang Stammer, and Kristian Kersting. Human-in-the-loop or AI-in-the-loop? Automate or Collabo rate?, December 2024. arXiv:2412.14232 [cs] version: 1

  13. [17]

    Argyle, Ethan C

    Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gu bler, Christopher Rytting, and David Wingate. Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3):337–351, July 2023. Publisher: Cambridge Universit y Press

  14. [18]

    Clinton, Cassy Dorff, Brenton Ke nkel, and Jennifer M

    James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Ke nkel, and Jennifer M. Larson. Synthetic Replacements for Human Survey Data? The Perils of Large Lang uage Models. Political Analysis, pages 1–16, May 2024

  15. [19]

    AI-Augmented Surveys: Lev eraging Large Language Models for Opinion Prediction in Nationally Representative Surveys, May 2023

    Junsol Kim and Byungkyu Lee. AI-Augmented Surveys: Lev eraging Large Language Models for Opinion Prediction in Nationally Representative Surveys, May 2023 . arXiv:2305.09620 [cs]

  16. [20]

    Roy Burstein, Eric Mafuta, and Joshua L. Proctor. Large language models for analyzing open text in global health surveys: why children are not accessing vacci ne services in the Democratic Republic of the Congo. International Health , page ihaf015, March 2025

  17. [21]

    Language Models Trained on Media Diets Can Predict Public Opinion, March 2023

    Eric Chu, Jacob Andreas, Stephen Ansolabehere, and Deb Roy. Language Models Trained on Media Diets Can Predict Public Opinion, March 2023. arXiv:2303.1 6779 [cs]

  18. [22]

    Goldberg, Seth A

    Sanguk Lee, Tai-Quan Peng, Matthew H. Goldberg, Seth A. Rosenthal, John E. Kotcher, Edward W. Maibach, and Anthony Leiserowitz. Can large language model s estimate public opinion about global warming? An empirical assessment of algorithmic fidelity an d bias. PLOS Climate , 3(8...

  19. [23]

    Simulating Climate Change Discussion with Large Language Models: Cons iderations for Science Communication at Scale

    Ha Nguyen, Victoria Nguyen, Saríah López-Fierro, Sara Ludovise, and Rossella Santagata. Simulating Climate Change Discussion with Large Language Models: Cons iderations for Science Communication at Scale. In Proceedings of the Eleventh ACM Conference on Learning @ Sca le, L@S ...

  20. [24]

    Donald Trumps in the Virtual Polls: Simulating and Predicting Public Opinions in Surveys Using Large Language Models, February 2025

    Shapeng Jiang, Lijia Wei, and Chen Zhang. Donald Trumps in the Virtual Polls: Simulating and Predicting Public Opinions in Surveys Using Large Language Models, February 2025. arXiv:2411.01582 [econ]

  21. [25]

    Synthesizing Public Opinions with LLMs: Role Creation, Imp acts, and the Future to eDemorcacy, March 2025

    Rabimba Karanjai, Boris Shor, Amanda Austin, Ryan Kenn edy, Yang Lu, Lei Xu, and Weidong Shi. Synthesizing Public Opinions with LLMs: Role Creation, Imp acts, and the Future to eDemorcacy, March 2025. arXiv:2504.00241 [cs]. 20 Recalibrating the Compass

  22. [26]

    Can Large Language Mo dels Accurately Predict Public Opinion? A Review

    Arnault Pachot and Thierry Petit. Can Large Language Mo dels Accurately Predict Public Opinion? A Review. September 2024

  23. [27]

    Performance and biases of Large Lang uage Models in public opinion simulation

    Yao Qu and Jue Wang. Performance and biases of Large Lang uage Models in public opinion simulation. Humanities and Social Sciences Communications , 11(1):1–13, August 2024. Publisher: Palgrave

  24. [28]

    Large language models display human-like soci al desirability biases in Big Five personality surveys

    Aadesh Salecha, Molly E Ireland, Shashanka Subrahmany a, João Sedoc, Lyle H Ungar, and Johannes C Eichstaedt. Large language models display human-like soci al desirability biases in Big Five personality surveys. PNAS Nexus , 3(12):pgae533, December 2024

  25. [29]

    Fairness i n LLM-Generated Surveys, January 2025

    Andrés Abeliuk, Vanessa Gaete, and Naim Bro. Fairness i n LLM-Generated Surveys, January 2025. arXiv:2501.15351 [cs]

  26. [30]

    Aligning Language Models to User Opinions, May 2023

    EunJeong Hwang, Bodhisattwa Prasad Majumder, and Nike t Tandon. Aligning Language Models to User Opinions, May 2023. arXiv:2305.14929 [cs]

  27. [31]

    Xue, Peter S

    Mohammad Atari, Mona J. Xue, Peter S. Park, Damián Blasi , and Joseph Henrich. Which Humans?, September 2023

  28. [32]

    Whose Opinions Do Language Models Reflect?, March 2023

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo L ee, Percy Liang, and Tatsunori Hashimoto. Whose Opinions Do Language Models Reflect?, March 2023. arXi v:2303.17548 [cs]

  29. [33]

    More human than human: measuring ChatGPT political bias

    Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigue s. More human than human: measuring ChatGPT political bias. Public Choice, August 2023

  30. [34]

    What shapes your bias?

    Jisu Shin, Hoyun Song, Huije Lee, Soyeong Jeong, and Jon g Park. Ask LLMs Directly, “What shapes your bias?”: Measuring Social Bias in Large Language Models . In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics :...

  31. [36]

    Wichmann

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michae lis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neu ral networks. Nature Machine Intelligence, 2:665–673, 2020

  32. [37]

    Thomas McCoy, Ellie Pavlick, and Tal Linzen

    R. Thomas McCoy, Ellie Pavlick, and Tal Linzen. Right fo r the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 3428–3448, 2019

  33. [39]

    Do LLMs exhibit human-like response biases? A case study in survey d esign, November 2023

    Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Am eet Talwalkar, and Graham Neubig. Do LLMs exhibit human-like response biases? A case study in survey d esign, November 2023. arXiv:2311.04076 [cs]

  34. [40]

    Watts, and Mark E

    Tuan Dung Nguyen, Duncan J. Watts, and Mark E. Whiting. E mpirically evaluating commonsense intelligence in large language models with large-scale hum an judgments, May 2025. arXiv:2505.10309 [cs]

  35. [41]

    D. A. Dillman. Mail and internet surveys : the tailored design method . John Wiley & Sons, New York, 2nd edition, 2000

  36. [42]

    Two-Component Models of Socially Desi rable Responding

    Delroy Paulhus. Two-Component Models of Socially Desi rable Responding. Journal of Personality and Social Psychology, 46:598–609, March 1984

  37. [43]

    Argyle, Christopher A

    Lisa P. Argyle, Christopher A. Bail, Ethan C. Busby, Jos hua R. Gubler, Thomas Howe, Christopher Ryt- ting, Taylor Sorensen, and David Wingate. Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale. Proceedings of the Na...

  38. [44]

    Costello, Gordon Pennycook, and David G

    Thomas H. Costello, Gordon Pennycook, and David G. Rand . Durably reducing conspiracy beliefs through dialogues with AI. Science, 385(6714):eadq1814, September 2024. Publisher: America n Asso- ciation for the Advancement of Science

  39. [45]

    Engagement- Driven Content Generation with Large Language Models, Nove mber 2024

    Erica Coppolillo, Federico Cinus, Marco Minici, Franc esco Bonchi, and Giuseppe Manco. Engagement- Driven Content Generation with Large Language Models, Nove mber 2024. arXiv:2411.13187 [cs]. 21 Recalibrating the Compass

  40. [46]

    Fisher, Jennifer Pan, Yulia Tsvetkov, and Katharina Reinecke

    Jillian Fisher, Shangbin Feng, Robert Aron, Thomas Ric hardson, Yejin Choi, Daniel W. Fisher, Jennifer Pan, Yulia Tsvetkov, and Katharina Reinecke. Biased AI can I nfluence Political Decision-Making, November 2024. arXiv:2410.06415 [cs]

  41. [47]

    Evaluating the per suasive influence of political microtargeting with large language models

    Kobi Hackenburg and Helen Margetts. Evaluating the per suasive influence of political microtargeting with large language models. Proceedings of the National Academy of Sciences , 121(24):e2403116121, June 2024

  42. [48]

    Collaborating with AI Agents: Field Experiments on Teamwork, Produc- tivity, and Performance, March 2025

    Harang Ju and Sinan Aral. Collaborating with AI Agents: Field Experiments on Teamwork, Produc- tivity, and Performance, March 2025. arXiv:2503.18238 [cs ]

  43. [49]

    S. C. Matz, J. D. Teeny, S. S. Vaid, H. Peters, G. M. Harari , and M. Cerf. The potential of generative AI for personalized persuasion at scale. Scientific Reports , 14(1):4692, February 2024. Publisher: Nature Publishing Group

  44. [50]

    Bakker, Daniel Jarre tt, Hannah Sheahan, Martin J

    Michael Henry Tessler, Michiel A. Bakker, Daniel Jarre tt, Hannah Sheahan, Martin J. Chadwick, Raphael Koster, Georgina Evans, Lucy Campbell-Gillingham , Tantum Collins, David C. Parkes, Matthew Botvinick, and Christopher Summerfield. AI can help humans find common ground in dem...

  45. [51]

    Taly Reich and Jacob D. Teeny. Does artificial intellige nce cause artificial confidence? Generative artificial intelligence as an emerging social referent. Journal of Personality and Social Psychology , April 2025

  46. [52]

    Auditing multimodal large language m odels for context-aware content moderation, February 2025

    Thomas Davidson. Auditing multimodal large language m odels for context-aware content moderation, February 2025

  47. [53]

    Generative AI may backfire for counter- speech, November 2024

    Dominik Bär, Abdurahman Maarouf, and Stefan Feuerrieg el. Generative AI may backfire for counter- speech, November 2024. arXiv:2411.14986 [cs]

  48. [54]

    Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Th omas L. Griffiths. Explicitly unbiased large language models still form biased associations. Proceedings of the National Academy of Sciences , 122(8):e2416228122, February 2025. Publisher: Proceedin gs of the National Academ...

  49. [55]

    How persuasive is AI-generated propaganda? PNAS Nexus , 3(2):pgae034, February 2024

    Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Sta mos, and Michael Tomz. How persuasive is AI-generated propaganda? PNAS Nexus , 3(2):pgae034, February 2024

  50. [56]

    Large Language Models Understand and Can be Enhanced by Emotional Stimuli, November 2023

    Cheng Li, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxi n Hou, Jianxun Lian, Fang Luo, Qiang Yang, and Xing Xie. Large Language Models Understand and Can be Enhanced by Emotional Stimuli, November 2023. arXiv:2307.11760 [cs]

  51. [57]

    Artificial In telligence in Deliberation: The AI Penalty and the Emergence of a New Deliberative Divide, March 2025

    Andreas Jungherr and Adrian Rauchfleisch. Artificial In telligence in Deliberation: The AI Penalty and the Emergence of a New Deliberative Divide, March 2025. arXi v:2503.07690 [cs]

  52. [58]

    Yidan Yin, Nan Jia, and Cheryl J. Wakslak. AI can help peo ple feel heard, but an AI label diminishes this impact. Proceedings of the National Academy of Sciences , 121(14):e2319112121, April 2024. Publisher: Proceedings of the National Academy of Sciences

  53. [60]

    Jac kson

    Qiaozhu Mei, Yutong Xie, Walter Yuan, and Matthew O. Jac kson. A Turing test of whether AI chatbots are behaviorally similar to humans. Proceedings of the National Academy of Sciences , 121(9):e2313925121, February 2024. Publisher: Proceedin gs of the National Academy of Sciences

  54. [61]

    Arriaga, and Adam Tauman Kalai

    Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies, July 2 023. arXiv:2208.10264 [cs]

  55. [64]

    Maria del Rio-Chanona, Marco Pangallo, and Cars Homm es

    R. Maria del Rio-Chanona, Marco Pangallo, and Cars Homm es. Can Generative AI agents behave like humans? Evidence from laboratory market experiments, May 2 025. arXiv:2505.07457 [econ]

  56. [65]

    OASIS: Open Agent Social Interaction Simulations wit h One Million Agents, November 2024

    Ziyi Yang, Zaibin Zhang, Zirui Zheng, Yuxian Jiang, Ziy ue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, Prateek Gupta, Shuyue Hu, Zhenfe i Yin, Guohao Li, Xu Jia, Lijun 22 Recalibrating the Compass Wang, Bernard Ghanem, Huchuan Lu, Chaochao Lu, Wanli Ouyan...

  57. [66]

    Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddh arth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T. Rogers. Simulating Op inion Dynamics with Networks of LLM-based Agents, April 2024. arXiv:2311.09618 [physics]

  58. [67]

    Bernstein

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredit h Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative Agents: Interactive Simulacra of Hu man Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technol ogy, pages 1–22, ...

  59. [68]

    Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy, October 2024

    Zhenyu Guan, Xiangyu Kong, Fangwei Zhong, and Yizhou Wa ng. Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy, October 2024. arXiv:2407.06813 [c s]

  60. [69]

    War and Peace (WarAgent): Large Language Mo del-based Multi-Agent Simulation of World Wars, January 2024

    Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji , Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang. War and Peace (WarAgent): Large Language Mo del-based Multi-Agent Simulation of World Wars, January 2024. arXiv:2311.17227 [cs]

  61. [70]

    Modelling Political Coalition Negotiations Using LLM-based Agents, February 2 024

    Farhad Moghimifar, Yuan-Fang Li, Robert Thomson, and G holamreza Haffari. Modelling Political Coalition Negotiations Using LLM-based Agents, February 2 024. arXiv:2402.11712 [cs]

  62. [71]

    Artificial Leviathan: Exploring Social Evolu tion of LLM Agents Through the Lens of Hobbesian Social Contract Theory, July 2024

    Gordon Dai, Weijia Zhang, Jinhan Li, Siqi Yang, Chidera Onochie lbe, Srihas Rao, Arthur Caetano, and Misha Sra. Artificial Leviathan: Exploring Social Evolu tion of LLM Agents Through the Lens of Hobbesian Social Contract Theory, July 2024. arXiv:2406.1 4373 [cs]

  63. [72]

    What if LLMs Have Different World View s: Simulating Alien Civilizations with LLM-based Agents, December 2024

    Mingyu Jin, Beichen Wang, Zhaoqian Xue, Suiyuan Zhu, We nyue Hua, Hua Tang, Kai Mei, Mengnan Du, and Yongfeng Zhang. What if LLMs Have Different World View s: Simulating Alien Civilizations with LLM-based Agents, December 2024. arXiv:2402.13184 [c s]

  64. [73]

    Large language models empowered agent-based modeling and s imulation: a survey and perspectives

    Chen Gao, Xiaochong Lan, Nian Li, Yuan Yuan, Jingtao Din g, Zhilun Zhou, Fengli Xu, and Yong Li. Large language models empowered agent-based modeling and s imulation: a survey and perspectives. Humanities and Social Sciences Communications , 11(1):1–24, September 2024. Publish...

  65. [74]

    From Individ ual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents, De cember 2024

    Xinyi Mou, Xuanwen Ding, Qi He, Liang Wang, Jingcong Lia ng, Xinnong Zhang, Libo Sun, Jiayu Lin, Jie Zhou, Xuanjing Huang, and Zhongyu Wei. From Individ ual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents, De cember 2024. arXiv:2412.03563 [cs]

  66. [75]

    A survey on large language model based autonomous agents

    Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Ji ngsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong We n. A survey on large language model based autonomous agents. Frontiers of Computer Science , 18(6):186345, March 2024

  67. [76]

    Christopher A. Bail. Can Generative AI improve social s cience? Proceedings of the National Academy of Sciences, 121(21):e2314021121, May 2024. Publisher: Proceedings o f the National Academy of Sciences

  68. [77]

    Yeager, Christop her J

    Dorottya Demszky, Diyi Yang, David S. Yeager, Christop her J. Bryan, Margarett Clapper, Susannah Chandhok, Johannes C. Eichstaedt, Cameron Hecht, Jeremy Ja mieson, Meghann Johnson, Michaela Jones, Danielle Krettek-Cobb, Leslie Lai, Nirel JonesMitc hell, Desmond C. Ong, Carol S...

  69. [78]

    Generative AI Mee ts Open-Ended Survey Responses: Research Participant Use of AI and Homogenization

    Simone Zhang, Janet Xu, and AJ Alvero. Generative AI Mee ts Open-Ended Survey Responses: Research Participant Use of AI and Homogenization. Sociological Methods & Research, page 00491241251327130, May 2025. Publisher: SAGE Publications Inc

  70. [79]

    Michael E. W. Varnum, Nicolas Baumard, Mohammad Atari, and Kurt Gray. Large Language Models based on historical text could offer informative tools for be havioral science. Proceedings of the National Academy of Sciences , 121(42):e2407639121, October 2024. Publisher: Proceedi n...

  71. [80]

    Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks, August 2023

    Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyu an Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim. Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks, August 2023. arXiv:2 307.02477 [cs]

  72. [81]

    People Make Better Edits: Measuring the Efficacy o f LLM-Generated Counterfactually Augmented Data for Harmful Language Detection, November 20 23

    Indira Sen, Dennis Assenmacher, Mattia Samory, Isabel le Augenstein, Wil van der Aalst, and Clau- dia Wagner. People Make Better Edits: Measuring the Efficacy o f LLM-Generated Counterfactually Augmented Data for Harmful Language Detection, November 20 23. arXiv:2311.01270 [cs]....

  73. [82]

    Integrating Genera tive Artificial Intelligence into Social Sci- ence Research: Measurement, Prompting, and Simulation

    Thomas Davidson and Daniel Karell. Integrating Genera tive Artificial Intelligence into Social Sci- ence Research: Measurement, Prompting, and Simulation. Sociological Methods & Research , page 00491241251339184, May 2025. Publisher: SAGE Publication s Inc

  74. [83]

    Large AI models are cultural and social technologies

    Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. Large AI models are cultural and social technologies. Science, 387(6739):1153–1156, March 2025. Publisher: American As sociation for the Advancement of Science

  75. [84]

    J. M. Wing. Computational thinking. Communications of the ACM , 49(3):33–35, 2006. 24

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.