Pith. sign in

REVIEW 3 major objections 5 minor 20 references

Feeding GPT-4.1 demographic personas predicted the winning candidate in eight of nine 2024 states and reached 0.94 accuracy on a childhood-vaccine question, the paper reports.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 10:22 UTC pith:QPLFSCQY

load-bearing objection The negative findings on persona simulation hold up, but the 0.94 vaccine accuracy is uninterpretable: the authors added 30 unspecified features after seeing low test scores, and label leakage is the likely explanation. the 3 major comments →

arxiv 2607.20589 v1 pith:QPLFSCQY submitted 2026-07-22 cs.CL

Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1

classification cs.CL
keywords persona simulationlarge language modelsGPT-4.1opinion predictionelection forecastingvaccine opinionsdemographic biasdialogue generation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper attempts to establish that GPT-4.1, when given demographic personas, can predict human opinions—state-level election winners and vaccine beliefs—with high raw accuracy, and that persona simulation is therefore a promising but bias-prone tool for social-science research. Using census-based personas for nine states, the authors report that GPT-4.1 picked the winning 2024 presidential candidate in eight of nine states. Using a large national survey panel, it reached 0.94 accuracy on a childhood-vaccine opinion question after enriching each persona with 30 additional healthcare-related features. The authors' central caveat is that these numbers hide systematic bias: predicted vote distributions are far from actual results, simulated racial groups are over-homogenized, and generated conversations reflect demographic labels but not individual personalities.

Core claim

The paper's central claim is that GPT-4.1 can, to a meaningful degree, simulate U.S. public opinion from written demographic profiles. In election forecasting, census-only meta personas correctly identified the winner in eight of nine states, with the miss being a swing state; richer objective personas performed worse, and all predicted distributions deviated sharply from actual vote shares. On healthcare, after adding 30 features describing healthcare habits and beliefs, GPT-4.1 reached 0.94 accuracy and 0.85 F1 on the childhood-vaccines question, while scoring far lower (0.75 accuracy) on an evenly split question about whether medical treatments are worth their costs. The paper also report

What carries the argument

The persona simulation framework: each persona is a JSON object of demographic and psychographic attributes, and GPT-4.1 is prompted with a forced multiple-choice opinion question, returning a predicted choice, an explanation, and the three features it says most influenced the choice. The paper uses four persona types: meta personas (census race/age/sex), objective personas (adding income, education, marital status), subjective personas (adding Big Five traits, religion, politics), and dataset personas built from actual survey respondents. The extracted feature-importance lists are what let the authors compare GPT-4.1's reasoning with logistic-regression coefficients.

Load-bearing premise

The 0.94 vaccine-accuracy score rests on 30 extra features the paper never lists; if those features carry the same vaccine opinions being predicted, the result is circular, and the eight-of-nine election score assumes nine states are a fair test.

What would settle it

Disclose the 30 added features: if any ask about childhood vaccines, COVID-19 vaccines, or past vaccination decisions, the 0.94 is not predictive. A second check is to re-run the survey-panel experiment with features chosen only on a randomly selected half of respondents; if held-out accuracy falls to the 90% majority baseline, the earlier gain was overfitting.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For lopsided survey questions, GPT-4.1 can reach high accuracy largely by predicting the dominant answer; the harder test is evenly split questions, where accuracy drops to around 0.75.
  • State-level election winners can sometimes be recovered from coarse census-like personas, but simulated vote distributions are too biased to serve as a substitute for polling.
  • Comparing GPT-4.1's self-reported feature importance with a logistic regression's coefficients exposes which demographic attributes the model leans on, aiding transparency.
  • Dialogue simulation with demographic personas captures occupations and hobbies but fails to differentiate personality traits, so it is not yet a replacement for human focus groups.
  • The same methodology could in principle extend to other opinion domains such as public health, lawmaking, and economics, as the authors propose.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The 0.94 vaccine-accuracy figure should be treated as preliminary until the 30 added features are disclosed; if any of them ask about vaccine opinions or past vaccination behavior, the score would reflect label leakage rather than simulation skill.
  • A cleaner experiment would select features using only a training split of the survey panel and then evaluate on a held-out split; the paper's before/after design cannot rule out selection bias.
  • The framework could be turned into a stereotype-audit tool by comparing GPT-4.1's per-demographic opinion distributions against real survey cross-tabulations.
  • If the added features genuinely are non-leaking lifestyle variables, then mundane consumption and habit data may be enough to infer sensitive opinions, a privacy-relevant implication the paper does not explore.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper evaluates GPT-4.1 persona simulation for opinion prediction. It builds meta, objective, subjective, and dataset personas from the Personas dataset, ANES, and Pew's American Trends Panel Wave 123. For the 2024 U.S. election, it reports that GPT-4.1 with demographic meta personas predicts the state winner in eight of nine states, while objective personas yield only five of nine; with ANES data, GPT-4.1 obtains 0.610 accuracy versus 0.648 for logistic regression. For healthcare opinions, it uses four W123 questions and reports that adding 30 extra 'healthcare habits and beliefs' features raises accuracy/F1, with CHILD reaching accuracy 0.94 and F1 0.8468. It also qualitatively evaluates dialogues among three personas. The paper concludes that persona simulation is promising but biased.

Significance. If the quantitative claims were robust, the paper would provide useful evidence about LLM persona simulation for survey and election prediction, especially the comparison with logistic regression and the analysis of which persona features matter. Strengths include the use of publicly available datasets and a clear before/after experimental contrast. However, the headline claims are currently not supported: the 0.94 vaccine accuracy rests on an undisclosed post-hoc feature addition, and the election claim is based on nine states with severely miscalibrated vote distributions. As it stands, the paper is an interesting preliminary study rather than a convincing quantitative evaluation.

major comments (3)
  1. [Section IV-B and Table II] The paper's highest quantitative claim ('up to 0.94' for CHILD) rests on a before/after comparison that is not an independent prediction. The authors state that 30 features about 'healthcare habits and beliefs' were added 'when scores remained low'; the 30 features are never listed. Because W123 is an opinion survey, some of the added features may be attitudinal and semantically equivalent to the target questions (e.g., prior vaccine-safety beliefs). If so, GPT-4.1 is copying label information, and the F1 jump from 0.6036 to 0.8468 is exactly the signature of leakage. This is load-bearing: without a full list of the 30 features and a demonstration that none are the outcome or near-proxies, the 0.94 accuracy cannot be interpreted as evidence that persona simulation anticipates vaccine beliefs. The authors should also report results with only the 33 objective features as the primary evalua
  2. [Section V-A and Fig. 2] The 'eight out of nine states' claim is based on nine state-level outcomes with no confidence intervals or calibration. The reported vote distributions are severely miscalibrated (e.g., California 98.2% blue vs. 58.7% Harris/38.5% Trump; similar gaps in Maryland and Wyoming). With objective personas, only five of nine states are correct. The paper should report per-state absolute error, calibration plots, and uncertainty estimates, and characterize the result as a directional match on a small convenience sample rather than 'accurate prediction.'
  3. [Section V-A (ANES) and Section IV-B (W123)] Several accuracy claims have no uncertainty quantification. The ANES comparison (GPT-4.1 0.610 vs. logistic regression 0.648) relies on a single split, after listwise deletion reduces the sample from 5,521 to 2,882, raising representativeness concerns. The W123 experiment uses 100 randomly selected personas with no repeated seeds or confidence intervals. The authors should provide confidence intervals, seed variation, and a justification for listwise deletion (e.g., comparison of demographic distributions before/after deletion).
minor comments (5)
  1. [Table I] The table uses 'patients' for survey respondents; 'respondents' is more accurate. The '% of 1' column should explicitly define the reference option for each question, especially for CHOICE and HEALTH, where '1' is not self-evident.
  2. [Section V-C and Abstract] The dialogue evaluation is described as 'adhered well to personas' in the abstract, but the conclusion says 'personalities were not captured well.' This inconsistency should be resolved. The three-persona sample is anecdotal; the paper should say so and ideally include exact prompts and transcripts.
  3. [General] Typos and formatting: 'V oting' in Section V-A, 'V A' in the affiliation, inconsistent 'PERSONAS'/'Personas.' Figure captions should include sample sizes and ground-truth sources.
  4. [Reproducibility] No data/code availability statement is provided. The W123 feature list, the prompt templates, and the random seeds used for persona subsampling should be released for reproducibility.
  5. [References] References should be harmonized; for example, [13] is a website without a formal citation and [6]/[7] have incomplete author lists. Please follow journal style.

Circularity Check

1 steps flagged

The 0.94 vaccine-belief accuracy is an in-sample post-hoc feature-addition result; the added inputs are described as healthcare beliefs, making the headline 'prediction' target-related by the paper's own description.

specific steps
  1. fitted input called prediction [Section IV-B (Healthcare Opinion Prediction); Section V-B (Healthcare Opinion Prediction, Table II)]
    "When scores remained low, we incorporated an additional 30 features for each persona, making a total of 63 features. These features provided more information about personas’ healthcare habits and beliefs, improving GPT-4.1’s performance. ... After adding more information about each persona’s vaccination habits and other healthcare-related beliefs, scores skyrocketed, as shown in Table II. At its peak, GPT-4.1 was able to anticipate beliefs about childhood vaccines with an accuracy of up to 0.94 and an F1 score of up to 0.8468."

    The four outcomes (CHILD, COVID, CHOICE, HEALTH) are healthcare beliefs. The added predictors are described as 'healthcare habits and beliefs' and 'vaccination habits and other healthcare-related beliefs' — the same construct class as the target. They were added only after initial scores were low, and the evaluation appears to reuse the same sample, so the before/after comparison is in-sample tuning. Because the 30 features are never enumerated and no exclusion test shows they omit the target or semantic equivalents, the 0.94 accuracy cannot be read as predicting beliefs from demographic personas; if any added column is the CHILD response or a near-paraphrase, GPT-4.1 is copying the label from the input. The headline claim thus reduces, by the paper's own description, to feeding target-rel

full rationale

The election-forecasting component is evaluated against external state results and ANES individual vote labels, so it is not circular even though the 9-state sample is small. The dialogue component is qualitative and not presented as a derivation. The only load-bearing circularity risk is the W123 healthcare claim: Section IV-B states that 30 unlisted features about 'healthcare habits and beliefs' were added after low scores, and Section V-B attributes the jump to this addition. Because the target questions are themselves healthcare beliefs, the input and output overlap by construction unless the withheld feature list proves otherwise. The paper provides no such list or leakage test, and the before/after F1 improvement (CHILD 0.60→0.85) is exactly the pattern expected if target-related answers were added as inputs. This makes the highest quantitative claim (accuracy up to 0.94) partially circular and in-sample rather than a validated prediction. The rest of the paper's results are not circular.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

No fundamentally new theoretical entities; the 'personas' are existing datasets. The main unaccounted cost is the hidden feature list and the arbitrary design choices.

free parameters (4)
  • Added 30 healthcare features = 30 features (33 -> 63)
    Introduced after seeing low W123 scores; no pre-registration; drives Table II improvements.
  • Nonresponse exclusion threshold = 50%
    Questions with >50% nonresponse and patients with any missing answers removed; arbitrary cutoff.
  • Persona and subsample sizes = 1,000 per state; 100 for W123; 3 for dialogue
    No power analysis or justification; small N for headline claims.
  • Random selection seeds = not reported
    Random samples not reproducible.
axioms (4)
  • domain assumption Meta personas built from census distributions represent each state's voting population
    Election inference relies on this; objective/subjective personas generated by Llama may add bias (Section III).
  • ad hoc to paper The 30 added W123 features are not the target opinions or near-proxies
    Features called 'healthcare habits and beliefs'; if they include vaccine sentiment, CHILD/COVID predictions are leaked (Section IV-B).
  • domain assumption GPT-4.1 responses reflect the persona rather than the model's prior opinions
    All accuracy scores treat the model's forced-choice output as persona belief; no debiasing or counterfactual control.
  • domain assumption Self-reported survey answers in ANES/W123 are ground truth
    Standard but unstated; no discussion of response bias.

pith-pipeline@v1.3.0-alltime-deepseek · 6439 in / 14863 out tokens · 120547 ms · 2026-08-01T10:22:31.355448+00:00 · methodology

0 comments
read the original abstract

Persona simulation involves utilizing large language models (LLMs) to anticipate human choices or interactions based on specific characteristic information. To further understand current limitations and future directions, we tested persona simulation in opinion prediction with GPT-4.1 (knowledge cutoff by June 2024). Using personas from nine U.S. states provided by Columbia University's Personas dataset, GPT-4.1 accurately predicted 2024 election outcomes in eight out of the nine states, only failing in one of the swing states. We then focused on opinions related to medicine and healthcare. With the American Trends Panel Wave 123 dataset from Pew Research Center, GPT-4.1 was able to anticipate beliefs about childhood vaccines with an accuracy of up to 0.94. Furthermore, we applied GPT-4.1 to generate conversations among personas and observed that the simulated dialogues and opinions adhered well to personas' personalities and backgrounds, albeit lacking natural human-like flow. Persona simulation proves to be a promising application of artificial intelligence as long as biases are addressed. In the near future, it will be beneficial to apply it to opinion analysis and reaction prediction in diverse fields ranging from public health to lawmaking to economics.

Figures

Figures reproduced from arXiv: 2607.20589 by Sarah Y. Li, Ziyu Yao.

Figure 1
Figure 1. Figure 1: The persona simulation framework. 2016, and 2020 elections [2]. Li et al., on the other hand, created personas for each state using census data and used those to predict election results for 2016, 2020, and 2024 [6]. However, both did not highlight state-by-state voting distributions, which is a key focus in our work. We also noted important features that impacted GPT-4.1’s predictions, providing insight i… view at source ↗
Figure 3
Figure 3. Figure 3: Confusion matrices and feature importances for ANES using logistic [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 2
Figure 2. Figure 2: Voting distributions for nine states and across races. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

20 extracted references · 2 linked inside Pith

  1. [1]

    Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines,

    M. Sarstedt, S. Adler, L. Rau, and B. Schmitt, “Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines,” inPsychology & Marketing, 2024, pp. 1254–1270

  2. [2]

    Out of One, Many: Using Language Models to Simulate Human Samples,

    L. P. Argyle, E. C. Busby, N. Fulda, J. Gubler, C. Rytting, and D. Wingate, “Out of One, Many: Using Language Models to Simulate Human Samples,” inPolitical Analysis, 2023, pp. 337–351

  3. [3]

    LLMs Generate Structurally Realistic Social Networksbut Overestimate Political Homophily,

    S. Chang, A. Chaszczewicz, E. Wang, M. Josifovska, E. Pierson, and J. Leskovec, “LLMs Generate Structurally Realistic Social Networksbut Overestimate Political Homophily,” inProceedings of the Nineteenth International AAAI Conference on Web and Social Media, 2025, pp. 341–371

  4. [4]

    Generative agents: Interactive simulacra of human behavior,

    J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” inProceedings of the 36th annual acm symposium on user interface software and technology, 2023, pp. 1–22

  5. [5]

    MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education,

    M. Yue, W. Lyu, W. Mifdal, J. Suh, Y . Zhang, and Z. Yao, “MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education,”arXiv preprint arXiv:2404.06711, 2024

  6. [6]

    LLM Generated Persona is a Promise with a Catch,

    A. Li, H. Chen, H. Namkoong, and T. Peng, “LLM Generated Persona is a Promise with a Catch,” inarXiv preprint arXiv:2503.16527, 2025

  7. [7]

    Generative Agent Simulations of 1,000 People,

    J. S. Parket al., “Generative Agent Simulations of 1,000 People,” in arXiv preprint arXiv:2411.10109, 2024

  8. [8]

    (2025) Introducing GPT-4.1 in the API

    OpenAI. (2025) Introducing GPT-4.1 in the API. [Online]. Available: https://openai.com/index/gpt-4-1/

  9. [9]

    PAARS: Persona Aligned Agentic Retail Shoppers,

    S. Mansour, L. Perelli, L. Mainetti, G. Davidson, and S. D’Amato, “PAARS: Persona Aligned Agentic Retail Shoppers,” inProceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025), 2025, pp. 143–159

  10. [10]

    Artificial Intelligence Simulation of Adolescents’ Responses to Vaping-Prevention Messages,

    P. Sheeranet al., “Artificial Intelligence Simulation of Adolescents’ Responses to Vaping-Prevention Messages,” inJAMA Pediatrics, 2024

  11. [11]

    From Persona to Personalization: A Survey on Role- Playing Language Agents,

    J. Chenet al., “From Persona to Personalization: A Survey on Role- Playing Language Agents,”Transactions on Machine Learning Re- search, 2024

  12. [12]

    PATIENTSIM: A Persona-Driven Simulator for Re- alistic Doctor-Patient Interactions,

    D. Kyunget al., “PATIENTSIM: A Persona-Driven Simulator for Re- alistic Doctor-Patient Interactions,” inarXiv preprint arXiv:2505.17818, 2025

  13. [13]

    Character.AI,

    N. Shazeer and D. de Freitas, “Character.AI,” https://character.ai/, 2025

  14. [14]

    In- Context Impersonation Reveals Large Language Models’ Strengths and Biases,

    L. Salewski, S. Alaniz, I. Rio-Torto, E. Schulz, and Z. Akata, “In- Context Impersonation Reveals Large Language Models’ Strengths and Biases,” inProceedings of the 37th International Conference on Neural Information Processing Systems, 2023, pp. 72 044 – 72 057

  15. [15]

    (2025) Personas

    Tianyi-Lab. (2025) Personas. [Online]. Available: https://huggingface. co/datasets/Tianyi-Lab/Personas

  16. [16]

    (2025) ANES 2024 Time Series Study Full Release [dataset and documentation]

    American National Election Studies. (2025) ANES 2024 Time Series Study Full Release [dataset and documentation]. [Online]. Available: https://electionstudies.org/data-center/2024-time-series-study/

  17. [17]

    (2023) American Trends Panel Wave 123

    Pew Research Center. (2023) American Trends Panel Wave 123. [Online]. Available: https://www.pewresearch.org/dataset/ american-trends-panel-wave-123/

  18. [18]

    Llama-3.1-70b large language model,

    Meta AI, “Llama-3.1-70b large language model,” https://huggingface. co/meta-llama/Llama-3.1-70B, 2024, released July 23, 2024

  19. [19]

    (2024) Election 2024: Presidential results

    CNN Politics. (2024) Election 2024: Presidential results. [Online]. Avail- able: https://www.cnn.com/election/2024/results/president?election-data

  20. [20]

    V oting patterns in the 2024 election,

    H. Hartig, S. Keeter, A. Daniller, and T. Van Green, “V oting patterns in the 2024 election,” inPew Research Center, 2025