REVIEW 3 major objections 5 minor 20 references
Feeding GPT-4.1 demographic personas predicted the winning candidate in eight of nine 2024 states and reached 0.94 accuracy on a childhood-vaccine question, the paper reports.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 10:22 UTC pith:QPLFSCQY
load-bearing objection The negative findings on persona simulation hold up, but the 0.94 vaccine accuracy is uninterpretable: the authors added 30 unspecified features after seeing low test scores, and label leakage is the likely explanation. the 3 major comments →
Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that GPT-4.1 can, to a meaningful degree, simulate U.S. public opinion from written demographic profiles. In election forecasting, census-only meta personas correctly identified the winner in eight of nine states, with the miss being a swing state; richer objective personas performed worse, and all predicted distributions deviated sharply from actual vote shares. On healthcare, after adding 30 features describing healthcare habits and beliefs, GPT-4.1 reached 0.94 accuracy and 0.85 F1 on the childhood-vaccines question, while scoring far lower (0.75 accuracy) on an evenly split question about whether medical treatments are worth their costs. The paper also report
What carries the argument
The persona simulation framework: each persona is a JSON object of demographic and psychographic attributes, and GPT-4.1 is prompted with a forced multiple-choice opinion question, returning a predicted choice, an explanation, and the three features it says most influenced the choice. The paper uses four persona types: meta personas (census race/age/sex), objective personas (adding income, education, marital status), subjective personas (adding Big Five traits, religion, politics), and dataset personas built from actual survey respondents. The extracted feature-importance lists are what let the authors compare GPT-4.1's reasoning with logistic-regression coefficients.
Load-bearing premise
The 0.94 vaccine-accuracy score rests on 30 extra features the paper never lists; if those features carry the same vaccine opinions being predicted, the result is circular, and the eight-of-nine election score assumes nine states are a fair test.
What would settle it
Disclose the 30 added features: if any ask about childhood vaccines, COVID-19 vaccines, or past vaccination decisions, the 0.94 is not predictive. A second check is to re-run the survey-panel experiment with features chosen only on a randomly selected half of respondents; if held-out accuracy falls to the 90% majority baseline, the earlier gain was overfitting.
If this is right
- For lopsided survey questions, GPT-4.1 can reach high accuracy largely by predicting the dominant answer; the harder test is evenly split questions, where accuracy drops to around 0.75.
- State-level election winners can sometimes be recovered from coarse census-like personas, but simulated vote distributions are too biased to serve as a substitute for polling.
- Comparing GPT-4.1's self-reported feature importance with a logistic regression's coefficients exposes which demographic attributes the model leans on, aiding transparency.
- Dialogue simulation with demographic personas captures occupations and hobbies but fails to differentiate personality traits, so it is not yet a replacement for human focus groups.
- The same methodology could in principle extend to other opinion domains such as public health, lawmaking, and economics, as the authors propose.
Where Pith is reading between the lines
- The 0.94 vaccine-accuracy figure should be treated as preliminary until the 30 added features are disclosed; if any of them ask about vaccine opinions or past vaccination behavior, the score would reflect label leakage rather than simulation skill.
- A cleaner experiment would select features using only a training split of the survey panel and then evaluate on a held-out split; the paper's before/after design cannot rule out selection bias.
- The framework could be turned into a stereotype-audit tool by comparing GPT-4.1's per-demographic opinion distributions against real survey cross-tabulations.
- If the added features genuinely are non-leaking lifestyle variables, then mundane consumption and habit data may be enough to infer sensitive opinions, a privacy-relevant implication the paper does not explore.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates GPT-4.1 persona simulation for opinion prediction. It builds meta, objective, subjective, and dataset personas from the Personas dataset, ANES, and Pew's American Trends Panel Wave 123. For the 2024 U.S. election, it reports that GPT-4.1 with demographic meta personas predicts the state winner in eight of nine states, while objective personas yield only five of nine; with ANES data, GPT-4.1 obtains 0.610 accuracy versus 0.648 for logistic regression. For healthcare opinions, it uses four W123 questions and reports that adding 30 extra 'healthcare habits and beliefs' features raises accuracy/F1, with CHILD reaching accuracy 0.94 and F1 0.8468. It also qualitatively evaluates dialogues among three personas. The paper concludes that persona simulation is promising but biased.
Significance. If the quantitative claims were robust, the paper would provide useful evidence about LLM persona simulation for survey and election prediction, especially the comparison with logistic regression and the analysis of which persona features matter. Strengths include the use of publicly available datasets and a clear before/after experimental contrast. However, the headline claims are currently not supported: the 0.94 vaccine accuracy rests on an undisclosed post-hoc feature addition, and the election claim is based on nine states with severely miscalibrated vote distributions. As it stands, the paper is an interesting preliminary study rather than a convincing quantitative evaluation.
major comments (3)
- [Section IV-B and Table II] The paper's highest quantitative claim ('up to 0.94' for CHILD) rests on a before/after comparison that is not an independent prediction. The authors state that 30 features about 'healthcare habits and beliefs' were added 'when scores remained low'; the 30 features are never listed. Because W123 is an opinion survey, some of the added features may be attitudinal and semantically equivalent to the target questions (e.g., prior vaccine-safety beliefs). If so, GPT-4.1 is copying label information, and the F1 jump from 0.6036 to 0.8468 is exactly the signature of leakage. This is load-bearing: without a full list of the 30 features and a demonstration that none are the outcome or near-proxies, the 0.94 accuracy cannot be interpreted as evidence that persona simulation anticipates vaccine beliefs. The authors should also report results with only the 33 objective features as the primary evalua
- [Section V-A and Fig. 2] The 'eight out of nine states' claim is based on nine state-level outcomes with no confidence intervals or calibration. The reported vote distributions are severely miscalibrated (e.g., California 98.2% blue vs. 58.7% Harris/38.5% Trump; similar gaps in Maryland and Wyoming). With objective personas, only five of nine states are correct. The paper should report per-state absolute error, calibration plots, and uncertainty estimates, and characterize the result as a directional match on a small convenience sample rather than 'accurate prediction.'
- [Section V-A (ANES) and Section IV-B (W123)] Several accuracy claims have no uncertainty quantification. The ANES comparison (GPT-4.1 0.610 vs. logistic regression 0.648) relies on a single split, after listwise deletion reduces the sample from 5,521 to 2,882, raising representativeness concerns. The W123 experiment uses 100 randomly selected personas with no repeated seeds or confidence intervals. The authors should provide confidence intervals, seed variation, and a justification for listwise deletion (e.g., comparison of demographic distributions before/after deletion).
minor comments (5)
- [Table I] The table uses 'patients' for survey respondents; 'respondents' is more accurate. The '% of 1' column should explicitly define the reference option for each question, especially for CHOICE and HEALTH, where '1' is not self-evident.
- [Section V-C and Abstract] The dialogue evaluation is described as 'adhered well to personas' in the abstract, but the conclusion says 'personalities were not captured well.' This inconsistency should be resolved. The three-persona sample is anecdotal; the paper should say so and ideally include exact prompts and transcripts.
- [General] Typos and formatting: 'V oting' in Section V-A, 'V A' in the affiliation, inconsistent 'PERSONAS'/'Personas.' Figure captions should include sample sizes and ground-truth sources.
- [Reproducibility] No data/code availability statement is provided. The W123 feature list, the prompt templates, and the random seeds used for persona subsampling should be released for reproducibility.
- [References] References should be harmonized; for example, [13] is a website without a formal citation and [6]/[7] have incomplete author lists. Please follow journal style.
Circularity Check
The 0.94 vaccine-belief accuracy is an in-sample post-hoc feature-addition result; the added inputs are described as healthcare beliefs, making the headline 'prediction' target-related by the paper's own description.
specific steps
-
fitted input called prediction
[Section IV-B (Healthcare Opinion Prediction); Section V-B (Healthcare Opinion Prediction, Table II)]
"When scores remained low, we incorporated an additional 30 features for each persona, making a total of 63 features. These features provided more information about personas’ healthcare habits and beliefs, improving GPT-4.1’s performance. ... After adding more information about each persona’s vaccination habits and other healthcare-related beliefs, scores skyrocketed, as shown in Table II. At its peak, GPT-4.1 was able to anticipate beliefs about childhood vaccines with an accuracy of up to 0.94 and an F1 score of up to 0.8468."
The four outcomes (CHILD, COVID, CHOICE, HEALTH) are healthcare beliefs. The added predictors are described as 'healthcare habits and beliefs' and 'vaccination habits and other healthcare-related beliefs' — the same construct class as the target. They were added only after initial scores were low, and the evaluation appears to reuse the same sample, so the before/after comparison is in-sample tuning. Because the 30 features are never enumerated and no exclusion test shows they omit the target or semantic equivalents, the 0.94 accuracy cannot be read as predicting beliefs from demographic personas; if any added column is the CHILD response or a near-paraphrase, GPT-4.1 is copying the label from the input. The headline claim thus reduces, by the paper's own description, to feeding target-rel
full rationale
The election-forecasting component is evaluated against external state results and ANES individual vote labels, so it is not circular even though the 9-state sample is small. The dialogue component is qualitative and not presented as a derivation. The only load-bearing circularity risk is the W123 healthcare claim: Section IV-B states that 30 unlisted features about 'healthcare habits and beliefs' were added after low scores, and Section V-B attributes the jump to this addition. Because the target questions are themselves healthcare beliefs, the input and output overlap by construction unless the withheld feature list proves otherwise. The paper provides no such list or leakage test, and the before/after F1 improvement (CHILD 0.60→0.85) is exactly the pattern expected if target-related answers were added as inputs. This makes the highest quantitative claim (accuracy up to 0.94) partially circular and in-sample rather than a validated prediction. The rest of the paper's results are not circular.
Axiom & Free-Parameter Ledger
free parameters (4)
- Added 30 healthcare features =
30 features (33 -> 63)
- Nonresponse exclusion threshold =
50%
- Persona and subsample sizes =
1,000 per state; 100 for W123; 3 for dialogue
- Random selection seeds =
not reported
axioms (4)
- domain assumption Meta personas built from census distributions represent each state's voting population
- ad hoc to paper The 30 added W123 features are not the target opinions or near-proxies
- domain assumption GPT-4.1 responses reflect the persona rather than the model's prior opinions
- domain assumption Self-reported survey answers in ANES/W123 are ground truth
read the original abstract
Persona simulation involves utilizing large language models (LLMs) to anticipate human choices or interactions based on specific characteristic information. To further understand current limitations and future directions, we tested persona simulation in opinion prediction with GPT-4.1 (knowledge cutoff by June 2024). Using personas from nine U.S. states provided by Columbia University's Personas dataset, GPT-4.1 accurately predicted 2024 election outcomes in eight out of the nine states, only failing in one of the swing states. We then focused on opinions related to medicine and healthcare. With the American Trends Panel Wave 123 dataset from Pew Research Center, GPT-4.1 was able to anticipate beliefs about childhood vaccines with an accuracy of up to 0.94. Furthermore, we applied GPT-4.1 to generate conversations among personas and observed that the simulated dialogues and opinions adhered well to personas' personalities and backgrounds, albeit lacking natural human-like flow. Persona simulation proves to be a promising application of artificial intelligence as long as biases are addressed. In the near future, it will be beneficial to apply it to opinion analysis and reaction prediction in diverse fields ranging from public health to lawmaking to economics.
Figures
Reference graph
Works this paper leans on
-
[1]
Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines,
M. Sarstedt, S. Adler, L. Rau, and B. Schmitt, “Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines,” inPsychology & Marketing, 2024, pp. 1254–1270
2024
-
[2]
Out of One, Many: Using Language Models to Simulate Human Samples,
L. P. Argyle, E. C. Busby, N. Fulda, J. Gubler, C. Rytting, and D. Wingate, “Out of One, Many: Using Language Models to Simulate Human Samples,” inPolitical Analysis, 2023, pp. 337–351
2023
-
[3]
LLMs Generate Structurally Realistic Social Networksbut Overestimate Political Homophily,
S. Chang, A. Chaszczewicz, E. Wang, M. Josifovska, E. Pierson, and J. Leskovec, “LLMs Generate Structurally Realistic Social Networksbut Overestimate Political Homophily,” inProceedings of the Nineteenth International AAAI Conference on Web and Social Media, 2025, pp. 341–371
2025
-
[4]
Generative agents: Interactive simulacra of human behavior,
J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” inProceedings of the 36th annual acm symposium on user interface software and technology, 2023, pp. 1–22
2023
-
[5]
MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education,
M. Yue, W. Lyu, W. Mifdal, J. Suh, Y . Zhang, and Z. Yao, “MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education,”arXiv preprint arXiv:2404.06711, 2024
arXiv 2024
-
[6]
LLM Generated Persona is a Promise with a Catch,
A. Li, H. Chen, H. Namkoong, and T. Peng, “LLM Generated Persona is a Promise with a Catch,” inarXiv preprint arXiv:2503.16527, 2025
Pith/arXiv arXiv 2025
-
[7]
Generative Agent Simulations of 1,000 People,
J. S. Parket al., “Generative Agent Simulations of 1,000 People,” in arXiv preprint arXiv:2411.10109, 2024
Pith/arXiv arXiv 2024
-
[8]
(2025) Introducing GPT-4.1 in the API
OpenAI. (2025) Introducing GPT-4.1 in the API. [Online]. Available: https://openai.com/index/gpt-4-1/
2025
-
[9]
PAARS: Persona Aligned Agentic Retail Shoppers,
S. Mansour, L. Perelli, L. Mainetti, G. Davidson, and S. D’Amato, “PAARS: Persona Aligned Agentic Retail Shoppers,” inProceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025), 2025, pp. 143–159
2025
-
[10]
Artificial Intelligence Simulation of Adolescents’ Responses to Vaping-Prevention Messages,
P. Sheeranet al., “Artificial Intelligence Simulation of Adolescents’ Responses to Vaping-Prevention Messages,” inJAMA Pediatrics, 2024
2024
-
[11]
From Persona to Personalization: A Survey on Role- Playing Language Agents,
J. Chenet al., “From Persona to Personalization: A Survey on Role- Playing Language Agents,”Transactions on Machine Learning Re- search, 2024
2024
-
[12]
PATIENTSIM: A Persona-Driven Simulator for Re- alistic Doctor-Patient Interactions,
D. Kyunget al., “PATIENTSIM: A Persona-Driven Simulator for Re- alistic Doctor-Patient Interactions,” inarXiv preprint arXiv:2505.17818, 2025
arXiv 2025
-
[13]
Character.AI,
N. Shazeer and D. de Freitas, “Character.AI,” https://character.ai/, 2025
2025
-
[14]
In- Context Impersonation Reveals Large Language Models’ Strengths and Biases,
L. Salewski, S. Alaniz, I. Rio-Torto, E. Schulz, and Z. Akata, “In- Context Impersonation Reveals Large Language Models’ Strengths and Biases,” inProceedings of the 37th International Conference on Neural Information Processing Systems, 2023, pp. 72 044 – 72 057
2023
-
[15]
(2025) Personas
Tianyi-Lab. (2025) Personas. [Online]. Available: https://huggingface. co/datasets/Tianyi-Lab/Personas
2025
-
[16]
(2025) ANES 2024 Time Series Study Full Release [dataset and documentation]
American National Election Studies. (2025) ANES 2024 Time Series Study Full Release [dataset and documentation]. [Online]. Available: https://electionstudies.org/data-center/2024-time-series-study/
2025
-
[17]
(2023) American Trends Panel Wave 123
Pew Research Center. (2023) American Trends Panel Wave 123. [Online]. Available: https://www.pewresearch.org/dataset/ american-trends-panel-wave-123/
2023
-
[18]
Llama-3.1-70b large language model,
Meta AI, “Llama-3.1-70b large language model,” https://huggingface. co/meta-llama/Llama-3.1-70B, 2024, released July 23, 2024
2024
-
[19]
(2024) Election 2024: Presidential results
CNN Politics. (2024) Election 2024: Presidential results. [Online]. Avail- able: https://www.cnn.com/election/2024/results/president?election-data
2024
-
[20]
V oting patterns in the 2024 election,
H. Hartig, S. Keeter, A. Daniller, and T. Van Green, “V oting patterns in the 2024 election,” inPew Research Center, 2025
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.