{"id":"32367f53-8633-4d41-bedb-089187d4bd06","arxiv_id":"2412.03162","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LLM-generated survey responses, conditioned on a respondent's prior answers or a generated persona, align with human responses at the distributional level and, less reliably, at the individual level.","lead":"This paper tests whether giving a large language model a survey respondent's demographics and earlier answers lets it predict that person's later answers, and whether those simulated answers reproduce the statistical patterns of the real survey. If it works, researchers could pre-test surveys on simulated respondents instead of running multiple expensive pilot rounds.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The individual-level replication claim is supported only by coarse agreement rates with no chance or majority baseline; distributional and path-coefficient metrics cannot establish per-person correspondence.","rationale":"The reader's weakest-assumption analysis already identifies the missing chance and majority baselines for the coarse agreement rates in Appendix B, and the present stress test finds that this is the most load-bearing gap in the paper. The paper's stated contribution is individual-level replication; the aggregate PLS-SEM and distributional results are informative but cannot establish that contribution. The consistency rates, as reported, could be matched by a trivial predictor if response distributions are imbalanced. Because this is a gap in evidence rather than a demonstrated contradiction, a conditional verdict remains appropriate: the authors should supply baselines, per-item breakdowns, and per-individual correlation metrics before the individual-level claim can be accepted. The verdict is therefore left unchanged from the reader's CONDITIONAL.","tokens_in":10686,"tokens_out":2027,"duration_ms":21680,"concrete_test":"Recompute the Appendix B consistency analysis per survey item with (a) a majority-class baseline that always predicts the modal human category, and (b) a chance-agreement baseline obtained by independently drawing from the human and LLM marginal distributions. Also compute, for each item, the exact match rate on the original 1-7 scale and the per-individual Spearman correlation between human and LLM responses. If LLM-Mirror or Omni-prompt agreement does not clearly exceed the majority baseline, or if per-individual correlations are near zero, the individual-level replication claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim is that LLM-Mirror responses \"closely follow human responses at the individual level.\" The evidence offered in the main text consists of Jensen-Shannon divergence, Wasserstein distance, and PLS-SEM path coefficients. All three are aggregate or distributional quantities: they compare marginal response distributions or covariance structures, not whether the LLM's prediction matches the same person's answer. Many different joint distributions can produce the same marginals and the same path coefficients, so these metrics cannot by themselves support a claim about individual-level replication. The only individual-level evidence is the consistency analysis in Appendix B (Tables 13-15), which groups the 7-point responses into three coarse categories (disagree 1-3, neutral 4, agree 5-7) and reports mean agreement rates of roughly 52-73%. These rates are not compared with a chance baseline, a majority-class baseline, or any per-individual correlation. If human responses and LLM responses both skew toward the modal category, always choosing the modal category would yield high coarse agreement without mirroring any individual. The 3-bin aggregation also makes high agreement easier to achieve than exact 7-point matching, so the reported percentages overstate the degree of replication. Additionally, the path-coefficient tables in Case 2 show that LLM-Mirror produces a significantly wrong sign for LIKE-to-LOY, a significant COMP-to-LOY where humans are insignificant, and an insignificant COMP-to-SAT where humans are significant; these divergences are acknowledged in the text but are difficult to reconcile with the claim of close individual-level replication. The central claim therefore hinges on an uncalibrated, coarse consistency metric. This is a correctness-risk concern about the strength of the evidence, not a disagreement with the general premise that LLMs can mimic some human response patterns.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LLM-Mirror, a generated-persona approach for pre-testing surveys. Using two existing survey datasets (an ad-blocker study and a bank-customer study), the authors compare four prompting strategies: Baseline (survey context only), Demo (adds demographics), Omni (adds prior questions and answers), and LLM-Mirror (uses a generated persona). They evaluate alignment through Jensen-Shannon divergence, Wasserstein distance, PLS-SEM path coefficients, and coarse category agreement rates in Appendix B. The paper claims that LLMs provided with respondent-specific information can reproduce individual human responses and that LLM-Mirror responses closely follow human responses at the individual level, with implications for using LLMs to pre-test surveys and structural models.","tokens_in":10920,"tokens_out":4065,"duration_ms":39949,"significance":"If the individual-level replication claim were established, the paper would make a useful practical contribution to survey pre-testing and to the literature on LLMs as social-science agents. The authors are right to move beyond distributional comparisons, and they deserve credit for reporting Case 2 failures openly and for releasing code and data. However, the current evidence does not support the headline individual-level claim: the main metrics are aggregate/distributional, and the only individual-level analysis is a coarse agreement rate with no chance or majority-class baseline. The paper's value is therefore more modest than claimed, though the underlying idea and datasets may be salvageable with additional validation.","major_comments":[{"comment":"The central claim that LLM-Mirror 'closely follows human responses at the individual level' rests on coarse agreement rates of 52.56% to 73.01%. These rates are reported as means with no chance baseline, no majority-class baseline, and no per-individual correlation or exact-match rate. Since the 7-point responses are collapsed into three bins (1-3, 4, 5-7), a predictor that always chooses the modal bin could achieve substantial agreement without mirroring any individual. The paper must report exact 7-point agreement, per-individual correlation coefficients (e.g., rank correlation), and comparison against a trivial majority-class predictor. Without these, the individual-level claim is unsupported.","section":"Appendix B, Tables 13-15"},{"comment":"Jensen-Shannon divergence, Wasserstein distance, and PLS-SEM path coefficients are aggregate or covariance-level quantities. They compare marginal distributions or structural relations, not whether the LLM's answer matches the same person's answer. Many joint distributions can share marginals and path coefficients while disagreeing at the individual level. The paper should either present individual-level metrics in the main text or explicitly restrict its conclusions to distributional and structural alignment, removing the 'individual-level' phrasing from the abstract.","section":"Evaluation Metrics"},{"comment":"In Study 1 and Case 1, the Omni-prompt condition provides the respondent's own prior answers to the explanatory-variable items and then asks the LLM to predict that same respondent's outcome items. High agreement may partly reflect statistical coupling between the provided and predicted variables rather than faithful replication of an individual's decision process. A control condition that pairs prior answers from one respondent with outcomes from another, or a held-out evaluation, is needed to quantify this effect. The LLM-Mirror construction also needs clarification: the introduction says personas are based on demographics and prior responses, but Table 1 lists both as absent; if prior responses are used to generate the persona, the same coupling concern applies.","section":"Methodology, Table 1 and Experimental Design"},{"comment":"The Case 2 results contain notable mismatches: LLM-Mirror and Omni produce a significant negative LIKE-to-LOY path where the human path is positive and significant, a significant COMP-to-LOY path where the human path is insignificant, and an insignificant COMP-to-SAT path where the human path is significant. These are not minor deviations; they reverse or change the qualitative conclusions of the structural model. The text acknowledges limitations, but the abstract and conclusion still claim that 'LLM-Mirror responses closely follow human responses at the individual level.' The claims need to be tempered to reflect that the approach works for some structures and fails for others, and the failure cases should be treated as boundary conditions.","section":"Table 6, Case 2"},{"comment":"No information is given about the number of LLM generations, sampling temperature, top-p, or other decoding parameters. The PLS-SEM standard errors appear to come from bootstrapping within a single set of LLM-generated responses, which does not capture variability across LLM samples. The agreement rates in Appendix B could change substantially across random draws. The paper should report repeated sampling statistics (e.g., means and confidence intervals over multiple generations) and the exact decoding settings used.","section":"Experiment Settings"}],"minor_comments":[{"comment":"The significance notation 'p < 0.5' is almost certainly a typo for 'p < 0.05'; please correct it.","section":"Table 2"},{"comment":"The relationship between the LLM-Mirror persona and the inputs used to construct it is unclear. Please specify precisely what information the generated persona contains and whether any of the outcome-related prior responses are encoded in the persona text.","section":"Methodology, Table 1"},{"comment":"The phrase 'Consistent Analysis' is vague; the paper should define whether the reported percentages are averaged over items, respondents, or both, and should report the full distribution of agreement rates rather than only the mean.","section":"Appendix B"},{"comment":"There is a typo: 'Jesen-Shannon divergence' should be 'Jensen-Shannon divergence'.","section":"Page near Table 9"},{"comment":"The caption 'Fraction of Theoretical model' is unclear; it should likely read 'Theoretical model' or 'Fraction of the theoretical model tested'.","section":"Figure 3"},{"comment":"The appendix heading says 'Damberg, Svenja, and Ringle (2023)' while the rest of the paper cites Damberg, Schwaiger, and Ringle (2022); please align the year and author order.","section":"Appendix D"},{"comment":"Please report the LLM version, inference hyperparameters (temperature, top-p, max tokens), and number of independent generations per participant so that the results are reproducible.","section":"Experiment Settings"}],"recommendation":"major_revision","confidential_remarks":"The paper's main claim is considerably stronger than the evidence. The distributional and path-coefficient analyses are reasonable as exploratory alignment checks, but the individual-level conclusion needs either much stronger evidence (exact-match rates, per-individual correlations, chance baselines, repeated sampling) or explicit removal from the abstract. Given that the authors have released code and data, adding these analyses should be feasible. If the authors cannot supply such evidence, the manuscript should be repositioned as a study of distributional and structural alignment only, in which case its novelty relative to prior work would need to be reassessed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — read this paper on LLM-generated personas for survey pre-testing. The headline is overstated: the paper claims individual-level replication, but the evidence only supports distributional and structural alignment. Still, there is something real here. The authors extend earlier LLM-survey work from aggregate distribution matching to PLS-SEM path-coefficient replication and introduce a generated-persona condition (LLM-Mirror) that is practical for pre-testing. Study 1 and Case 1 show path coefficients that line up reasonably with human data, and the distributional metrics improve clearly when prior responses are included. The authors also report Case 2 failures honestly, which counts for something.\n\nThe soft spot is the central claim. The only individual-level evidence is the three-bin consistency analysis in Appendix B (Tables 13-15), with mean agreement rates of 52-73%. No chance baseline, no majority-class baseline, no per-individual correlation. With Likert data skewed to one end, always choosing the modal category could hit those numbers; the 3-bin grouping makes the task much easier than exact matching. Distributional metrics and path coefficients are aggregate quantities; they cannot show that the model predicted the same person's answer. On top of that, the Omni and LLM-Mirror prompts feed the respondent's own prior answers into the prompt that generates the same person's outcomes, so agreement partly reflects statistical coupling, not faithful mirroring. Case 2 makes the overclaim concrete: generated paths include a wrong sign for LIKE-to-LOY, a significant COMP-to-LOY where humans are insignificant, and an insignificant COMP-to-SAT where humans are significant. The authors acknowledge these, but they sit awkwardly with the abstract's claim that responses 'closely follow human responses at the individual level.'\n\nFor a pre-testing tool, the aggregate results may already be useful—if you want to know whether a questionnaire will produce expected structural relationships, this is a cheap screen. But the individual-level claim needs much better evidence: per-individual metrics (correlation or accuracy against a chance baseline), repeated sampling, and ideally a no-prior-answers ablation. This paper deserves a serious referee, not because the claims are proven, but because the question is important and the method is promising. The referee should ask for the missing baselines and a more careful abstract.","headline":"Useful aggregate-level survey simulation, but the individual-level replication claim is not supported by the evidence.","tokens_in":11554,"tokens_out":3366,"would_cite":false,"duration_ms":28761,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that supplying a language model with a survey respondent's demographics and prior answers yields a generated persona whose responses track the real person's individual choices, and that this makes LLM-generated panels…","keywords":["LLM-Mirror","survey pre-testing","persona generation","individual-level response replication","PLS-SEM","GPT-4o","response distributions","survey simulation"],"falsifier":"Re-run the same per-question evaluation with a baseline that always picks the most common human response for each item, or samples from the human marginal distribution; if that baseline matches or exceeds the reported 52–73 percent consistency, individual-level mirroring is not demonstrated. A complementary test is to compute per-respondent correlations between LLM-Mirror and human answer vectors; near-zero or weakly positive correlations would falsify the claim that the model reproduces individual decision patterns.","tokens_in":10463,"feed_emoji":"📋","tokens_out":7486,"duration_ms":62938,"temperature":0.7,"pith_summary":"Surveys are expensive to pilot, so the authors ask whether a large language model can stand in for human respondents during pre-testing. Their central claim is that when the model is given each respondent's demographics and prior answers, it produces responses that match the real respondents not just on average but at the level of individual choices, and that the same holds when those details are condensed into a generated 'LLM-Mirror' persona. If true, researchers could test questionnaire items and structural hypotheses before fielding a survey, saving the cost and time of repeated pilots. The evidence is drawn from two published PLS-SEM datasets—an ad-blocker attitude survey and a bank-customer loyalty survey—and compares path coefficients, distribution distances, and per-question agreement.","feed_headline":"LLM personas can mirror individual survey responses","feed_subtitle":"Demographics plus prior answers let personas track individual replies, a step toward cheaper survey pre-tests.","key_machinery":"The load-bearing object is the LLM-Mirror persona: a short user profile generated by feeding the model a respondent's demographic attributes and their prior answers to survey items for the explanatory latent variables, then using that persona as the prompt for generating answers to the remaining items. The persona condenses respondent-specific information into a reusable prompt, which is what distinguishes the method from generic baseline or demographic-only prompting. The evaluation machinery is PLS-SEM, a partial-least-squares structural equation model that estimates path coefficients between latent variables; the authors compare coefficients estimated from human responses with those from each prompting condition, and they supplement that comparison with Jensen-Shannon divergence, Wasserstein distance, and per-question agreement.","core_discovery":"On the paper's own terms, the discovery is that a persona-based prompt—constructed from a respondent's demographics together with their answers to the questions that define the explanatory latent variables—lets GPT-4o generate survey responses that track the actual respondent's answers and reproduce the path coefficients of the original PLS-SEM model. In Study 1 and in Case 1 of Study 2, the Omni prompt (full prior questions and answers) and the LLM-Mirror prompt (persona only) both yield coefficients that match the human estimates in sign, significance, and rough magnitude, whereas prompts without prior responses diverge. The paper also reports Jensen-Shannon divergences and Wasserstein distances that are much smaller for these two conditions, and per-question consistency rates between 52 and 73 percent. In the more complex Case 2 model, the approach still tracks most paths but misses some insignificant relationships, which the authors attribute to model access and suggest fine-tuning could improve.","pith_inferences":["The 52–73 percent consistency rates are not benchmarked against a majority-class or chance predictor; until they are, 'individual-level replication' is an upper-bound claim.","PLS-SEM path coefficients are estimated from aggregate covariance structure, so close coefficient alignment can occur even if the LLM is not tracking each individual; per-respondent correlations across items would be a stricter test.","A natural next experiment is to use LLM-Mirror personas to generate a synthetic pilot, run the actual survey on a small human sample, and compare whether the pre-test would have caught the same design problems.","If the effect holds, the approach could be tested on underrepresented groups by conditioning personas on demographic strata, though representation errors in the LLM's prior would propagate into the mirror."],"forward_implications":["Survey pre-testing could be run on LLM-generated persona panels before human pilots, flagging weak items or non-significant paths early.","Because the Omni and LLM-Mirror prompts reproduce most path coefficients, researchers could use the approach to sanity-check which structural relationships are likely to survive data collection.","Demographics alone are not enough; the comparisons show that prior response information is what moves LLM outputs toward human answers, so persona construction should include substantive prior answers.","Even in the complex Case 2 model, most significant paths are recovered, suggesting the method scales to multi-mediator models with caveats.","The LLM-Mirror persona uses only a compact persona rather than full question-by-question history, making it a practical option when detailed prior responses are unavailable."],"supporting_citations":[{"why":"Supplies the Study 1 dataset on ad-blocker attitudes and the PLS-SEM model that the paper replicates.","marker":"Redondo and Aznar (2018)"},{"why":"Supplies the Study 2 dataset on bank customers and the structural model used in Case 1 and Case 2.","marker":"Damberg, Schwaiger, and Ringle (2022)"},{"why":"Provides the caution that distribution-level agreement does not imply individual-level agreement, motivating the paper's individual-level test.","marker":"Tjuatja et al. (2024)"},{"why":"Establishes the simulated-agent approach that this paper extends from aggregate patterns to individual responses.","marker":"Horton (2023)"},{"why":"Shows that demographic prompting can simulate group-level political preferences, the background that LLM-Mirror builds on.","marker":"Argyle et al. (2023)"},{"why":"Cited as evidence that fine-tuning improves LLM-human alignment and as the proposed remedy for the complex-model limitations reported in Case 2.","marker":"Gao et al. (2024)"}],"fun_headline_variants":["LLM-Mirror personas mimic individual survey answers","Persona prompts let LLMs mirror respondent replies","LLM personas from prior answers match survey responses","Survey pre-testing via LLM personas that mirror respondents","GPT-4o personas track individual survey answers closely"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that agreement rates of 52 to 73 percent between LLM-Mirror and human answers, with no chance-level or majority-class baseline, demonstrate individual-level replication; if a trivial predictor matches those rates, the central claim does not follow.","fun_headline_variants_meta":{"raw":{"variants":["LLM-Mirror personas mimic individual survey answers","Persona prompts let LLMs mirror respondent replies","LLM personas from prior answers match survey responses","Survey pre-testing via LLM personas that mirror respondents","GPT-4o personas track individual survey answers closely"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000168,"raw_usage":{"total_tokens":1267,"prompt_tokens":961,"completion_tokens":306,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":231}},"tokens_in":577,"tokens_out":306,"duration_ms":3718,"temperature":1.0,"reasoning_tokens":231,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:41:56.874744+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same per-question evaluation with a baseline that always picks the most common human response for each item, or samples from the human marginal distribution; if that baseline matches or exceeds the reported 52–73 percent consistency, individual-level mirroring is not demonstrated. A complementary test is to compute per-respondent correlations between LLM-Mirror and human answer vectors; near-zero or weakly positive correlations would falsify the claim that the model reproduces individual decision patterns.","supporting_citations":[],"review_version":1}