{"id":"20042bdd-7f1d-42bb-a7aa-0ec639b78912","arxiv_id":"2505.17479","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new public 2,058-person, 500-question, four-wave dataset with a test-retest benchmark, plus initial LLM digital twin evaluations at 71.7% individual-level accuracy.","lead":"This paper releases a public dataset of 2,058 US adults who each answered over 500 questions across four waves, including personality, cognitive, economic, and behavioral measures. It also tests whether LLM-based digital twins built from these answers can predict held-out responses, reaching about 72% accuracy versus an 81.7% human test-retest benchmark.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing unpersonalized LLM baseline means the 71.72% twin accuracy and the 87.67% relative-accuracy headline cannot be attributed to personalization; Table 3's normative-default failures suggest generic priors drive much of the result.","rationale":"After reading the paper in good faith, I believe the dataset itself is a valuable public contribution: it is large, richly annotated, and the authors include honest discussion of failures. The reader's conditional verdict is appropriate. My concern is not about data fabrication or internal inconsistency; it is about the interpretation of the headline accuracy. The paper's own aggregate results in Table 3 show systematic departures from human behavior in the direction of normative defaults, which strongly suggests the LLM is relying on priors rather than individual personas. This makes the missing unpersonalized baseline the most load-bearing gap for the central claim about digital twin fidelity. The representativeness concern raised by the reader is real but secondary: the paper's core contribution is a benchmark, and even a non-representative large sample is useful; moreover, the authors provide demographic tables that allow users to assess coverage. The lack of an unpersonalized control, by contrast, directly undermines the sentence 'highlighting the value of personalization and LLM-based simulation,' which is a key interpretive claim. A single control condition would resolve this. I therefore agree partially with the reader's weakest-assumption identification and recommend keeping the CONDITIONAL verdict: the dataset stands, but the headline performance claims should be re-analyzed with the control.","tokens_in":38808,"tokens_out":5916,"duration_ms":55709,"concrete_test":"Run the identical evaluation pipeline (same system prompt, temperature=0, post-processing) on the 88 holdout questions using GPT-4.1-mini under three conditions: (a) no persona, just the new question; (b) persona containing only the 14 demographic answers; (c) the full text persona used in the paper. Compare per-task and overall accuracy to the reported 71.72%. If conditions (a) or (b) are within 1-2 points of (c), personalization contributes little to the headline accuracy. Also report the chance-corrected ratio (twin - random)/(test-retest - random) for the full persona and the no-persona condition, to contextualize the 87.67% headline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim in the abstract—digital twin accuracy of 71.72% and a ratio of 87.67% to human test-retest accuracy—depends on interpreting the gap over random guessing (59.17%) as evidence that personalization works. Section 5.1 states 'The improvement over the baseline is consistent across all question types, highlighting the value of personalization and LLM-based simulation.' But Table 2 contains no LLM condition without a persona. In every LLM row the model receives at least some persona input, so the improvement over random could reflect the LLM's generic priors about how humans answer these classic tasks, not the specific individual's profile. Section 5.2 provides direct evidence for this concern: twins unanimously chose the normative option in the Allais problem, 98.8% gave the correct UN country count, only 4.0% refused the vaccine versus 45% of humans, and 100% of twins chose the maximizing strategy in probability matching. These are cases where a generic, well-informed LLM would answer 'correctly' without any persona. If an unpersonalized GPT-4.1-mini achieves accuracy near 71.72% on the same 88 holdout questions, the digital twin adds little over the base model. Moreover, the 87.67% ratio is a ratio of absolute accuracies; relative to chance, the twins capture (71.72-59.17)/(81.72-59.17)=55.6% of the test-retest advantage, far below the headline number.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Twin-2K-500, a publicly released dataset of 2,058 US participants who each answered 500+ questions across four waves, covering demographics, personality, cognitive ability, economic preferences, heuristics-and-biases experiments, and a pricing study. The authors construct LLM-based digital twins by feeding each participant's non-holdout responses as a text or JSON persona to GPT-4.1-mini (and several ablations), then evaluate the twins on 88 holdout questions from waves 1–3. They report a digital-twin accuracy of 71.72% against a human test-retest accuracy of 81.72%, a ratio they describe as 87.67%, and they assess aggregate replication of classic behavioral-economics effects. The paper also reports data-quality checks: correlations with face validity, replication of most known effects, and a test-retest baseline. The dataset and simulation code are publicly available.","tokens_in":39115,"tokens_out":4396,"duration_ms":47026,"significance":"If the central claims hold, this is a valuable community resource: it is, to my knowledge, the largest public dataset explicitly designed for digital-twin benchmarking, combining rich psychological profiles, behavioral tasks, and a test-retest baseline. The data collection appears careful, the authors honestly report failures (e.g., base-rate fallacy not replicating), and the holdout design is a genuine out-of-sample prediction rather than a circular fit. The public release of persona, evaluation, and retest JSON blocks will facilitate standardized comparisons. However, the headline quantitative claim about digital-twin fidelity currently lacks a crucial control—an unpersonalized LLM baseline—and the reported ratio overstates the twins' performance relative to chance. These issues weaken, but do not invalidate, the dataset contribution.","major_comments":[{"comment":"No unpersonalized LLM baseline is reported: every LLM row in Table 2 receives at least some persona input, so the 12.55-point improvement over random guessing could reflect generic LLM priors about how humans answer classic tasks rather than the specific individual's profile. The aggregate results in Table 3—twins unanimously choosing the normative option in the Allais problem, 98.8% giving the correct UN country count, 4.0% refusing the vaccine versus 45% of humans, and 100% choosing the maximizing strategy—suggest that a competent unpersonalized LLM would already achieve high accuracy on many holdout questions. Please add a condition with the same model and questions but no persona, and report the incremental accuracy attributable to personalization. This is load-bearing for the abstract's claim that the digital twins 'predict human behavior well.'","section":"§5.1, Table 2"},{"comment":"The headline '87.67% relative accuracy' is computed as a ratio of absolute accuracies (71.72/81.72 ≈ 87.76%, with a small arithmetic discrepancy in the reported value). This ratio does not account for the chance baseline. Measured relative to random guessing, the twins capture (71.72–59.17)/(81.72–59.17) ≈ 55.6% of the test-retest advantage, a much less favorable number. Please report both the absolute ratio and the chance-relative ratio, and use the more conservative framing in the abstract and Figure 1.","section":"§5.1, Abstract"},{"comment":"The sentence in §5.1—'The improvement over the baseline is consistent across all question types, highlighting the value of personalization and LLM-based simulation'—is contradicted by Table 3, where the twins fail to replicate the outcome bias, sunk cost fallacy, Allais problem, omission bias, and probability matching, and only partially replicate anchoring and nonseparability. Please qualify the individual-level claim in light of these aggregate replication failures, and discuss the possibility that the individual-level accuracy is driven by tasks where behavioral variation is small relative to generic priors.","section":"§5.2, Table 3 vs. §5.1"},{"comment":"The abstract and Section 2 describe the sample as 'representative US respondents,' but no attrition analysis is provided: 2,509 participants completed Wave 1 and 2,058 completed all four waves (82%). Without a comparison of completers versus dropouts on baseline demographics and key measures, the representativeness claim is not empirically supported. Please report such an analysis, or soften the representativeness claim to 'a Prolific sample recruited with demographic targets.'","section":"§2, Abstract"}],"minor_comments":[{"comment":"The abstract and Section 5.1 report 87.67% relative accuracy, while Figure 1 and the conclusion say '88%.' Please unify these numbers.","section":"Abstract, Figure 1"},{"comment":"The random-guessing baseline is defined only as 'chooses each answer from a random uniform distribution.' Please specify whether this is uniform over all response options or calibrated to empirical marginal distributions, since this affects the magnitude of the claimed improvement over baseline.","section":"§3, Figure 2"},{"comment":"The twins' row shows a partial replication ('✓✗') for anchoring, but the text only explains the UN country count. Please state explicitly which of the two anchoring items (redwood height or UN countries) replicated and which did not.","section":"Table 3, Anchoring and adjustment row"},{"comment":"The text states that digital twins 'always selected the normative option' in probability matching; please clarify whether this holds for both the card and dice versions, and whether any 'OTHER' strategies were observed.","section":"§5.2, Probability matching"},{"comment":"The system prompt instructs the model to answer as the persona but does not specify whether repeated sampling with temperature is used; for binary and numerical tasks, please state the number of samples per question and whether the reported accuracy is averaged over them.","section":"Technical Appendix A.1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the dataset is the contribution and it is real. 2,058 people, 500+ questions across four waves, psychological batteries, cognitive tests, economic preferences, behavioral economics replications, pricing data, and a test-retest benchmark on 88 holdout questions. Public release with code, and the paper is refreshingly honest about failures (base-rate fallacy not replicating, some non-separability nulls). That alone justifies referee time.\n\nThe paper does less well on its interpretive claim. The 71.72% twin accuracy / 87.67% ratio headline is presented as evidence for personalization, but no LLM condition without a persona is included. Every LLM row in Table 2 receives the persona. The aggregate results in Table 3 show exactly why this matters: twins unanimously choose the normative option in Allais, 98.8% know 54 UN countries, only 4% refuse the vaccine, and all maximize in probability matching. A generic, well-trained LLM would produce many of these responses from priors. Without a no-persona baseline, the improvement over random guessing cannot be attributed to the persona. That is a load-bearing gap for the paper's narrative, though not for the dataset.\n\nAlso worth flagging: the 87.67% is a ratio of absolute accuracies. Relative to chance, twins capture (71.72-59.17)/(81.72-59.17)=55.6% of the test-retest advantage. The paper does state the ratio to test-retest, and it reports the absolute numbers and chance baseline, but the abstract's '88% average accuracy relative to a test-retest benchmark' sits on the more favorable ratio. A careful reader will notice; a less careful one won't.\n\nMinor but real: final sample of 2,058 from 2,509 invited, and no attrition analysis. The paper calls the sample representative; that needs at least a demographics comparison.\n\nOn the central dataset claim, the evaluation design is sound: holdout questions are excluded from the persona, so the predictive test is genuine. The correlation checks are standard but appropriate. The citation pattern is fine. I would not desk-reject; I would send to a serious referee, with the explicit ask to require an unpersonalized LLM baseline and a reframed headline. For anyone working on LLM simulation, this is a resource worth having on the shelf.","headline":"A genuinely valuable public dataset for digital-twin research, but the headline accuracy claim overstates what is shown because no unpersonalized LLM baseline is run.","tokens_in":39651,"tokens_out":2005,"would_cite":true,"duration_ms":16971,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces a public dataset of 2,058 US adults who each answered 500+ questions, and reports that LLM-based digital twins predict individual holdout responses with 71.72% accuracy, reaching 87.67% of the human test-retest…","keywords":["digital twins","LLM simulation","persona","test-retest accuracy","representative sample","behavioral economics","survey dataset","benchmark"],"falsifier":"Compare the wave-1 demographics and key measures of the 451 participants who did not complete all four waves with those of the 2,058 completers; if the dropouts are systematically different (e.g., younger, lower-income, or lower-scoring), the 'representative US sample' claim behind the dataset fails. The paper reports no such comparison, so this is the most direct check.","tokens_in":38636,"feed_emoji":"🧠","tokens_out":7112,"duration_ms":49699,"temperature":0.7,"pith_summary":"This paper introduces Twin-2K-500, a public dataset of 2,058 US participants who each answered over 500 questions across four survey waves, covering demographics, personality, cognitive ability, economic preferences, replications of classic behavioral-economics experiments, and a product-pricing study. The final wave repeats earlier questions to measure how consistently humans answer the same tasks, giving a test-retest accuracy baseline of 81.72%. Using the non-repeated answers to build LLM-based digital twins, the paper reports that the twins predict individual holdout responses with 71.72% accuracy, which is 87.67% of the human test-retest accuracy. At the group level, the twin simulations replicate most of the classic experimental effects, with clear failures in cases where humans deviate from normative or 'rational' behavior, such as outcome bias, omission bias, and the Allais paradox. The dataset is offered as a public benchmark for developing and validating persona simulations and for social-science research more broadly.","feed_headline":"LLM digital twins hit 88% of human test-retest accuracy","feed_subtitle":"New public dataset of 2,058 people and 500-question profiles gives AI personas a human consistency benchmark.","key_machinery":"The load-bearing structure is the dataset's three-way split of each participant's answers: a persona record (all non-holdout responses from waves 1-3, formatted as text or JSON), an evaluation answer block (the wave 1-3 answers to the 88 holdout questions, used as ground truth), and a retest answer block (the same questions answered again in wave 4, used to compute the human test-retest benchmark). On top of this, the metric that makes the headline number meaningful is the accuracy definition: for each of the 17 tasks, accuracy is 1 minus the absolute deviation between predicted and ground-truth answer, normalized by the answer range, then averaged across respondents; for binary items this reduces to exact match. This lets the authors express twin performance as a single ratio, 71.72% / 81.72% = 87.67%, so that the human's own consistency is the ceiling against which machine imitation is measured.","core_discovery":"The paper's central claim is that a sufficiently rich individual-level survey record can be turned into an LLM persona that reproduces a large share of that person's behavioral responses, and that the field now has a public, large-scale resource for measuring that share honestly. The authors construct, for each of 2,058 participants, a structured record of roughly 412 non-holdout answers (the persona), hold out 88 questions covering 17 behavioral tasks, and have GPT-4.1-mini answer those questions as if it were the participant. The twin accuracy of 71.72% compares against a human test-retest accuracy of 81.72% on the same tasks, meaning the simulations capture roughly 88% of the consistency humans show with themselves across a two-week gap. On aggregate treatment effects, the twins replicate 6 of 10 between-subject and 2 of 5 within-subject classic results; the failures cluster in domains where humans show bias (outcome bias, omission bias, probability matching) or where the model cannot 'unlearn' textbook facts (e.g., the number of African countries in the UN). The paper concludes that the dataset is unique in combining breadth, representativeness, behavioral tasks, and a test-retest benchmark, and that it can accelerate transparent benchmarking of digital-twin methods.","pith_inferences":["If the representativeness holds, the dataset could serve as a calibration set for correcting LLM opinion surveys toward true population distributions, since it contains both the psychological profiles and the actual political and medical judgments of the same people.","The two-week gap between wave 1-3 and wave 4 gives an upper bound on what any deterministic mapping from a persona to behavior can achieve; methods that exceed ~82% are not 'better' than the human-human consistency, but may be exploiting task-specific regularities rather than true personalization.","A direct comparison of the 451 week-1 dropouts with the 2,058 completers on demographics and wave-1 measures would quantify the attrition bias and tell whether the 'representative' target is met; this analysis is not reported.","Because the twins reproduced the framing, conjunction, and myside effects but not the medical-domain biases, one testable hypothesis is that LLM personas are systematically more 'normative' and more trustful of medical authority than the US population; future surveys that vary the medical scenario could falsify this."],"forward_implications":["Future persona methods can be benchmarked on the same persona/holdout/retest split, with 71.72% as a strong baseline reported here.","The 87.67% ratio implies LLM twins capture most of the predictable, person-specific variance in these tasks; the remaining gap is likely a mix of model limitations and the irreducible noise in human re-answers.","The aggregate failures (outcome bias, omission bias, Allais, probability matching) show that some classic effects will not be reproduced by LLM twins, and that these are precisely the cases where human behavior is non-normative.","The public release of all four waves, including the retest block, lets researchers separate measurement error from genuine behavioral change.","Across the many prompt/format/model variations tested, accuracies cluster in a narrow band (67.9-71.9%), suggesting the architecture matters less than the quality of the persona data."],"supporting_citations":[{"why":"Establishes the 1,000-person digital-twin paradigm and reports the 85% replication figure that this paper's 87.67% ratio is compared against.","marker":"Park et al. (2024)"},{"why":"Source of most of the heuristics-and-biases tasks used as holdout questions and as replication targets (sunk cost, less-is-more, myside bias, WTA/WTP, omission bias, probability matching, dominator neglect).","marker":"Stanovich and West (2008)"},{"why":"Provides the anchoring-and-adjustment tasks that the twins fail to replicate for the UN item.","marker":"Tversky and Kahneman (1974)"},{"why":"Provides the outcome-bias and dictator-game tasks; the outcome bias is one of the between-subject effects the twins fail to replicate.","marker":"Baron and Hershey (1988)"},{"why":"Provides the base-rate problem, which was not replicated by humans in this sample nor by the twins.","marker":"Kahneman and Tversky (1973)"},{"why":"Provides the disease-framing task, one of the effects the twins do replicate.","marker":"Tversky and Kahneman (1981)"},{"why":"Source of the 40-product pricing study that generates the demand-curve comparison in wave 3, wave 4, and the twins.","marker":"Gui and Toubia (2023)"},{"why":"Supplies the economic-preference measures (discount, present bias, risk and loss aversion, trust) used to build the persona profiles.","marker":"Dean and Ortoleva (2019)"}],"fun_headline_variants":["LLM twins hit 88% of human test-retest consistency","2,000-person survey dataset benchmarks digital twin AI","Twin-2K-500: public data for building LLM personas","Human consistency becomes the bar for AI doublegangers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 2,058 people who completed all four waves still represent the US adult population, even though about 18% of the initially recruited 2,509 dropped out and the paper does not analyze whether completers differ from dropouts.","fun_headline_variants_meta":{"raw":{"variants":["LLM twins hit 88% of human test-retest consistency","2,000-person survey dataset benchmarks digital twin AI","Twin-2K-500: public data for building LLM personas","Human consistency becomes the bar for AI doublegangers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000293,"raw_usage":{"total_tokens":1776,"prompt_tokens":1082,"completion_tokens":694,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":698,"completion_tokens_details":{"reasoning_tokens":623}},"tokens_in":698,"tokens_out":694,"duration_ms":7555,"temperature":1.0,"reasoning_tokens":623,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:45:36.670895+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the wave-1 demographics and key measures of the 451 participants who did not complete all four waves with those of the 2,058 completers; if the dropouts are systematically different (e.g., younger, lower-income, or lower-scoring), the 'representative US sample' claim behind the dataset fails. The paper reports no such comparison, so this is the most direct check.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of most of the heuristics-and-biases tasks used as holdout questions and as replication targets (sunk cost, less-is-more, myside bias, WTA/WTP, omission bias, probability matching, dominator neglect)."},{"cited_title":"and Kahneman, D","cited_arxiv_id":null,"evidence_quote":"Provides the anchoring-and-adjustment tasks that the twins fail to replicate for the UN item."},{"cited_title":"and Hershey, J","cited_arxiv_id":null,"evidence_quote":"Provides the outcome-bias and dictator-game tasks; the outcome bias is one of the between-subject effects the twins fail to replicate."},{"cited_title":"and Tversky, A","cited_arxiv_id":null,"evidence_quote":"Provides the base-rate problem, which was not replicated by humans in this sample nor by the twins."},{"cited_title":"and Kahneman, D","cited_arxiv_id":null,"evidence_quote":"Provides the disease-framing task, one of the effects the twins do replicate."},{"cited_title":"and Ortoleva, P","cited_arxiv_id":null,"evidence_quote":"Supplies the economic-preference measures (discount, present bias, risk and loss aversion, trust) used to build the persona profiles."}],"review_version":1}