{"id":"0cbf01a7-a5b7-47b6-a507-e8da6d8c55c6","arxiv_id":"2608.08058","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LLM-based value simulation is systematically more accurate for wealthy, high-governance, individualist countries, and common interventions rarely fix the imbalance.","lead":"Across 59 countries and 9 open-source LLMs, this study measures how accurately models reproduce World Values Survey responses and finds large, systematic country-level gaps: richer, more digitally connected countries are simulated much better. It also tests four intervention strategies and shows that improving average accuracy or a target language's accuracy does not guarantee more equal performance across countries.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Contamination of pretraining corpora with WVS/ISSP survey data is the load-bearing unverified premise: if memorized items drive country-level accuracy, the headline inequality is an exposure artifact, not simulation capability.","rationale":"The reader's weakest_assumption correctly identifies contamination as the most load-bearing unverified premise. The paper's headline conclusion is that current open LLMs simulate human values very unevenly across countries, and the reported correlations with economic, technological, governance, and cultural indicators are interpreted as structural simulation bias. All of that hinges on the accuracy numbers being measures of simulation rather than memorized reproduction of public survey data. The paper provides no membership test, no exact-match probe, and no held-out post-cutoff benchmark; the ISSP replication does not fix the gap because its data are equally public and within plausible training windows. This is why the concern is load-bearing. I credit the paper for its internal robustness work: multi-model consistency, sensitivity analyses in Appendix B.3, the first-token versus repeated-generation validation in Appendix C.1, and the ISSP replication in Appendix F all support the internal validity of the measured accuracy differences. However, none of those checks addresses the training-data exposure question. The additional-information circularity is also real and should be disclosed more sharply, but it is secondary because it affects the intervention-level claims rather than the foundational inequality result. Since the reader already assigned CONDITIONAL on substantially these grounds, my stress-test pass does not change the verdict; it sharpens the reason why the condition is not merely a style preference but a prerequisite for interpreting the central claim.","tokens_in":42120,"tokens_out":4294,"duration_ms":50654,"concrete_test":"Re-run the foundational analysis of Section 5 on a survey benchmark released after the training cutoff of every evaluated model, such as a WVS Wave 8 or ISSP module fielded after 2024, using the same prompt protocol and aggregation. If country-level accuracy remains above chance and the GDP/Internet/GII/WGI and power-distance correlations replicate on this post-cutoff benchmark, the contamination concern is refuted for the headline inequality pattern. If accuracy for lower-resource countries drops toward chance while high-resource countries retain accuracy, memorization of WVS/ISSP data is implicated as the driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim — that open LLMs simulate human values unevenly across countries, with higher accuracy for wealthier, higher-governance, lower-power-distance populations — is framed as a property of simulation capability. The load-bearing premise is that WVS Wave 7 (2017–2022) ground truth, adopted in Section 4.1, has not leaked into the training corpora of the nine evaluated models, whose training cutoffs span 2023–2024. No contamination analysis is provided, despite WVS codebooks, wave files, and country-level value statistics being publicly available before those cutoffs. If models have memorized item texts and answer distributions, the first-token probabilities extracted in Appendix A.2 can reproduce survey margins without 'simulating' values. The ISSP robustness check in Appendix F does not settle this concern because ISSP 2020–2023 modules are also public and plausibly inside the training windows. The inequality pattern itself may survive contamination, but its interpretation as evidence of simulation capability, and the statement that LLM-based simulation reproduces structural inequalities, would reduce to a claim about training-data exposure. This is an unverified external premise rather than an internal inconsistency. A secondary but genuine issue is that the additional-information intervention feeds empirical WVS distributions for the same demographic subgroup into the prompt via the historical-memory module (Section 3.4), making its reported joint accuracy-equality gains partly circular; that affects the intervention guidance, not the foundational inequality result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces \"representational equality\" as a metric for LLM-based cross-country value simulation, defined as the evenness of simulation accuracy across countries and operationalized primarily through the coefficient of variation of country-level accuracies (EqCV). The authors evaluate nine open-source LLMs on 59 countries using World Values Survey Wave 7 as ground truth, report substantial cross-country inequality, and show that accuracy correlates positively with GDP per capita, Internet use, innovation, and governance quality and negatively with power distance. They then compare two intervention pathways—contextual adaptation (native-language prompting, additional information) and parametric modification (continued post-training, preference alignment)—and find that additional information most consistently improves both accuracy and equality, while preference alignment yields no systematic gains. Extensive sensitivity analyses include leave-one-dimension-out, alternative weighting, ISSP replication, and first-token versus repeated-generation validation.","tokens_in":42397,"tokens_out":5022,"duration_ms":53987,"significance":"If the central finding holds, this is a significant contribution to LLM-based social simulation and cross-cultural NLP, providing a formal evaluation framework, a broad empirical map of inequality, and a systematic comparison of intervention strategies. The paper's strengths are substantial: a 59-country, 2,420-subpopulation evaluation; nine open models spanning scales; four complementary equality indices; BH-FDR-corrected correlations; and careful robustness checks (LODO, weighting, ISSP, first-token validation). These elements make the descriptive inequality finding well supported. However, two premises are load-bearing and currently unverified: that the public WVS/ISSP ground truth has not leaked into training corpora, and that the additional-information intervention does not feed the target survey's own joint distribution into the prompt. Both affect the interpretation of the results rather than the internal consistency of the measurements.","major_comments":[{"comment":"The evaluation assumes that WVS Wave 7 (2017-2022) and ISSP 2020-2023 responses have not been memorized by the evaluated models, whose training cutoffs span 2023-2024. No contamination analysis is provided anywhere in the manuscript. This is load-bearing because the headline claim is about simulation capability; if models reproduce memorized survey margins, then the reported accuracy levels and their correlations with wealth, governance, and cultural indicators reflect training-data exposure rather than generalizable simulation. The ISSP robustness check in Appendix F does not settle this because those modules are equally public and fall inside the training windows. I request a contamination analysis, for example: evaluating on paraphrased question variants, comparing accuracy on items published after each model's cutoff, or running membership-inference probes on question-answer pairs from WVS/ISSP.","section":"Section 4.1, Section 5"},{"comment":"The additional-information intervention feeds the model the empirical WVS response distributions of the same demographic subgroup for non-target questions via the historical-memory module. Because the target questions come from the same WVS questionnaire, this gives the model direct information about the joint response distribution for that subgroup. The reported improvements under this intervention (e.g., GLM-4-9B EqCV dropping from 0.0486 to 0.0255 under BM25) are therefore partly by construction: the model is given the very survey's subgroup-level answer patterns as context. This does not establish a generalizable inference-time intervention for realistic settings where such subgroup-level survey data are unavailable. Please either reframe the additional-information setting as an oracle or upper-bound analysis, or remove it from the main intervention claims and conclusions.","section":"Section 3.4, Appendix D.1, Section 6.2"},{"comment":"The claim that some models demonstrate 'higher' or 'lower' equality (e.g., ChatGLM3-6B with EqCV=0.0286 versus Qwen2.5-72B with EqCV=0.0546) is made without uncertainty quantification for the equality indices. Table E.7 reports 95% confidence intervals for country-level accuracy but not for EqCV, EqGini, or the AE composite. Given that country-level accuracies have overlapping confidence intervals across models, it is unclear whether these equality differences are statistically significant. Please provide bootstrap confidence intervals or a direct significance test for the equality indices, or temper the model-ranking claims to a descriptive level.","section":"Section 5.1, Table E.7"}],"minor_comments":[{"comment":"The correlation heatmap cells are too small to read the bolding and correlation values; please enlarge the figure or split it into per-factor-family panels.","section":"Figure 3"},{"comment":"The notation for the primary index alternates between 'EqCV' and 'Eq CV'; please standardize to one form and use it consistently in the text, tables, and equations.","section":"Throughout"},{"comment":"The AE composite metric multiplies a distance (JSD) by a dispersion (CV) and takes the geometric mean; the text would benefit from a clearer explanation of why this particular combination is appropriate and how to interpret its scale.","section":"Equation (2)"},{"comment":"The minimum-support threshold of 10 respondents is stated but not justified; please add a sensitivity analysis over this threshold or explain why 10 is sufficient for stable subgroup-level distributions.","section":"Appendix B.1"},{"comment":"The limitations section is candid about several threats but does not mention the possibility of training-data contamination; please add a paragraph acknowledging this risk and the need for contamination checks in future work.","section":"Section 8 (Limitations)"},{"comment":"Several reference entries contain typos or incomplete information (e.g., Rokeach 1973, 'Free peess'; some URLs in Appendix A.3). Please run a careful copyedit of the reference list.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The contamination concern is the same class of issue that has affected many LLM benchmarks, and the authors should be expected to address it directly with at least a basic analysis. The additional-information leakage is more serious because it affects the validity of one of the paper's main intervention conclusions; reframing that setting as an oracle is the most straightforward fix. I do not see a reason to reject, as the core descriptive inequality finding appears robust to the sensitivity checks already reported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. The core empirical claim — open LLMs simulate WVS value responses much more accurately for wealthy, well-governed, individualist countries — is well supported and looks real. EqCV across nine models and 59 countries, the extensive sensitivity analyses (LODO, weighting, ISSP replication, first-token validation), and the BH-FDR corrections all point the same way. The US/Iraq case study is concrete and informative. That is a meaningful addition to prior work that measured average accuracy or directional bias rather than cross-country dispersion as an equality property.\n\nThe soft spots are real but unequal in weight. First, contamination: WVS Wave 7 and ISSP 2020–2023 were public before the evaluated models' training cutoffs, and the paper provides no contamination analysis. That matters because the headline interprets accuracy as 'simulation capability' and concludes that LLM-based simulation reproduces structural inequalities. If models have memorized survey margins, the inequality could be an exposure artifact. This does not undermine the descriptive inequality result — whether memorized or not, the model's outputs are systematically uneven across countries — but it does undermine the simulation-capability interpretation. The omission is genuine; the authors should add a contamination check.\n\nSecond, the additional-information intervention feeds empirical WVS response distributions for the same demographic subgroup into the prompt via the historical-memory module. The reported accuracy and equality gains are therefore partly circular: the model is given answer distributions on related items from the same subgroup. That affects the intervention guidance, not the foundational inequality result.\n\nMinor points: the AE composite takes a geometric mean of JSD and CV, which are on different scales; a short justification would help. The native-language results use official WVS translations, which is appropriate but also means the 'language cue' is confounded with the exact item wording that may appear in training data.\n\nThe paper is carefully done, and the conditional verdict is fair. It deserves a serious referee. I would send it out and ask for a contamination analysis plus a reframing of the intervention-level claims. The audience is computational social science and LLM fairness evaluation; they will get real value from the measurement framework even if the capability interpretation needs revising.","headline":"Solid, well-robustness-checked inequality finding; the 'simulation capability' framing is hostage to an unexamined benchmark-contamination premise, and the additional-information intervention is partly circular.","tokens_in":42939,"tokens_out":1376,"would_cite":true,"duration_ms":41670,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current open LLMs simulate human values unevenly across countries, systematically better for wealthier, better-governed, more individualist populations, and standard interventions do not reliably fix the gap.","keywords":["representational equality","value simulation","cross-country bias","large language models","World Values Survey","simulation accuracy","contextual adaptation","parametric modification"],"falsifier":"Run a contamination-controlled replication: measure n-gram or paraphrase overlap between the 160 benchmark questions and each model's training corpus, then re-score accuracy only on questions with no measurable overlap. If the GDP, Internet-use, and power-distance correlations with accuracy vanish or reverse on uncontaminated items, the inequality pattern is a memorization artifact; if they persist, it is a genuine property of cross-country simulation capability.","tokens_in":41913,"feed_emoji":"🌍","tokens_out":12334,"duration_ms":111008,"temperature":0.7,"pith_summary":"This paper sets out to show that large language models, asked to simulate how people in different countries answer value-survey questions, perform far better for some populations than for others, and that this unevenness is systematic rather than random. Across nine open models and 59 countries, simulation accuracy is consistently higher for wealthier, more technologically advanced, better-governed, lower-power-distance, and more individualist populations, a pattern that survives across model families and scales. The paper introduces 'representational equality' as a metric distinct from average accuracy, and shows that interventions that raise average or target-group accuracy—native-language prompting, extra retrieved context, continued post-training, preference alignment—do not reliably shrink cross-country inequality; supplying retrieved survey context at inference is the one pathway that most consistently improves both accuracy and equality at once. The stakes: LLM-based opinion simulation, now proposed as a scalable proxy for cross-national social science, would quietly reproduce global structural inequalities in any downstream application that takes its outputs at face value.","feed_headline":"59-nation test finds AI simulates rich countries' values best","feed_subtitle":"Across nine open models, richer and more individualist populations get simulated more accurately.","key_machinery":"The load-bearing object is the Representational Equality Index, defined as the coefficient of variation of country-level simulation accuracies, $Eq_{CV} = \\sigma_{A_c}/\\mu_{A_c}$, selected after four fairness-style indices (max-min difference, min-max ratio, CV, Gini) were shown to rank models consistently. Underneath it sits the simulation pipeline: subpopulations are cells defined by country plus two demographic attributes (gender, age, education, income), kept when they contain at least ten WVS respondents, for 2,420 cells across 59 countries. Each cell's simulated value distribution comes from a first-token probability method, in which the model's log-probabilities over the top-20 first-position tokens are filtered to valid answer letters and normalized, with a capped residual-mass approximation for missing options; accuracy is $1 - \\mathrm{JSD}$ between that distribution and the empirical WVS distribution, aggregated with equal weight across 11 value dimensions. The diagnostic arm runs Spearman correlations of country accuracy against PEST-family macro indicators (Political: six Worldwide Governance Indicators; Economic: GDP per capita; Socio-cultural: Hofstede's six dimensions; Technological: Internet use and the Global Innovation Index), with Benjamini–Hochberg false-discovery-rate correction. A supplementary accuracy-equality composite, the geometric mean of average JSD and $Eq_{CV}$, lets the paper rank models on both axes at once.","core_discovery":"The paper's central claim is that representational equality—the evenness of simulation accuracy across populations—is a measurable, structurally patterned property of current LLMs, and that current models fail it. Using first-token probabilities to approximate each model's response distribution over answer-option letters for 2,420 country-by-demography subpopulations, and scoring accuracy as $1-\\mathrm{JSD}$ against World Values Survey Wave 7 ground truth, the authors find country-level accuracy dispersion with $Eq_{CV}$ values from 0.0286 to 0.0546 across nine models; GLM-4-9B is the most accurate (0.739) but not the most equal, while ChatGLM3-6B is the most equal. Country-level accuracy correlates positively with GDP per capita, Internet use, the Global Innovation Index, and the six Worldwide Governance Indicators, and negatively with Hofstede power distance, with individualism and indulgence positively associated. In a case study, GLM-4-9B's simulated Iraqi response distributions are more often closer to the US empirical distribution than to the Iraqi one, indicating anchoring toward better-represented countries. The paper also claims the two intervention pathways diverge: native-language prompting improves accuracy unevenly and can relocate inequality; retrieved additional information usually improves accuracy and equality together; language-specific continued post-training raises average accuracy for target languages but unevenly across countries and questions; and preference alignment via DPO or GRPO yields no systematic gains on either axis, with human-annotated preference data preserving accuracy better than AI-annotated data. An ISSP-based supplementary benchmark is offered to show the main patterns are not specific to one survey instrument.","pith_inferences":["A direct test the paper leaves implicit: if the wealth-accuracy gradient tracks training-corpus exposure, the same correlations should be predictable from country-level web-text frequencies, and equality audits could be run before any simulation by auditing pretraining data coverage.","The US-versus-Iraq anchoring suggests a general collapse of low-resource-country simulations toward the dominant global distribution; a per-question 'closeness to the best-represented country' diagnostic would make that mechanism visible in future audits.","The $Eq_{CV}$ framework transfers to other partitioning axes—language, dialect, rural/urban, within-country regions—so the same metric could audit representational equality below the country level.","Because the equality gains of additional-information retrieval likely scale with retrieved-item relevance, a testable extension is to measure how $Eq_{CV}$ changes with retrieval set size, retrieval quality, and the empirical support of retrieved subgroup distributions."],"forward_implications":["Cross-national LLM opinion simulation cannot be treated as representative of global populations unless it passes an equality audit alongside an accuracy check.","The highest-accuracy model is not the highest-ranked joint accuracy-equality model: GLM-4-9B leads on accuracy (0.739) yet trails ChatGLM3-6B and Mistral-7B-v0.3 on the composite metric.","Supplying retrieved empirical context at inference is the most reliable lever for raising accuracy and cutting cross-country inequality simultaneously.","Gains concentrated in particular language communities can relocate rather than reduce inequality, so interventions must be evaluated as distributions, not as group averages.","Preference alignment (DPO or GRPO, human- or AI-annotated) should not be expected to improve value simulation, and human-annotated preference data is the safer choice for preserving accuracy."],"supporting_citations":[{"why":"Supplies the first-token probability extraction method and the capped residual-mass approximation used to recover missing answer-option probabilities.","marker":"Santurkar et al. 2023"},{"why":"Provides the country-level simulation accuracy definition (1 − JSD against human response distributions) that the equality index aggregates.","marker":"Durmus et al. 2023"},{"why":"Supplies the accuracy-equality principle from the fairness literature that motivates measuring comparable performance across groups.","marker":"Berk et al. 2017"},{"why":"Supplies the fairness-definition taxonomy from which the max-min, min-max ratio, CV, and Gini equality indices are drawn.","marker":"Verma and Rubin 2018"},{"why":"Provides the six cultural dimensions (power distance, individualism, and the rest) used as socio-cultural correlates of accuracy.","marker":"Hofstede 2011"},{"why":"Provides the six Worldwide Governance Indicators used as political-governance correlates.","marker":"Kaufmann, Kraay, and Mastruzzi 2010"},{"why":"Documents systematic cross-country geographic bias in LLMs, motivating country as the evaluation unit.","marker":"Manvi et al. 2024"},{"why":"Shows alignment can introduce or intensify cultural biases, motivating the preference-alignment experiments.","marker":"Ryan, Held, and Yang 2024"},{"why":"Supplies the false-discovery-rate correction applied to all correlation significance tests.","marker":"Benjamini and Hochberg 1995"}],"fun_headline_variants":["AI value simulations favor richer nations, 59-country study shows","Richer countries get more accurate AI opinion simulations","Wealth gap in AI's cross-country value predictions","Study: LLMs simulate wealthy populations more precisely","AI's opinion forecasts less reliable for poorer nations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes the WVS Wave 7 responses used as ground truth have not already been memorized by the evaluated models, and because no contamination check is reported, the country-accuracy gradient could partly measure training-data exposure rather than simulation capability.","fun_headline_variants_meta":{"raw":{"variants":["AI value simulations favor richer nations, 59-country study shows","Richer countries get more accurate AI opinion simulations","Wealth gap in AI's cross-country value predictions","Study: LLMs simulate wealthy populations more precisely","AI's opinion forecasts less reliable for poorer nations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000837,"raw_usage":{"total_tokens":3727,"prompt_tokens":1101,"completion_tokens":2626,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":717,"completion_tokens_details":{"reasoning_tokens":2551}},"tokens_in":717,"tokens_out":2626,"duration_ms":19428,"temperature":1.0,"reasoning_tokens":2551,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:29:35.789778+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a contamination-controlled replication: measure n-gram or paraphrase overlap between the 160 benchmark questions and each model's training corpus, then re-score accuracy only on questions with no measurable overlap. If the GDP, Internet-use, and power-distance correlations with accuracy vanish or reverse on uncontaminated items, the inequality pattern is a memorization artifact; if they persist, it is a genuine property of cross-country simulation capability.","supporting_citations":[{"cited_title":"Online readings in psychology and culture , volume=","cited_arxiv_id":null,"evidence_quote":"Provides the six cultural dimensions (power distance, individualism, and the rest) used as socio-cultural correlates of accuracy."}],"review_version":1}