{"id":"9967acda-10e9-403a-941b-e518580a77a3","arxiv_id":"2605.27401","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"Zero-shot LLM-generated survey data captures broad state-level health contrasts but shows variable-dependent errors that propagate unevenly through IPF-based population synthesis.","lead":"This paper tests whether zero-shot LLM-generated health survey data can replace real surveys in population synthesis pipelines. It finds LLMs capture broad geographic contrasts but fail on complex variables like insurance, making them a supplement, not a replacement.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Memorization of BRFSS structure and state-level statistics in LLM pretraining data is the most load-bearing untested assumption; the prompt explicitly references BRFSS by name, making it impossible to distinguish geographic reasoning from retrieval of seen distributions.","rationale":"The reader correctly identified the most load-bearing concern: the memorization of BRFSS data in LLM pretraining. This is the right concern because the entire evaluation framework uses BRFSS as both the generation target and the validation benchmark, and the prompt explicitly names BRFSS. The reader's verdict of CONDITIONAL with MODERATE confidence is appropriate. The paper is honest about its limitations and does not overclaim—it explicitly states the data is 'not yet a replacement for real survey data.' The empirical work is sound for what it tests, and the spatial validation against ACS and CDC PLACES provides some independent evidence. However, the generalizability claim remains untested, and the joint distribution gap is a real methodological limitation given the paper's own framing. The verdict should remain CONDITIONAL: the results are promising but preliminary, and the memorization concern must be addressed before stronger claims can be made. I note the joint distribution concern as a secondary issue that the paper itself acknowledges in its limitations section, so it does not independently warrant a verdict change but should be flagged for future work.","tokens_in":12528,"tokens_out":2974,"duration_ms":69069,"concrete_test":"Run the same generation pipeline with a modified prompt that (a) removes all mention of BRFSS by name, asking only for 'a health survey representative of [state] adults' with the same 14 variables, and (b) includes a third state with less prominent health-disparity discussion (e.g., Nebraska or Iowa). Compare JS divergence and spatial correlation against the current results. If performance drops substantially in condition (a) or for the less-prominent state, memorization is the likely mechanism and the generalizability claim weakens. Additionally, compute JS divergence on the bivariate joint distribution of at least one variable pair (e.g., income × general health) to directly test whether multivariate relationships are preserved, not just marginals.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that zero-shot LLM-generated data captures meaningful geographic contrasts—rests on the assumption that the models are reasoning about state-level health context rather than reproducing memorized BRFSS statistics. This is undermined by the prompt design itself (Section 2.1): the LLMs are explicitly instructed to 'use their knowledge of the Behavioral Risk Factor Surveillance System Survey.' Since BRFSS is a widely disseminated public dataset with published state-level tables, and both GPT-4.1 and Gemini-2.5-Pro have training cutoffs that likely include 2023 BRFSS data (released in 2024), the models may be retrieving known marginal distributions rather than synthesizing them from general geographic knowledge. The paper acknowledges this risk in Section 4 but does not test it. If memorization is the mechanism, the approach will not generalize to less-prominent surveys, smaller geographies, or novel question instruments—the exact use cases where 'supplementary data' would be most valuable. The strongest evidence against pure memorization is the variable-dependent accuracy (some variables like insurance are poorly reproduced), but this could also reflect incomplete memorization rather than genuine reasoning. A secondary concern: the evaluation checks only marginal distributions per variable (JS divergence in Tables 1-3), not the joint distributions that IPF preserves and that the paper's own framing (Section 1) identifies as the real challenge. Incorrect joint distributions could produce incorrect synthetic populations even when marginals appear accurate, and the two-variable spatial validation provides only indirect evidence on this point.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This manuscript evaluates whether zero-shot LLM-generated health survey data (using GPT-4.1 and Gemini-2.5-Pro) can serve as input to an IPF-based geographically explicit population synthesis workflow. The authors generate synthetic BRFSS-like survey records for Colorado and Mississippi, use them in an IPF pipeline fitted to ACS marginal controls, and evaluate the resulting census tract-level synthetic populations against external benchmarks (ACS for insurance, CDC PLACES for general health). The evaluation uses JS divergence for marginal distributions and Pearson correlation for spatial agreement. The authors find that LLMs capture broad state-level contrasts but performance is strongly variable-dependent, with downstream IPF effects being mixed (sometimes amplifying, sometimes reducing errors). The paper is transparent about limitations, including cases where LLM data fails entirely (e.g., no uninsured individuals generated by GPT).","tokens_in":12747,"tokens_out":1168,"duration_ms":95125,"significance":"The paper addresses a practically important question at the intersection of LLM-generated synthetic data and population synthesis. The contribution is well-scoped: rather than evaluating LLM-generated records in isolation, the authors evaluate them as inputs to an established synthesis pipeline, which is the context where joint distributions matter. The use of external benchmarks (ACS, CDC PLACES) for spatial validation is a strength, as is the transparent reporting of variable-dependent failures. The code and prompts are publicly available (OSF DOI), supporting reproducibility. The two-state design (CO vs. MS) provides a meaningful geographic contrast. The finding that IPF partially regularizes but does not fix LLM-generated data errors is useful for the community.","major_comments":[{"comment":"Section 2.1: The prompt explicitly instructs LLMs to 'use their knowledge of the Behavioral Risk Factor Surveillance System Survey.' Since BRFSS is a widely disseminated public dataset with published state-level tables, and both models have training cutoffs that likely include 2023 BRFSS data, the evaluation may conflate memorization of seen distributions with genuine geographic reasoning. The authors acknowledge this risk in Section 4 but do not test it. This is load-bearing because the paper's claim that zero-shot generation 'produces geographically differentiated survey data' (Abstract) could be an artifact of the models reproducing memorized state-level marginals. A concrete test would be to run the same pipeline on a less-prominent survey instrument or a geography with less public data, or to compare results with and without the BRFSS name in the prompt. Without such a test, the 'ge","section":null}],"minor_comments":[{"comment":"Table 1: The row mean for 'Insurance' (0.129) is reported as the highest divergence, but the text on the same page says 'insurance has the maintains the lowest mean divergence of 0.070' — this appears to be a typo conflating Table 1 and Table 2 values. Please clarify.","section":null},{"comment":"Section 3.2.1, paragraph discussing Table 3: 'insurance has the maintains the lowest mean divergence of 0.070' is grammatically broken. Also, 0.070 is the highest row mean in Table 2, not the lowest, so the claim appears incorrect.","section":null},{"comment":"Figure 1 caption: 'A negative residual indicates that the LLM overestimates the category relative to the ground truth data and a positive residual indicates that the LLM underestimates.' This sign convention is non-intuitive (negative = overestimate). Consider clarifying or reversing the sign for reader intuition.","section":null},{"comment":"Section 2.1: The batch size of 75 was selected based on experiments testing sizes of 50, 75, 100, 150, and 200, but no quantitative results from these experiments are reported. A brief table or sentence summarizing the trade-offs would strengthen the justification.","section":null},{"comment":"Section 2.3: The JS divergence is defined with values 'approaching 1' for maximum dissimilarity, but JS divergence using log base 2 has an upper bound of 1 only for distributions over 2 categories. For 14-category variables, the maximum is still 1 (since JS is bounded by log(2) = 1 in base 2), but this should be stated explicitly for clarity.","section":null},{"comment":"Table 2: Several BRFSS-based population divergences are non-zero (e.g., Education 0.038 for CO BRFSS, Flu Vaccination 0.041 for CO BRFSS). The text explains these arise from the fitting and expansion process, but a brief note in the table caption would help readers interpret why the 'ground truth' reference is not zero.","section":null},{"comment":"Section 4: The phrase 'the BRFSS is a prominent public survey, so some of its structure is likely reflected in LLM pretraining data' is an important limitation. Consider elevating this to a more prominent position (e.g., in the Introduction or Methods) rather than burying it in the Discussion, as it affects interpretation of all results.","section":null}],"recommendation":"major_revision","confidential_remarks":"The memorization concern is the most significant issue. The prompt design (naming BRFSS explicitly) makes it impossible to distinguish geographic reasoning from retrieval of memorized distributions. This is fixable within the manuscript's scope: the authors could add a control experiment (e.g., removing the BRFSS name from the prompt, or testing on a less-prominent survey) or, at minimum, reframe the claims to acknowledge that the mechanism (memorization vs. reasoning) cannot be distinguished with the current design. The paper is otherwise methodologically sound and the topic fits the journal's scope well."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive review. The referee raises one major concern about the potential confounding of memorization with genuine geographic reasoning. We agree this is an important issue and outline below how we will address it.","responses":[{"response":"We agree this is the most important limitation of the study, and we appreciate the referee framing it so precisely. The concern is valid: because BRFSS is a prominent public dataset with widely available state-level tables, the models may be reproducing memorized marginals rather than performing genuine geographic reasoning. We already flag this in Section 4, but the referee is right that acknowledging a limitation is not the same as testing it. We will take two concrete steps in revision. First, we will run an ablation in which the BRFSS name is removed from the prompt—replacing the instruction to 'use knowledge of BRFSS' with a neutral instruction to generate realistic health survey responses for the specified state population. This directly tests whether naming the instrument is driving the results. Second, we will soften the abstract claim from 'produces geographically differentiated survey data' to language that does not presuppose the mechanism—e.g., 'produces survey data that differs across states'—and we will add a sentence in the abstract noting that the role of memorization cannot be ruled out with the current design. We note honestly that we cannot fully resolve the memorization question within this paper. The ablation will provide evidence about whether naming BRFSS matters, but even without the name, the models may have internalized BRFSS distributions from pretraining. A definitive test would require a less-prominent survey instrument or a geography with minimal public data, which we identify as a priority for future work and will state explicitly. We believe the ablation and the revised framing meaningfully address the referee's concern without overclaiming what the current study can establish.","revision_made":"partial","referee_comment":"Section 2.1: The prompt explicitly instructs LLMs to 'use their knowledge of the Behavioral Risk Factor Surveillance System Survey.' Since BRFSS is a widely disseminated public dataset with published state-level tables, and both models have training cutoffs that likely include 2023 BRFSS data, the evaluation may conflate memorization of seen distributions with genuine geographic reasoning. The authors acknowledge this risk in Section 4 but do not test it. This is load-bearing because the paper's claim that zero-shot generation 'produces geographically differentiated survey data' (Abstract) could be an artifact of the models reproducing memorized state-level marginals. A concrete test would be to run the same pipeline on a less-prominent survey instrument or a geography with less public data, or to compare results with and without the BRFSS name in the prompt."}],"tokens_in":12105,"tokens_out":1043,"duration_ms":19829,"standing_objections":["We cannot definitively distinguish memorization from genuine geographic reasoning, even with the proposed ablation, because both models may have internalized BRFSS distributions during pretraining regardless of whether the instrument is named in the prompt. A fully convincing test would require a survey instrument or geography not represented in the training data, which is beyond the scope of the current revision."]},"desk_editor":{"model":"glm-5.2","letter":"This paper asks a practical question: can zero-shot LLM-generated survey data serve as input to IPF-based population synthesis? The answer is a qualified yes—some variables reproduce well, others don't, and IPF has a complicated relationship with upstream error. That's a useful, honest finding for people building synthetic populations in data-sparse settings. The code and prompts are public, which matters here. The evaluation uses external benchmarks (ACS, CDC PLACES) for spatial validation rather than relying solely on internal consistency, and the authors transparently report failures—GPT generating zero uninsured individuals is a good example of the kind of honesty that makes the results credible. The finding that IPF sometimes amplifies and sometimes reduces errors in the generated data is genuinely useful and not something you'd get from evaluating the LLM outputs alone. The two-state contrast (Colorado vs. Mississippi) is a reasonable design choice for showing whether the models produce geographically differentiated data. What's new here is specifically the downstream evaluation: prior work looked at LLM-generated survey data in isolation, but this paper traces what happens when you feed it through a standard synthesis pipeline. That's a real contribution. The main soft spot is the memorization concern, and it's real. The prompt explicitly tells the LLMs to use their knowledge of BRFSS, which is a widely disseminated public dataset. So when the models reproduce state-level contrasts, we can't tell whether they're reasoning about geographic context or retrieving memorized marginals. The authors flag this in Section 4, which is good, but they don't test it. A replication with a less prominent survey or a novel geography would go a long way. The variable-dependent accuracy (insurance is poorly reproduced) offers weak evidence against pure memorization, but as the stress-test notes, incomplete memorization could produce the same pattern. A secondary concern: the evaluation focuses on marginal distributions per variable, but the paper's own framing identifies joint distributions as the real challenge for IPF. The spatial validation provides indirect evidence on this, but a direct check on bivariate or higher-order joint distributions would strengthen the claims substantially. The two-state, two-model scope limits generalizability, though the authors are upfront about this. This is a well-executed empirical study that fills a genuine gap. It's not a breakthrough, but it's the kind of careful, honest work that practitioners in spatial population synthesis will read and use. It deserves a serious referee who can push the authors to address the memorization confound directly—ideally with a less prominent survey—and to add joint distribution checks.","headline":"Solid empirical study with an honest reporting style; the main soft spot is the memorization confound, which the authors acknowledge but don't test.","tokens_in":13301,"tokens_out":1060,"would_cite":false,"duration_ms":24702,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"LLM-Generated Health Surveys Capture State Contrasts","keywords":["LLM-generated survey data","population synthesis","iterative proportional fitting","BRFSS","geographically explicit","synthetic populations","zero-shot generation","spatial validation"],"falsifier":"If LLM-generated survey data for a survey instrument and geography not present in pretraining data showed the same level of state-level contrast reproduction, the geographic reasoning claim would be strengthened; if performance collapsed, the memorization concern would be confirmed.","tokens_in":12649,"feed_emoji":"🗺️","tokens_out":998,"duration_ms":1884691,"temperature":0.7,"pith_summary":"This paper tests whether zero-shot LLM-generated health survey responses can replace real survey data as input to iterative proportional fitting (IPF), a standard method for building geographically explicit synthetic populations. The authors prompt GPT-4.1 and Gemini-2.5-Pro to generate synthetic BRFSS-like survey records for Colorado and Mississippi, feed those records into an IPF pipeline, and validate the resulting census tract-level populations against external benchmarks (ACS and CDC PLACES). The central finding is that zero-shot LLM-generated survey data captures broad state-level health contrasts and sometimes produces spatial patterns that correlate reasonably well with ground truth, but performance is highly variable-dependent: some variables are reproduced almost perfectly while others diverge substantially. A key mechanism finding is that IPF does not simply propagate upstream errors in a predictable direction—it sometimes amplifies them, sometimes reduces them, and occasionally an LLM-based population outperforms a real-survey-based one on specific variables. The paper concludes that zero-shot LLM-generated survey data is a promising supplementary input for population synthesis when real survey data is unavailable, but not yet a drop-in replacement.","feed_headline":"LLM-Generated Health Surveys Capture State Contrasts, Not Ready to Replace Real Data","feed_subtitle":"Zero-shot LLM survey data shows promise as supplementary input for population synthesis but variable-dependent accuracy and mixed downstream","key_machinery":"The central machinery is the IPF (iterative proportional fitting) pipeline: LLM-generated individual survey records serve as the joint-distributional template, which IPF then reweights to match census tract-level demographic marginals from the American Community Survey. The evaluation uses Jensen-Shannon divergence to measure distributional similarity and Pearson correlation against external benchmarks (ACS for insurance, CDC PLACES for general health) to assess spatial accuracy.","core_discovery":"The paper establishes two findings. First, zero-shot LLMs can generate state-conditioned health survey data that reproduces broad geographic contrasts (e.g., Colorado being healthier than Mississippi), but accuracy varies dramatically by variable: sex and age are near-perfect while health insurance and income diverge significantly. Second, the relationship between survey-data accuracy and downstream synthetic-population accuracy is not monotonic—IPF partially regularizes differences between LLM-generated datasets, sometimes reducing divergence from ground truth and sometimes worsening it, meaning that evaluating LLM-generated survey data in isolation gives an incomplete picture of its real合成","pith_inferences":["If the BRFSS survey structure and state-level health statistics are present in LLM pretraining data, the apparent geographic differentiation may partly reflect memorization rather than reasoning, which would limit generalizability to less-prominent surveys or geographies not well-represented in training corpora.","The variable-dependent accuracy pattern may correlate with category cardinality and rarity: variables with many categories or rare outcomes (insurance, heart disease) are harder for LLMs to reproduce, suggesting that few-shot or constrained generation could disproportionately improve the weakest variables.","The occasional outperformance of LLM-based populations over BRFSS-based ones for specific variables (education, flu vaccination) may indicate that LLMs smooth noisy sampling distributions, which could be a feature rather than a bug for small-sample survey contexts."],"forward_implications":["Population synthesis practitioners in data-sparse regions or domains could use LLM-generated survey data as a provisional input when real surveys are unavailable, provided they validate the specific variables of interest downstream rather than at the generation stage alone.","Variable-level validation is essential: some health variables (insurance, income, heart disease) are systematically harder for LLMs to reproduce and may require targeted correction or constrained generation rather than zero-shot prompting.","The finding that IPF can either amplify or reduce upstream errors suggests that the choice of synthesis method interacts non-trivially with input data quality, and future work should characterize which method-error combinations are self-correcting versus error-amplifying."],"fun_headline_variants":["Zero-shot LLM surveys capture state contrasts but accuracy varies by variable","LLM-generated health surveys map state differences, but miss key variables","Synthetic LLM surveys show geographic promise, but variable accuracy limits use","IPF synthesis mixes results for zero-shot LLM-generated health surveys","LLM surveys reproduce state health contrasts, but not ready to replace real data"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The paper assumes that the LLMs are generating plausible synthetic survey data from general geographic knowledge rather than regurgitating memorized distributions from the BRFSS survey, which is a prominent public dataset likely present in their pretraining data. If the models are recalling seen examples rather than reasoning about state-level context, the generalizability to less-prominent surveys or geographies would be undermined.","fun_headline_variants_meta":{"raw":{"variants":["Zero-shot LLM surveys capture state contrasts but accuracy varies by variable","LLM-generated health surveys map state differences, but miss key variables","Synthetic LLM surveys show geographic promise, but variable accuracy limits use","IPF synthesis mixes results for zero-shot LLM-generated health surveys","LLM surveys reproduce state health contrasts, but not ready to replace real data","Zero-shot LLM survey accuracy varies by variable in population synthesis"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1110,"prompt_tokens":540,"completion_tokens":570,"prompt_tokens_details":null},"tokens_in":540,"tokens_out":570,"duration_ms":10259,"temperature":1.0,"reasoning_tokens":542,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-04T20:02:48.937621+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If LLM-generated survey data for a survey instrument and geography not present in pretraining data showed the same level of state-level contrast reproduction, the geographic reasoning claim would be strengthened; if performance collapsed, the memorization concern would be confirmed.","supporting_citations":[],"review_version":1}