{"id":"05f2bc53-f474-4501-85b6-ba36ee3231ca","arxiv_id":"2607.03091","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Zero-shot LLMs reach 52% exact-match accuracy predicting unseen TEDS 2024 survey items from other answers, within 6 pp of a same-population random forest, with a stable construct hierarchy.","lead":"Zero-shot LLMs predict individual answers to unseen Taiwanese political survey items at 52% accuracy, within 6 points of a supervised random forest. The work offers a stricter individual-level test of silicon sampling and maps which attitudes transfer.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The 6 pp gap claim rests on an untested assumption that Set A anchors do not leak most of the predictive signal for Set B.","rationale":"The reader's weakest_assumption correctly isolates the single condition on which the strongest claim depends. The reported numbers themselves are clear and the baselines are appropriate; the soft spot is not fabrication or arithmetic error but the untested claim that the chosen anchors leave a genuine out-of-construct residual. Because the paper already flags the fixed partition as a limitation and because a simple ablation (strip anchors or re-partition without same-construct leakage) would settle the issue, the appropriate stance remains CONDITIONAL rather than REJECT or ACCEPT. No stronger internal inconsistency appears; the hierarchy and variance-collapse observations are secondary and do not underwrite the 6 pp headline. Code/prompt release would further help, but the partition-sensitivity check is the decisive one for the central claim.","tokens_in":10163,"tokens_out":575,"duration_ms":5341,"concrete_test":"Re-run the exact Table II protocol after removing all same-construct anchors from Set A (or after a random A/B re-partition that forbids same-construct items in A). If LLM exact-match falls by more than ~5–8 pp relative to the published 52% (or the LLM–RF gap widens substantially), the OOD framing and the 'within 6 pp' headline weaken; if accuracy is essentially unchanged, the concern does not land.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim (Abstract finding 1; Table II) is that zero-shot LLMs reach 52% exact-match on 'genuinely unseen' Set B items and close to within 6 pp of a same-population RF (58.3%). That claim is only as strong as the assertion that the fixed theory-driven A/B partition (§III, Table I) is a true out-of-construct test. Every Set B item is given explicit same-construct anchors in Set A (e.g., M5a sovereignty is anchored to N1a/N4a/N5a; B3 to B1/B2a; D11e to D11a,c). If those anchors already carry most of the signal that RF and the LLMs exploit, then the evaluation is residual within-construct prediction rather than the harder OOD transfer advertised, and the 6 pp gap is measured under a softer regime than claimed. The paper itself notes the partition is fixed and non-optimized (§VII) and never reports a no-anchor or cross-construct-only ablation, so the load-bearing condition remains unchecked.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces cross-survey transfer as an individual-level evaluation for silicon sampling: an LLM is conditioned on a respondent’s demographics plus answers to a theory-driven Set A of TEDS 2024 items and must predict that same respondent’s answers to a disjoint Set B of items. Using three open-weight models (27B–120B), abliterated variants, and supervised baselines (majority, logistic regression on demographics, random forest on demographics+Set A via 5-fold CV), the authors report that zero-shot LLMs reach ~52% exact-match accuracy (within 6 pp of the RF’s 58.3%), a stable construct-level predictability hierarchy (partisan attitudes ~67% down to sovereignty ~23%), and more nuanced patterns of variance collapse and alignment effects than previously claimed.","tokens_in":10516,"tokens_out":1044,"duration_ms":17909,"significance":"If the empirical claims hold under a genuinely out-of-construct regime, the work supplies a clearer diagnostic than distributional silicon-sampling evaluations and situates LLM performance relative to same-population supervised ceilings. Strengths include the non-WEIRD Mandarin political setting, open-weight models that permit abliteration, transparent item-level and construct-level tables (II–V), and the explicit comparison of variance ratios across LLMs and RF. These elements make the paper useful for both survey methodologists and the LLM social-simulation community, even if absolute accuracies remain modest.","major_comments":[{"comment":"Abstract finding 1 and the framing of “genuinely unseen / out-of-construct” transfer rest on the fixed theory-driven A/B partition (§III, Table I). Every Set B item is supplied with same-construct anchors in Set A (e.g., M5a sovereignty is anchored to N1a/N4a/N5a; B3 to B1/B2a; D11e to D11a,c). Without a no-anchor or pure cross-construct ablation, it is impossible to know how much of the 52% accuracy (and the 6 pp gap to RF) is residual within-construct leakage rather than the harder OOD transfer advertised. The Limitations section notes the partition is non-optimized but does not quantify the leakage; this is load-bearing for the central claim.","section":"§III, Table I, Abstract"},{"comment":"The prompt construction (§III) injects population-level response distributions “for base-rate calibration.” Majority vote already achieves 46.6% (Table II); supplying the same base rates to the LLM may inflate zero-shot exact-match figures relative to a pure persona-only condition. An ablation that removes the base-rate component is needed to isolate the contribution of individual Set A answers.","section":"§III (prompt components), Table II"},{"comment":"Interpretation of the 52% vs. 58% gap (and of the low-tier constructs) is limited by the absence of human test–retest reliability for TEDS items. The paper correctly flags this in §VII, yet without that ceiling it remains unclear whether the observed accuracies are near the irreducible noise floor or still far below it—especially for the 0–10 scales that are arithmetically disadvantaged under exact-match.","section":"§VII Limitations, Table III"}],"minor_comments":[{"comment":"Table III header contains the typo “CONPARISON”; §III heading is missing a space (“TRANSFERFRAMEWORK”).","section":"Table III, §III"},{"comment":"Fig. 1 is described but the visual encoding of information regimes (hatched vs. dotted) is not fully self-explanatory without the caption; a short legend inside the figure would help.","section":"Fig. 1"},{"comment":"Quantization/precision choices (Q4 vs. FP16) are listed but never ablated; a one-sentence note on whether they affect the ranking would strengthen reproducibility claims.","section":"§IV"},{"comment":"The within-±1 and MAE columns in Table II are useful; reporting them also by construct (or at least for the 0–10 items) would make the hierarchy in Table III easier to interpret.","section":"Tables II–III"}],"recommendation":"major_revision","confidential_remarks":"The core experimental design is clean and the non-WEIRD setting is a genuine plus, but the missing anchor ablation is the single issue that most directly undercuts the paper’s strongest claim. Once that (and the base-rate ablation) are supplied, the manuscript should be close to acceptance; I would not reject solely on the current framing."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is a cleaner individual-level test than most silicon-sampling work. Instead of demographics-only prompting and distributional match, they give the model a respondent’s Set A answers and ask for exact Set B answers on held-out items from the same TEDS 2024 survey. Zero-shot open-weight models hit ~52% exact match; a same-population RF with 5-fold CV hits 58.3%. That 6 pp gap is the headline number, and the tables are transparent about it.\n\nWhat is actually new: the disjoint A/B protocol with theory-driven anchors, the explicit information-asymmetry comparison to supervised baselines, the construct hierarchy (partisan items ~67%, sovereignty ~23%), and the multi-family abliteration results showing alignment effects are model-dependent rather than uniform. The non-WEIRD Mandarin setting and the variance-ratio numbers (LLMs often preserve more diversity than RF) are also real additions. Citations to Argyle, Bisbee, Santurkar, Converse, etc. look appropriate; no obvious citation games.\n\nSoft spots, in proportion. The stress-test concern is fair but not fatal: every Set B item has same-construct anchors in Set A (Table I), the partition is fixed and never ablated against a no-anchor or pure cross-construct condition, and the paper itself flags this in Limitations. So the “genuinely OOD / out-of-construct” framing is a bit stronger than the design strictly supports; residual within-construct leakage is possible. Exact-match also disadvantages the 0–10 items, and they lack a human test-retest ceiling for TEDS. Code/prompts are “available upon request,” which is a practical friction. None of these collapse the central accuracy claim under the design they actually ran; they just mean the rigor claim needs a qualifier.\n\nThis is for people who care about LLM social simulation evaluation and survey methodology. It is not a breakthrough, but it is careful enough and non-WEIRD enough to be worth a serious referee’s time. I would bring it to reading group, cite the protocol and the hierarchy numbers, and send it to peer review with a request for partition-sensitivity checks and public scripts.","headline":"Solid empirical methods paper that tightens silicon-sampling evaluation to individual-level cross-item prediction; the 6 pp gap is real under their design, but the OOD claim is softer than advertised because same-construct anchors are never ablated.","tokens_in":11105,"tokens_out":562,"would_cite":true,"duration_ms":5334,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Zero-shot LLMs predict individual survey answers on unseen questions to within 6 points of a supervised model trained on the same population.","keywords":["silicon sampling","large language models","survey simulation","cross-survey transfer","political attitudes","individual-level prediction","variance collapse","safety alignment"],"falsifier":"Re-run the identical models on a re-partitioned survey in which no Set B item shares any construct facet or anchor with Set A; if accuracy falls to majority-vote levels, the original claim of coherent cross-construct transfer is false.","tokens_in":11089,"feed_emoji":"📊","tokens_out":859,"duration_ms":12255,"temperature":0.7,"pith_summary":"Most tests of silicon sampling only check whether language models can reproduce population averages, not whether they can track a single person across different topics. This paper introduces cross-survey transfer: give a model one respondent’s answers to a first block of questions and ask it to predict that same person’s answers to a completely different block. On a national Taiwanese election survey, three open-weight models reach about 52 percent exact-match accuracy with no training examples from that population, closing most of the gap to a random forest that sees nearly a thousand labeled cases. A clear hierarchy appears: party-linked attitudes are far more predictable than personal or sovereignty items. Variance narrowing and safety-alignment effects turn out to be shared or model-specific rather than universal LLM flaws. The result shows both how far zero-shot simulation can go and where it still fails.","feed_headline":"Zero-shot LLMs hit 52% on unseen survey answers","feed_subtitle":"Within 6 points of a trained random forest, and a clear hierarchy of which attitudes transfer","key_machinery":"Cross-survey transfer: a fixed, theory-driven partition of survey items into disjoint Set A (persona context with construct anchors) and Set B (never-seen prediction targets), evaluated by individual-level exact-match accuracy rather than aggregate distributions.","core_discovery":"Zero-shot language models, conditioned only on demographics and a respondent’s answers to one set of survey items, can predict that respondent’s answers to disjoint items from the same survey at 52 percent exact-match accuracy, within six percentage points of a supervised random forest trained on same-population data. A stable construct hierarchy runs from roughly 67 percent for partisan attitudes down to 23 percent for sovereignty, and both variance collapse and alignment distortions prove less LLM-specific than commonly claimed.","pith_inferences":["The same hierarchy should appear in any multi-party democracy whose belief systems are structured by party identification.","Abliteration results imply that alignment removal is not a free lunch for political simulation and must be validated per model family.","Exact-match ceilings near 50–60 percent suggest that richer persona inputs (short interviews or longitudinal history) will be needed before LLMs can replace rather than merely augment surveys.","The sovereignty gap may flag a general limit: items that are both multi-dimensional and politically sensitive remain hard for zero-shot models even when statistical baselines succeed."],"forward_implications":["Silicon sampling can serve as a zero-shot prior for new populations when labeled survey data do not yet exist.","Researchers can rank survey constructs by predictability before fielding costly human samples.","Hybrid pipelines that blend LLM predictions with small human calibration sets become the practical next step.","Claims that variance collapse or safety alignment uniquely cripple LLMs must be re-checked against supervised baselines and across model families.","Non-English, non-WEIRD political surveys are viable test beds rather than afterthoughts."],"fun_headline_variants":["Zero-shot LLMs hit 52% predicting unseen survey answers","LLMs within 6pp of random forest on cross-survey transfer","Silicon sampling: 52% exact-match on disjoint survey items","Construct hierarchy: 67% partisan down to 23% sovereignty","Zero-shot LLMs near supervised baseline on respondent transfer"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The chosen Set A anchors already supply enough within-construct signal that the task is not a true out-of-construct test.","fun_headline_variants_meta":{"raw":{"variants":["Zero-shot LLMs hit 52% predicting unseen survey answers","LLMs within 6pp of random forest on cross-survey transfer","Silicon sampling: 52% exact-match on disjoint survey items","Construct hierarchy: 67% partisan down to 23% sovereignty","Zero-shot LLMs near supervised baseline on respondent transfer"]},"model":"grok-4.5","effort":"low","cost_usd":0.00502,"raw_usage":{"total_tokens":1434,"prompt_tokens":800,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":50200000,"prompt_tokens_details":{"text_tokens":800,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":561,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":800,"tokens_out":73,"duration_ms":4968,"temperature":1.0,"reasoning_tokens":561,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T04:57:47.797671+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the identical models on a re-partitioned survey in which no Set B item shares any construct facet or anchor with Set A; if accuracy falls to majority-vote levels, the original claim of coherent cross-construct transfer is false.","supporting_citations":[],"review_version":1}