{"id":"1c7e3701-dfc4-469d-a009-4aee72735930","arxiv_id":"2608.03044","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Base language models outperform post-trained models at generating individual survey responses that match human distributions, while post-trained models are more accurate when directly predicting population opinion distributions.","lead":"Large language models can either write fake survey answers that look like one person's, or directly guess how a whole population distributes across answer options. This paper shows these are separate skills: base models write more human-like answers, while instruction-tuned models are better at predicting overall opinion distributions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Base-vs-post-trained comparison confounded by unmatched prompt scaffolds and temperatures; the emulation claim requires matched-scaffold robustness testing.","rationale":"The reader identified the same load-bearing weakness: the base versus post-trained comparison is confounded by different prompt scaffolds and unmatched sampling temperatures. I agree this is the single most important threat to the central claim, because the paper's headline finding ('base models are stronger emulators') is a claim about model training status, not about prompt design. If the emulation advantage disappears when scaffolds are matched, the theoretical explanation in §5 (Assistant persona bottleneck) collapses, and the practical recommendation ('use base models for text generation, post-trained for distribution estimation') would need qualification. The estimation claim is also overgeneralized, as one matched pair shows the opposite direction on TVD. These issues do not invalidate the paper's useful conceptual distinction, and the emulation finding is plausible, so a conditional verdict remains appropriate. The required fix is a matched-scaffold robustness analysis, which the paper currently lacks. The reader's conditionality is therefore well-founded; I recommend no change to the verdict, but the authors must address the confound directly before the paper can be accepted.","tokens_in":10422,"tokens_out":5828,"duration_ms":50062,"concrete_test":"Re-run the emulation experiment on all six models using the identical dialogue scaffold (the base prompt in Figure 5) for both base and post-trained variants, with matched sampling temperature (e.g., set both to 1.0, and separately both to 1.5). If post-trained models still underperform base models on TVD and Wasserstein in at least 20 of 21 model-condition comparisons, the emulation claim survives; if the gap narrows or reverses, the main result is scaffold-dependent. Repeat the same cross for estimation using the partial-string scaffold (Figure 8) for all models, and report per-condition errors with bootstrap confidence intervals.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is that the paper's central comparison is not controlled for prompt scaffold, so the attribution of the emulation advantage to base-model training status (versus post-training) is not established. In §3.1 and Appendix A.2.2–A.2.3, base-model emulation uses an interviewer-participant dialogue with demographic conditioning inserted as a prior turn (Figure 5), while post-trained emulation uses a system instruction that names the attribute (Figure 7). These scaffolds differ in structure, tone, and placement of demographic information; any of these differences could drive the observed base advantage in emulation error (Table 1) and demographic correlation (Table 2). Similarly, §3.2 and Appendix A.3 give base models a partial distribution string (Figure 8) but require post-trained models to produce JSON (Figure 9), so the estimation comparison is also scaffold-dependent. Temperature is additionally unmatched: post-trained emulation samples at 1.5 (§4.1), while base temperature is not reported. The causal account in §5 attributes the difference to an 'Assistant persona bottleneck' in post-training, but this is only persuasive if the prompt confound is ruled out. The paper never tests whether post-trained models under the base dialogue scaffold still lose to base models, or whether base models under the post-trained system scaffold gain ground. As a secondary point, Table 1 shows one matched pair (Olmo-3-1025-7B vs. Olmo-3-7B-Instruct) where the base model achieves lower TVD on estimation (0.278 vs. 0.285), directly contradicting the unqualified claim that post-trained models are better estimators.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper distinguishes two tasks for LLM-based opinion simulation: emulation (generating individual responses that aggregate into a population distribution) and estimation (directly predicting the population distribution). Using three matched base/post-trained model pairs on 59 Pew American Trends Panel items, the authors report that base models are stronger emulators on both total variation and Wasserstein distance and better preserve demographic structure, while post-trained models are stronger estimators when asked explicitly for distributions. The authors propose that downstream task form should determine which model type to use, and they interpret the tradeoff through an 'Assistant persona bottleneck' in post-training.","tokens_in":10676,"tokens_out":4578,"duration_ms":42356,"significance":"The emulation–estimation distinction is a useful organizing principle and may help reconcile conflicting results in the LLM opinion-simulation literature. The paper has clear strengths: it evaluates frozen models against external Pew ground truth, fits no parameters, validates the LLM judge against human annotation, includes a first-token extraction corroboration in Appendix A.6, and makes explicit limitations. If the central claims hold, the paper offers actionable model-selection guidance. However, the current evidence is weakened by unmatched prompt scaffolds and sampling procedures across base and post-trained conditions, and by at least one pair-level reversal in the estimation result. The core claims are defensible but need additional robustness experiments before they can be considered established.","major_comments":[{"comment":"The central emulation comparison is confounded by prompt scaffold: base models are prompted with an interviewer-participant dialogue and demographic conditioning inserted as a prior turn (Figure 5), while post-trained models receive a system instruction that names the attribute (Figure 7). The causal account in §5 attributes the emulation gap to post-training, but any of the scaffold differences could drive the observed base advantage. In addition, post-trained emulation samples at temperature 1.5 while the base-model temperature is not reported, and temperature directly affects the output-diversity measures in Table 4. Please report the base sampling temperature and add matched-scaffold robustness tests, e.g., post-trained models under the base dialogue scaffold and base models under the post-trained system scaffold, with temperatures held fixed.","section":"§3.1, Appendices A.2.2–A.2.3, §4.1"},{"comment":"The estimation comparison is similarly scaffold- and procedure-dependent: base models complete a partial distribution string (Figure 8) while post-trained models must return a JSON object with integer percentages summing to 100 (Figure 9), and base estimates average multiple completions while post-trained estimates use a single greedy decode. The claim that post-trained models are stronger estimators is therefore not yet isolated from output-format and sampling effects. Please test both model types under a common output format and sampling procedure, for example by prompting post-trained models with the partial-string scaffold and base models with JSON instructions.","section":"§3.2, Appendices A.3.1–A.3.2"},{"comment":"The estimation claim is not consistent for one of the three pairs: Olmo-3-7B-Instruct has higher mean TVD (0.285) than its base counterpart Olmo-3-1025-7B (0.278), and Table 6 shows the reversal is concentrated in the Very Liberal and Very Conservative conditions (0.387 vs. 0.365 and 0.370 vs. 0.268). The abstract and §4.3 state that post-trained models are stronger estimators without this qualification. Please either weaken the claim to an average effect and report per-pair consistency, or provide an analysis explaining the reversal.","section":"Table 1 and Table 6"},{"comment":"For Olmo-3-7B-Instruct, the demographic-structure correlation is not statistically significant (ρ = 0.350, p = 0.201), yet the text groups this model with others under 'weaker structural alignment' and uses the aggregate interpretation to support the emulation claim. Because demographic structure preservation is part of the central emulation finding, this non-significant result should be explicitly acknowledged and the strength of the corresponding claim moderated.","section":"Table 2 and §4.2"}],"minor_comments":[{"comment":"The post-trained rows for Qwen3-14B and Olmo-3-1125-32B appear to be missing the 'Instruct' label, making the table difficult to parse; please add consistent labels for all rows.","section":"Table 1"},{"comment":"The phrase 'across all six models' is ambiguous because there are three matched pairs of model variants; please say 'across all six model variants (three matched pairs)'.","section":"§4.2 and Appendix A.1"},{"comment":"The human annotator is reported as a member of the research team; please state whether the LLM judge and the annotator were blinded to model type and whether any independent annotators were used, since the judge validation underlies the emulation pipeline.","section":"Appendix A.2.1"},{"comment":"Figure 11 is described as spanning '8B to 14B scale' but only Qwen3-14B appears in the main experiments; please clarify which additional Qwen sizes were evaluated and where the results are available.","section":"Appendix A.6"}],"recommendation":"major_revision","confidential_remarks":"The scaffold and temperature confounds in the base-vs-post-trained comparison are real and load-bearing, and the estimation-claim reversal for the Olmo-3 pair needs explicit treatment. I would not reject the paper: the distinction is valuable and the authors can address the concerns with matched-scaffold experiments. I would recommend requiring those experiments before publication rather than accepting the current version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Best to know: this paper gives a clean name to a real distinction — generating individual responses (emulation) vs directly predicting the population distribution (estimation) — and shows, across three matched model families, that base models are more faithful emulators while post-trained models are better estimators. That is a genuinely useful framing, and the open-response protocol is a sensible way to dodge positional bias. The paper is honest about its limits and does not overclaim on the mechanism.\n\nWhat is solid: six models, three matched base/post-trained pairs, 59 Pew items, seven demographic conditions. The emulation advantage is consistent in every condition, and the demographic-correlation results (base models compress inter-group differences to roughly 70% of human magnitude while post-trained models exaggerate by about 2x) are striking. No free parameters are fitted; evaluation is against external ground truth. The first-token appendix corroborates the direction. The Persona Selection Model is presented as interpretation, not evidence — fine.\n\nWhere it wobbles: the central comparison is not controlled for prompt scaffold. Base emulation uses an interviewer-participant dialogue with demographics in a prior turn; post-trained gets a system instruction naming the attribute. Estimation uses a partial string for base and JSON for post-trained. Temperature is unmatched (post-trained 1.5, base unreported). Any of these could drive part of the gap. The paper never tests the crossed conditions. Given that the entire causal story in Section 5 rests on an 'assistant persona bottleneck,' that is a real gap. Also, Table 1 shows the Olmo-7B base beats its instruct variant on estimation TVD (0.278 vs 0.285), so the estimation claim should be softened. The LLM judge's kappa of 0.66 against a single in-team annotator is moderate; a second annotator or a disclosed judge model would help.\n\nNone of this kills the paper. The emulation finding is likely to survive matched scaffolds — the self-similarity numbers (8–14x more collapse in post-trained outputs) are independent of scaffold. But as written, the attribution to post-training per se is not established.\n\nWho benefits: anyone picking a model for synthetic respondents or for survey-distribution prediction. It deserves a serious referee, but the revision should add crossed-scaffold robustness checks, report base-model temperature, disclose the judge model, and qualify the estimation exception.","headline":"Useful emulation/estimation distinction, but the base-vs-post-trained comparison needs matched scaffolds before the causal story holds.","tokens_in":11248,"tokens_out":2036,"would_cite":true,"duration_ms":17963,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Base LLMs emulate people; tuned LLMs estimate them.","keywords":["opinion simulation","emulation vs estimation","base models","post-trained models","demographic conditioning","distributional alignment","persona collapse","open-response generation"],"falsifier":"Re-run the emulation evaluation with both model types under identical prompt scaffolds and identical sampling temperatures (for example, giving post-trained models the interviewer-participant dialogue with demographic conditioning as a preceding turn). If a post-trained model then matches or beats its base counterpart on Wasserstein distance to human ground truth, the paper's central claim of a base-model emulation advantage is falsified.","tokens_in":10181,"feed_emoji":"🗳️","tokens_out":7986,"duration_ms":68077,"temperature":0.7,"pith_summary":"The paper argues that conflicting results in LLM opinion simulation come from conflating two distinct tasks. In emulation, a model generates individual survey answers and the population distribution emerges by aggregation; in estimation, the model directly predicts the distribution. Across three matched base/post-trained model pairs on the Pew American Trends Panel, base models are stronger emulators—closer to human ground truth and better at preserving demographic structure—while post-trained models are stronger estimators. This suggests that simulation users should choose the substrate according to the output form required: base models for generating text, post-trained models for predicting distributions.","feed_headline":"Base models emulate; post-trained models estimate","feed_subtitle":"Why past studies clashed: generating text and predicting distributions favor different model types.","key_machinery":"The central object is the emulation–estimation distinction itself, operationalized through two evaluation pipelines: open-response generation with an LLM judge mapping free text to answer categories (to eliminate positional bias), and verbalized-distribution estimation in which the model outputs a JSON probability vector. The load-bearing comparison runs across matched base and post-trained model pairs, using total variation distance and Wasserstein distance as complementary fidelity metrics, plus Spearman correlation of model and human pairwise inter-group distances to test demographic structure. The proposed explanatory mechanism is the 'assistant persona bottleneck': post-training steers the base distribution into a single helpful-assistant persona that is good at direct estimation but compresses the diversity needed for faithful emulation.","core_discovery":"On a benchmark of 59 four-option Pew survey items across seven demographic conditions, the authors find that base (pre-instruction-tuning) models generate individual responses whose aggregated distributions are closer to human ground truth on both total variation and Wasserstein distance in every model-condition comparison, and they better track the structure of demographic differences (Spearman rho 0.61–0.75 vs. 0.35–0.59 for post-trained models). Post-trained models give more accurate direct estimates of population distributions, with error decreasing at larger scale and the frontier reference model performing best. The authors interpret this as a tradeoff introduced by post-training: instruction tuning compresses output diversity (post-trained responses are 8.0–14.0 times more self-similar at the bigram level), so the model plays a single assistant persona simulating a demographic persona, whereas base models draw on the pre-training persona distribution directly. The practical conclusion is that emulation and estimation should be evaluated separately, and model choice should follow whether the task needs generated text or direct distribution prediction.","pith_inferences":["Editorial inference: the emulation–estimation split may generalize beyond opinion surveys to any demographic-conditioned generation task, such as persona-driven dialogue, creative writing, or agent simulation, where base models' broader output distribution could better preserve population diversity than instruct-tuned models.","Editorial inference: a testable extension would fine-tune a post-trained model to restore output diversity (e.g., via diversity-promoting training or multi-persona decoding) and check whether its estimation advantage survives; if the tradeoff is controlled by a single capability knob, it could be tuned rather than accepted.","Editorial inference: because frontier base checkpoints are not released, the paper's scale trend for estimation cannot currently be tested for emulation at the largest closed models; open-weight frontier base models would provide the decisive data.","Editorial inference: the Spearman results suggest emulation quality is not only about aggregate closeness but about the geometry of inter-group differences; a practical extension is to use the pairwise distance ratio (compression vs. exaggeration) as a design diagnostic for synthetic populations."],"forward_implications":["Model selection for human simulation becomes task-driven: use base models when the pipeline needs generated text (synthetic respondents, interactive agents), and post-trained models when it needs direct distributional estimates (polling-style predictions).","Conflicting prior results become explainable as a task mismatch: studies reporting good alignment typically evaluated base-model sampling, while studies reporting persona collapse and demographic insensitivity mostly tested post-trained models.","Emulation evaluations should use open-response generation with judge-based mapping to avoid the positional bias that confounds first-token probability extraction; the paper's first-token results corroborate the main finding.","Post-training is not uniformly better for simulation: it trades estimation accuracy against the output diversity required for faithful emulation, a tradeoff that should be measured and reported rather than assumed away.","If the claim holds, instruction-tuned models' exaggeration of inter-group differences (about 2x) relative to humans, versus base models' compression to about 70%, bears directly on any deployment of synthetic respondents for policy or market research."],"supporting_citations":[{"why":"Establishes the base-model distributional-alignment finding the paper extends and supplies the distributional-alignment evaluation convention.","marker":"Santurkar et al. (2023)"},{"why":"Introduces verbalized distribution estimation, the method that defines the estimation paradigm and the finding that models estimate better than they sample.","marker":"Meister et al. (2025)"},{"why":"Documents positional bias in first-token probabilities, motivating the open-response emulation pipeline.","marker":"Wang et al. (2024)"},{"why":"Shows ordering and label sensitivity in LLM survey answers, reinforcing the need for open-response generation.","marker":"Tjuatja et al. (2024)"},{"why":"Documents persona collapse in post-trained models, the phenomenon the paper's emulation results quantify.","marker":"Li et al. (2025)"},{"why":"Provides the Persona Selection Model used as the theoretical account of the emulation–estimation tradeoff.","marker":"Marks et al. (2026)"},{"why":"Provides the assistant-persona account of post-training that explains the diversity bottleneck.","marker":"Lu et al. (2026)"},{"why":"Supplies the American Trends Panel Wave 54 dataset and ground-truth human distributions for the evaluation.","marker":"Pew Research Center (2020)"}],"fun_headline_variants":["Task split resolves AI opinion simulation conflict","Base models mimic people, tuned models predict polls","Emulation with base, estimation with post-trained","Why LLM surveys clash: emulation vs estimation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the measured differences between base and post-trained models come from training status rather than from the different prompt scaffolds (dialogue vs. system-instruction) and sampling temperatures used for each model type; if post-trained models were re-run under the base models' dialogue scaffold, the emulation ranking could change.","fun_headline_variants_meta":{"raw":{"variants":["Task split resolves AI opinion simulation conflict","Base models mimic people, tuned models predict polls","Emulation with base, estimation with post-trained","Why LLM surveys clash: emulation vs estimation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000326,"raw_usage":{"total_tokens":1801,"prompt_tokens":894,"completion_tokens":907,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":849}},"tokens_in":510,"tokens_out":907,"duration_ms":9546,"temperature":1.0,"reasoning_tokens":849,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:58:36.726992+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the emulation evaluation with both model types under identical prompt scaffolds and identical sampling temperatures (for example, giving post-trained models the interviewer-participant dialogue with demographic conditioning as a preceding turn). If a post-trained model then matches or beats its base counterpart on Wasserstein distance to human ground truth, the paper's central claim of a base-model emulation advantage is falsified.","supporting_citations":[{"cited_title":"2026 , month = feb, day =","cited_arxiv_id":null,"evidence_quote":"Provides the Persona Selection Model used as the theoretical account of the emulation–estimation tradeoff."}],"review_version":1}