{"id":"166f5a3c-15db-45a9-9628-0b625ae1acf8","arxiv_id":"2505.16164","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Across 34 models and 45 configurations, no LLM individually or in ensemble reproduced the lexical diversity or retrieval structure of human phonemic fluency, with the best model producing fewer than half the unique words.","lead":"This study compared 34 large language models to 106 people on a one-minute word fluency task and found no model matched the variety of human responses. Even combining models did not help, because different models draw on the same common word pools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Count-constrained prompting is circular: models are told each participant's exact response count, and that same count is used to select configurations, so the observed rigidity may be an artifact of the constraint rather than an intrinsic LLM limitation.","rationale":"The reader's weakest_assumption is exactly the count-constraint circularity, and I agree it is the most load-bearing concern. The paper's broad conclusion ('none reproduced the scope of human variability') is used to make a general claim about LLM limitations, but the experimental protocol gives models the exact number of responses to produce and then filters on fidelity to that number. This conflates 'LLMs cannot simulate variability' with 'LLMs that are forced to produce a pre-specified number of typical responses do not simulate variability.' The excluded overproducing models are never analyzed for diversity, so the claim remains untested for unconstrained generation. The proposed test is direct: remove the count from the prompt and compare diversity at matched token counts. If diversity recovers, the paper's central claim—and its practical recommendation against using LLMs as human substitutes—is overstated. I keep the reader's CONDITIONAL verdict because the concern is serious but empirical: the test could confirm the paper's conclusion despite the flawed design. The paper's transparency about the prompt and retention criteria makes this test feasible, and the manuscript's own Limitations section acknowledges that other prompting strategies may yield different patterns, which further supports the need for this check.","tokens_in":14029,"tokens_out":8447,"duration_ms":73846,"concrete_test":"For Claude 3.7 Sonnet, o3-mini, GPT-5.2 (High Reasoning), and GPT-4 Turbo (an overproducer), run 106 API calls per model using the original participant age/education but with the 'Number of correct responses' line removed from the prompt (Appendix A, Figure A1). Compute types, TTR, ITTTR, and Zipf α on (a) the full unconstrained outputs and (b) token-matched subsamples: for each participant, randomly sample the same number of tokens as that participant's human count, repeated 1,000 times. Compare these distributions to the human values (476 types, TTR 0.27, ITTTR 0.42, α=0.89) and to the count-constrained results in Table 1. Also rerun the ensemble procedure in Section 6.1 pooling unconstrained outputs from all models that produce them.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—no LLM, individually or in ensemble, reproduces human behavioral variability—rests on a circular design. The prompt (Appendix A, Fig A1) includes the participant's exact number of correct responses ('Number of correct responses: {num_correct}'). Section 3.1 then uses MAE against that same number (MAE≤1.69) as the criterion for 'successful' simulation of production rates, and retains for all diversity analyses only the 21 configurations with no outlier where |LLM−Human|>5. The 12 configurations that overproduced (e.g., GPT-4 Turbo, MAE=57.17; o3, MAE=35.39) are excluded from the participant-level and item-level analyses, so the conclusion that 'no LLM reproduces the scope of human variability' is only demonstrated for models that were explicitly told the target count and complied. The count instruction likely truncates generation at ~17 words—exactly the region where LLM outputs are most stereotyped—and the retention filter then guarantees every analyzed output is count-matched. Appendix A states that without the count, models overproduce in pilot tests; but overproduction may reflect deeper lexical exploration that could yield more unique and idiosyncratic types. The paper also does not report sampling parameters (temperature, top-p, or reasoning effort details beyond named modes), so the role of stochastic sampling in generating cross-participant diversity is unknown. If unconstrained or token-matched outputs from any model (or an ensemble of unconstrained outputs) approach human TTR (0.27), type count (476), or Zipf α (0.89), the conclusion that LLMs intrinsically lack human-like variability collapses. The ensemble analysis in Section 6 pools only the 33 count-constrained configurations (with outliers removed), so it cannot test whether unconstrained outputs would diversify the pool.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether large language models can reproduce the inter-participant behavioral variability seen in a phonemic fluency task. Using 106 human participants who generated F-words, the authors prompt 34 LLMs in 45 configurations to role-play each participant, providing age, education, and the participant's exact number of correct responses in the prompt. They then compare participant-level response counts, lexical diversity (TTR, ITTTR), item-level Zipf distributions and linguistic correlates, word co-occurrence networks, and ensemble mixtures. The headline finding is that no single model or ensemble matches human-level variability; Claude 3.7 Sonnet is the most human-like but still produces fewer than half the unique types. The paper concludes that LLMs are systematically more rigid and convergent than humans and cautions against using them as substitutes for human participants.","tokens_in":14340,"tokens_out":6162,"duration_ms":49125,"significance":"The question of whether LLMs can simulate human behavioral variability is timely and important for cognitive modeling claims. The paper's strengths include unusually broad model coverage (34 models, 45 configurations), grounding in real participant metadata, and multiple complementary outcome measures (TTR, Zipf slopes, network metrics, ensemble sampling). The empirical finding that LLM Zipf slopes are steeper than human slopes and that lexicon overlap across models is very high is a concrete, falsifiable regularity. However, the central claim currently rests on a design that injects the exact human response count into every prompt and then uses adherence to that injected count as both a success criterion and a selection filter. If the diversity gap persists when models are allowed to choose their own stopping point or are otherwise not constrained by the exact count, the paper's conclusion would be substantially strengthened; if not, the observed rigidity may be an artifact of the instruction rather than an intrinsic LLM limitation.","major_comments":[{"comment":"The design is circular in a way that is load-bearing for the central claim. The prompt (Figure A1) includes 'Number of correct responses: {num_correct}', and Section 3.1 evaluates adherence using MAE against that same number, retaining only configurations with MAE≤1.69 and no outlier where |LLM−Human|>5. Low MAE is therefore a measure of instruction-following, not evidence of human-like production rates, and the retention filter selects on the injected value. Because every analyzed output is count-matched, the observed low TTR and steep Zipf slopes could be a direct consequence of count-constrained generation rather than an intrinsic property of LLMs. Appendix A reports that without the count, pilot models overproduced; yet overproduction may reflect deeper lexical exploration that yields more unique and idiosyncratic types. To support the conclusion that no LLM reproduces human variability, the paper needs a control condition without the count injection (or with a different target count) comparing TTR, ITTTR, and Zipf slopes, or an empirically grounded argument that the count constraint does not suppress diversity.","section":"Section 3.1 and Appendix A"},{"comment":"The conclusion that 'no LLM, individually or in ensemble, reproduces the scope of human variability' is only demonstrated for the 21 configurations that adhered exactly to the injected counts. The 12 overproducing configurations (e.g., GPT-4 Turbo with MAE=57.17, o3 with MAE=35.39) are excluded from all participant-level and item-level analyses. If overproduction tends to generate longer, more exploratory lists with more unique and idiosyncratic types, then excluding those configurations biases the diversity comparison in the direction of the paper's conclusion. The authors should either include these configurations in the variability analyses (e.g., by token-matching or by analyzing their diversity directly) or explicitly restrict the abstract and Section 7 claims to count-constrained, instruction-following configurations. As written, the claim 'no LLM' is too broad given the exclusions.","section":"Section 3.1 and Table A2"},{"comment":"The paper does not report any sampling parameters: temperature, top-p, or random seed are absent, and the 'thinking' / 'reasoning effort' modes are named but not defined numerically. Lexical diversity in LLM generation is strongly sensitive to sampling temperature; with zero temperature, repeated prompts can yield nearly identical outputs, which would artificially lower TTR and idiosyncratic type counts. Without reporting these parameters, the results are not reproducible, and the reader cannot determine whether the diversity gap reflects a model's capability or the authors' choice of decoding strategy. The authors should provide the exact API parameters used, or at least state that default parameters were used, and ideally vary temperature to show that the diversity gap is robust across sampling settings.","section":"Section 2.2 and Appendix A"}],"minor_comments":[{"comment":"The z-tests for deviation from random networks report negative z-values (e.g., human: z=−2460.1 for CC) while the text states that the networks exhibit 'substantially higher CCs and slightly higher ASPLs than random networks.' The sign is internally inconsistent; either the sign convention is reversed or the interpretation is incorrect. Please correct this discrepancy.","section":"Section 5.2"},{"comment":"The ensemble type range is reported as 143–214, but Claude 3.7 Sonnet alone produced 226 types, which is above the upper bound. The phrase 'values that fall within or near the upper tail of the ensemble distribution for types' is inaccurate for a value that exceeds the maximum; this should be reworded.","section":"Section 6.2"},{"comment":"Table 1 appears to list only a subset of the 21 retained configurations (the table shows 11 model rows), but the text refers to all retained configurations. The selection criterion for inclusion in Table 1 is not stated. If the table is intended to be representative or to show only the most diverse models, that should be explicitly indicated.","section":"Table 1"},{"comment":"The human data come from Qiu and Johns (2021), which involved two experiments; the paper should clarify whether the 106 participants are pooled across both experiments and whether experiment membership was included as a covariate or otherwise controlled for, since the two experiments may have had different procedures.","section":"Section 2.1"},{"comment":"The prompt uses 'Highest degree: {education}' while Section 2.1 describes education in years of schooling (M=14.92, SD=2.01). These are different operationalizations; please clarify how the 'education' variable was mapped to the prompt field, since the mapping could affect the fidelity of the role-play.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper tackles an important and well-timed question, and its breadth of models is a genuine contribution. The central concern is not unanswerable: a comparison of count-constrained versus unconstrained generation (with sampling parameters reported) would directly address whether the diversity gap is intrinsic. If the authors can provide that control and find the same gap, the paper would likely be a strong contribution. I also note that the manuscript's own Limitations section does not acknowledge the count-injection circularity, which is a worrying omission for a serious journal. The exclusion of 12 overproducing configurations from the main analyses should be prominently disclosed in the abstract or Section 1 if it cannot be remedied."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful paper for anyone using LLMs as stand-ins for human participants. The 34-model/45-config sweep on phonemic fluency is more extensive than prior single-model demonstrations, and the finding that ensembles don't help because of high vocabulary overlap is a nice, non-obvious result. The writing is clear, the network analysis is careful, and grounding prompts in real participant metadata is a good idea.\n\nThe soft spot is the count constraint. Every prompt includes the participant's exact number of correct responses (Figure A1). Then Section 3.1 uses MAE against that same number to declare 'successful' production-rate simulation, and the 12 overproducing configurations are excluded before diversity is measured. So the main claim—that no LLM reproduces human variability—is demonstrated only for models that were explicitly told the target count and complied. It is plausible, as the paper says in the appendix, that without the count models overproduce; but overproduction might come with more lexical exploration, and the paper never checks whether unconstrained outputs get closer to human TTR or Zipf slope. That makes the conclusion conditional in a way the abstract does not convey.\n\nThe paper also doesn't report sampling parameters (temperature, top-p), so we can't tell whether the cross-configuration differences are partly just decoding choices. And the exclusion of 12 MAE-failing plus 12 outlier-containing configurations leaves 21 of 45 for the core analyses; that's a lot of data swept under the rug.\n\nThat said, I don't think the paper is wrong about the direction. The lexical diversity gap is real for count-matched outputs, and the network and regression results are descriptive facts that stand independently. The issue is the inference from those facts to 'LLMs intrinsically lack human-like variability.' A revision that adds an unconstrained condition, reports diversity for overproducers, and releases code and temperature settings would make this a solid contribution. As is, it deserves peer review but with a request for substantial revision.\n\nWho's it for: anyone working on LLM-as-participant simulation, and cognitive scientists using fluency tasks. Worth a reading-group discussion as a methodological cautionary tale.","headline":"Broad model sweep with a real design flaw: the count-constrained prompt makes the headline claim about LLM variability weaker than the abstract suggests.","tokens_in":14909,"tokens_out":3028,"would_cite":false,"duration_ms":27348,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No LLM, individually or in an ensemble, reproduces the scope of human behavioral variability in a phonemic fluency task.","keywords":["large language models","phonemic fluency","behavioral variability","cognitive simulation","lexical diversity","verbal fluency","Zipf's law"],"falsifier":"Run the same 106 demographic prompts but omit the 'number of correct responses' line; if any single model or ensemble then produces a unique-type count near 476 or a Zipf slope statistically indistinguishable from 0.89, the claim that LLMs cannot reproduce the scope of human variability would be falsified.","tokens_in":13824,"feed_emoji":"🧠","tokens_out":6469,"duration_ms":49058,"temperature":0.7,"pith_summary":"This paper asks whether large language models can replace human participants in a cognitive experiment, using the phonemic fluency task in which people list words starting with F. It compares 34 models across 45 configurations with responses from 106 human participants. The central claim is that no model, alone or combined with others, reproduces the scope of human variability: humans produced 476 distinct words, the best model only 226, and every LLM's rank-frequency distribution was steeper than the human one. If this is right, LLMs systematically produce more rigid and convergent output than human behavior, which would undercut their use as stand-ins in behavioral research.","feed_headline":"Best LLM produced fewer than half the unique words humans did","feed_subtitle":"Ensembling 34 models kept outputs more rigid and convergent than 106 human participants did.","key_machinery":"The argument is carried by the phonemic fluency task itself paired with distributional indices of variability. Phonemic fluency requires an effortful, form-based search of the mental lexicon, which is unnatural in everyday language use, so it is a stringent probe of whether LLMs can mimic individual differences rather than central tendencies. The key measurements are the type-to-token ratio and the share of idiosyncratic types, the Zipf scaling exponent of the rank-frequency distribution, and word co-occurrence networks filtered with the Triangulated Maximally Filtered Graph. These measures are what separate human output's long tail of rare words from the LLMs' steeper, more convergent distributions.","core_discovery":"The paper's core discovery is that the variability of human word production in a phonemic fluency task is not approximated by any tested LLM. Claude 3.7 Sonnet came closest on averages and on which linguistic features predicted production, but it generated fewer than half of the human unique word types (226 vs. 476), a lower type-to-token ratio (0.13 vs. 0.27), and fewer idiosyncratic words (73 vs. 201). Every LLM followed Zipf's law with a steeper scaling exponent than humans (alpha 1.19 to 1.53 vs. 0.89), meaning high-frequency words dominated more and the long tail of rare words was thinner. Word co-occurrence networks also differed structurally: human output formed tighter local clusters with weaker global integration, while Claude's network was more evenly connected, and the two networks' similarity matrices correlated only modestly. Ensembling outputs from a random mix of models across 1,000 simulations did not recover human diversity, because model vocabularies overlapped so heavily (mean overlap 0.74).","pith_inferences":["Because the prompt supplied each participant's exact response count, the observed rigidity could partly be an artifact of that constraint; a version that lets models choose their own stopping point could yield more diverse output.","The high vocabulary overlap across providers suggests training-data and alignment pressures converge on high-frequency, prototypical responses, which may affect any open-ended generation task where rare or idiosyncratic content is valued.","A concrete extension would be to condition prompts on richer participant traits such as profession or reading history, and test whether that restores variability without sacrificing adherence to response counts.","If the network difference reflects true retrieval dynamics, one could test it behaviorally by comparing human switch costs or reaction times with LLM token-level latencies, though such a comparison is not in this paper."],"forward_implications":["Model selection matters: newer models and thinking-enabled modes can reduce lexical diversity, so evaluations should cover multiple providers and versions.","Ensemble sampling across diverse models is not a shortcut to human-like variability, because providers share a largely common word pool.","LLMs may be better treated as baselines that capture central tendencies, against which human uniqueness and flexibility can be measured.","Using LLMs as substitutes for human participants in fluency-based or similar open-ended cognitive tasks would likely underrepresent population variability."],"supporting_citations":[{"why":"Supplies the human F-fluency dataset of 106 participants whose age, education, and response counts are used as prompts and as the human benchmark.","marker":"Qiu and Johns (2021)"},{"why":"Establishes the Zipf's-law expectation for verbal fluency and provides the rank-frequency analysis that the paper uses to compare human and LLM distributions.","marker":"Taler et al. (2020)"},{"why":"Provides the English Lexicon Project variables (word frequency, neighborhood sizes, age of acquisition) used in the correlation and regression analyses.","marker":"Balota et al. (2007)"},{"why":"Supplies the co-occurrence network construction and bootstrap subnetwork comparison method used to compare human and LLM retrieval structure.","marker":"Borodkin et al. (2016)"},{"why":"Provides the TMFG network filtering algorithm that retains the strongest word associations in the co-occurrence networks.","marker":"Massara et al. (2017)"},{"why":"Prior finding that LLM semantic fluency networks differ from human networks, which motivates the phonemic task and the inclusion of response counts in prompts.","marker":"Wang et al. (2025b)"},{"why":"Proposed the idea of LLMs replacing human participants that the study challenges.","marker":"Dillion et al. (2023)"}],"fun_headline_variants":["No LLM matches human variability in phonemic fluency","LLMs converge, humans diverge in word fluency task","Ensembling 34 LLMs doesn't yield human-like word diversity","All 34 LLMs fall short on mimicking human word variety","Closest LLM still misses human word variability in test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on the assumption that giving the model the participant's exact number of correct responses does not itself shrink the variety of words it produces; if models were more diverse when free to choose their own stopping point, the reported gap would be partly an artifact of the prompt.","fun_headline_variants_meta":{"raw":{"variants":["No LLM matches human variability in phonemic fluency","LLMs converge, humans diverge in word fluency task","Ensembling 34 LLMs doesn't yield human-like word diversity","All 34 LLMs fall short on mimicking human word variety","Closest LLM still misses human word variability in test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000812,"raw_usage":{"total_tokens":3556,"prompt_tokens":936,"completion_tokens":2620,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":2550}},"tokens_in":552,"tokens_out":2620,"duration_ms":17416,"temperature":1.0,"reasoning_tokens":2550,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:05:47.244035+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 106 demographic prompts but omit the 'number of correct responses' line; if any single model or ensemble then produces a unique-type count near 476 or a Zipf slope statistically indistinguishable from 0.89, the claim that LLMs cannot reproduce the scope of human variability would be falsified.","supporting_citations":[],"review_version":1}