{"id":"ca5f934e-5a82-4f03-b471-93060d2b453e","arxiv_id":"2510.04391","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Human vividness-rating networks are correlated across populations and cluster by questionnaire context, whereas LLM-derived networks are mostly degenerate single-clusters, showing a human-LLM divergence in imagined-scene structure.","lead":"Researchers built networks from mental-imagery vividness ratings in humans and six large language models, and found human networks look similar across countries while LLM networks mostly collapse into a single cluster. The result suggests LLM text training does not reproduce the structured way people rate imagined scenes and sensations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The LLM-degeneracy result may be an EBICglasso artifact: single-cluster networks follow automatically from low conditional dependence, and no null model rules this out.","rationale":"The reader's weakest assumption precisely identifies the most load-bearing gap: the interpretation of empty/single-cluster LLM networks as evidence that LLMs lack human-like internal world models. The paper provides no null model, simulation, or power analysis to show that EBICglasso would recover edges in LLM rating data if any existed. In fact, the expected behavior of regularized partial-correlation estimation under a strong common factor or sparse covariance is an empty graph, so ARI = 0 follows automatically. Because the central claim of LLM failure is built on this empirical pattern, the conclusion is not yet secure. The proposed synthetic-data test would settle whether the single-cluster result is informative or artifactual. I do not see a stronger or more specific objection: the composite-sample overlap inflates some human-human correlations, but non-overlapping human pairs (e.g., Florida vs Poland-1) already show significant centrality correlations, so the human-alignment descriptive finding has independent support. The abstract's mischaracterization of the ARI analysis is a reporting error but secondary to the main statistical concern. The paper has real strengths: multiple independent human datasets, preregistered-style questionnaire administration, and a concrete computational pipeline; those warrant conditional acceptance pending the null-model check. Therefore the reader's CONDITIONAL verdict remains appropriate, and no verdict change is needed.","tokens_in":34714,"tokens_out":4254,"duration_ms":38813,"concrete_test":"Run the paper's exact network pipeline (EBICglasso with Spearman correlations, tuning parameter 0.5, walktrap community detection) on synthetic datasets calibrated to LLM response distributions. Specifically, for one representative LLM condition (e.g., Gemma3:12b independent, VVIQ-2), estimate the item covariance structure and generate 1000 simulated samples from a one-factor model with loadings matching the observed common variance, or from an independence model with the same item marginals. If the resulting networks are single-cluster and have ARI = 0 with the reported human partitions, the LLM degeneracy is a statistical artifact of low conditional dependence, not evidence about internal world model structure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that LLMs fail to replicate human imagination-network structure rests primarily on the observation that most LLM-derived networks are degenerate single-cluster topologies, yielding ARI = 0 with human partitions. But this outcome is exactly what EBICglasso produces when no partial-correlation edges survive regularization. If LLM vividness ratings are dominated by a single general factor (e.g., an overall vividness tendency induced by imagination-ability prompts) or have weak item-level covariances, then all pairwise partial correlations are near zero, the regularized graphical lasso returns an empty graph, and any empty graph is a single community. For a single-cluster partition versus any multi-cluster human partition, the adjusted Rand Index is identically zero because observed and expected pair-agreement coincide. The paper does not report raw item covariance structure for LLM responses, does not check whether EBICglasso has statistical power to recover edges in these particular data, and does not compare against synthetic null datasets. Table S1/S2 report only centrality stability coefficients, not edge-recovery simulations. Thus the 'single-cluster' result is not yet diagnostic about whether LLMs lack a structured internal world model; it may be an estimator artifact. A related but secondary issue is that the abstract's claim that clusters are 'aligned with VVIQ-2 scene contexts (ARI=0.27-0.40)' is not supported by the reported analyses, since the cited ARIs are alignments between different human networks, not alignments to the a priori scene labels. The null-model gap is the load-bearing concern because the LLM failure is the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces \"imagination networks\"—EBICglasso partial-correlation graphs of vividness ratings on the VVIQ-2 and PSIQ—and compares these networks across three human populations (Florida, Poland, London; N=2,743) and six LLM variants under independent and cumulative prompting. The central claims are: (1) human imagination networks show robust cross-population centrality correlations for expected influence, strength, and closeness, and community structure that is aligned across human groups; and (2) LLM networks largely fail to replicate this structure, with most LLM configurations producing degenerate single-cluster topologies (median ARI=0), which the authors interpret as evidence that LLMs lack a human-like internal world model. The paper also reports instability of betweenness centrality and task-dependent differences in LLM performance.","tokens_in":35033,"tokens_out":4165,"duration_ms":35163,"significance":"If the empirical claims were fully supported, this would be a valuable contribution to comparative cognitive modeling: it proposes a network-based, falsifiable operationalization of internal world models and applies it across multiple human samples and LLM families. Strengths include the use of multiple LLM architectures, two validated questionnaires, and the promise of public code and data. However, the two load-bearing claims—human cross-population robustness and LLM degeneracy—currently rest on partially non-independent comparisons and on an estimation procedure that may confound low covariance with absent structure. The manuscript's conclusions therefore outrun the evidence as presented.","major_comments":[{"comment":"The claim of robust human cross-population centrality correlations is overstated because the highest correlations come from overlapping composite groups. For instance, PSIQ expected influence Florida vs. London is r=0.425 with p_BH=0.12 (Fig. 4B-1), and VVIQ-2 closeness Florida vs. Poland2 is r=0.313 with p_BH=0.23 (Fig. 5A-1)—both non-significant after FDR. In contrast, comparisons such as Florida vs. Florida+London (r=0.914) and Poland1 vs. Poland All (r=0.931) are necessarily inflated by shared participants. The paper should report and interpret the non-overlapping comparisons separately and temper the abstract's 'robust cross-population' language accordingly.","section":"Results, Centrality correlations (Fig. 4 and Fig. 5)"},{"comment":"The abstract states that community detection 'recovered clusters aligned with VVIQ-2 scene contexts (ARI = 0.27–0.40)'. However, the reported ARI values are between data-driven clusterings of different human networks (e.g., Florida vs. Poland1), not between those clusterings and the questionnaire's a priori scene-context partition. The paper never computes ARI against the known eight VVIQ-2 contexts or seven PSIQ modalities. As a result, the claim of alignment with scene contexts is not supported by the analysis presented. Please either compute alignment against the true item-role partition or rephrase the claim.","section":"Results, Clustering (Fig. 6; Tables S3–S6)"},{"comment":"The conclusion that LLMs 'fail to replicate human network structure' rests primarily on the observation that most LLM networks are single-cluster (median ARI=0). This outcome is exactly what EBICglasso produces when no partial-correlation edges survive regularization. If LLM vividness ratings have low item covariance or are dominated by a single general factor, the estimated graph will be empty by construction, and any empty graph yields ARI=0 against any multi-cluster human partition. The manuscript does not provide raw item-covariance summaries for LLM responses, simulation-based power analyses, or synthetic null-network comparisons to rule out this estimator artifact. Supplementary Tables S1/S2 report CS-coefficients, which assess centrality stability under case-dropping, not edge-recovery accuracy. Thus the single-cluster finding is not yet diagnostic about whether LLMs lack a struct","section":"Methods, Network Analysis; Results, Clustering (LLM single-cluster findings)"}],"minor_comments":[{"comment":"The correlation range 'r = 0.31–0.93' mixes non-significant and overlapping-group comparisons; consider reporting a range restricted to non-overlapping, FDR-significant tests.","section":"Abstract / Fig. 4 caption"},{"comment":"The sentence 'See Table S1, S2 for clustering assignment' appears to refer to clustering, but Tables S1/S2 are CS-coefficient tables; the clustering assignments are in Tables S3/S4. Please correct this cross-reference.","section":"Methods, Network Analysis"},{"comment":"'EbicGlasso' is misspelled; the standard spelling is 'EBICglasso' as used in the text.","section":"Fig. 2 and Fig. 3 captions"},{"comment":"The paper states that 'we inferred that the imagined scenario was similarly involved across the networks' from centrality correlations. This linking of network correlations to internal-world-model similarity is partly definitional and should be flagged as an interpretive assumption rather than an empirical inference.","section":"Results, Overview of Network Estimation"},{"comment":"The claim that LLMs 'lack clear phenomenological structures' is stronger than the evidence supports, given the open possibility that the degeneracy reflects estimation limitations. Please temper this and related statements throughout the Discussion.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The core idea is interesting and potentially publishable, but the two main empirical pillars—human cross-population robustness and LLM degeneracy—need substantially stronger support. I would require a revision that (a) reanalyzes centrality/clustering using only non-overlapping human groups, (b) computes ARI against the true questionnaire context partitions, and (c) adds simulation or null-model analyses establishing that EBICglasso can recover structured graphs from LLM-like data with known cluster structure. Without these additions, the central claims are not yet established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look, but the headline claim doesn't survive close reading. The paper applies regularized partial-correlation network analysis to VVIQ-2 and PSIQ vividness ratings from three human samples and six LLM variants, and reports that human networks correlate across populations while LLM networks collapse to single clusters. The human-side descriptive finding is real and new: people in Florida, Poland, and London produce reasonably similar item-importance rankings and similar community structure for these questionnaires. That part is worth taking seriously.\n\nThe LLM side is where the logic goes soft. The paper interprets the single-cluster LLM topology as evidence that LLMs lack a structured internal world model. But single-cluster topologies are exactly what EBICglasso produces when no partial correlations survive regularization—which happens when LLM ratings have low item covariance or are dominated by a single general factor. The paper provides no edge-recovery simulations, no null-model baseline, and no report of the raw covariance structure for LLM responses. Until that gap is closed, the 'failure to replicate' claim is an estimator artifact in disguise. This is not a minor point; it is the paper's central conclusion.\n\nSecondary issues: the abstract says clusters were 'aligned with VVIQ-2 scene contexts (ARI = 0.27–0.40)', but those ARIs are alignments between human networks, not against a priori scene labels. And several non-overlapping human–human correlations fail to survive FDR (e.g., PSIQ expected influence Florida vs London r = 0.425, p_BH = 0.12); the strongest correlations come from overlapping composite groups. The word 'robust' overstates the evidence.\n\nWhat the paper does well: it treats questionnaire vividness ratings as multivariate network data, uses appropriate centrality measures, and checks stability with CS-coefficients. The idea of comparing humans and LLMs with the same instruments is sensible. But the analysis needs three things before it can support the conclusion: null models for LLM edge detection, a corrected abstract, and separation of overlapping-sample correlations.\n\nVerdict: send to peer review, but only with the expectation of major revision. The human data are a useful benchmark; the LLM claim is premature. A reader working on LLM world models should read this carefully, mainly to learn where the pitfalls are.","headline":"A useful human-side benchmark with an unexamined artifact at the core: the LLM single-cluster result may reflect empty network estimation rather than absent internal world structure.","tokens_in":35576,"tokens_out":4187,"would_cite":false,"duration_ms":31950,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human vividness ratings form shared imagination networks across populations, while six large language models largely fail to replicate this structure, collapsing to single clusters.","keywords":["imagination networks","internal world models","vividness ratings","psychological network analysis","large language models","regularized partial correlations","community detection","Adjusted Rand Index"],"falsifier":"Compute the inter-item covariance (or average pairwise Spearman correlation) of each LLM's raw vividness ratings and run EBICglasso on synthetic data with the same low variance but a planted multi-cluster structure; if the estimator returns one cluster for the synthetic data, the median ARI = 0 in this study is explained by estimation power, not by the absence of structure in LLMs.","tokens_in":34562,"feed_emoji":"🧠","tokens_out":8378,"duration_ms":76796,"temperature":0.7,"pith_summary":"The paper sets out to show that imagination has a measurable relational structure, and that people across different populations share it while large language models do not. Using vividness ratings from two imagery questionnaires, the authors build 'imagination networks' in which each imagined scenario is a node and edges capture how vividness ratings move together. Human networks show correlated node importance and clusters that line up with the questionnaires' scene and sensory categories. LLM networks, across model family, scale, and conversational memory, mostly collapse into a single cluster, with near-zero alignment to human clusters. The authors read this as evidence that human imagination draws on a world model organized by embodied memory, which text-only training does not reproduce.","feed_headline":"LLM imagination networks collapse; human ones align","feed_subtitle":"Human vividness ratings form similar network maps across countries; six LLMs each collapsed to one cluster.","key_machinery":"The key instrument is the imagination network: a regularized partial-correlation graph (EBICglasso) in which each questionnaire item is a node and edges are partial associations between vividness ratings, pruned of weak and spurious connections. Two levels of comparison carry the argument. At the micro level, node centrality measures (expected influence, strength, closeness, betweenness) are correlated across networks to ask whether the same imagined scenarios are similarly important in different populations. At the meso level, the walktrap community-detection algorithm recovers clusters and the Adjusted Rand Index quantifies how aligned those clusters are between two networks. Together they","core_discovery":"The central discovery is a systematic divergence between human and LLM imagination networks. Across three human populations, regularized partial-correlation networks estimated from vividness ratings show substantial cross-population correlations for expected influence, strength, and closeness (r = 0.31 to 0.93), and community detection returns clusters that align with the questionnaire contexts (ARI up to 0.40 for VVIQ-2 scenes and 0.87 to 1.0 for PSIQ sensory modalities); betweenness was the one unstable measure. For six LLM variants, correlations with human centrality were weak and mostly non-significant after correction, and most LLM networks had a single cluster, giving median ARI = 0 ag","pith_inferences":["One thing the paper leaves open: the single-cluster LLM result may be an estimation artifact. If LLM vividness ratings have very low inter-item covariance, the regularized graph estimator will produce an empty or fully connected graph by construction, which automatically yields ARI = 0 against any multi-cluster human partition. A simulation with LLM-like response variance but planted human-like cl","The framework implies a concrete alignment metric for AI development: fine-tune an LLM to match human centrality vectors and cluster assignments, then test whether downstream imagination-dependent behavior (scene description, planning) becomes more human-like. That would test whether network similarity is causally relevant or merely decorative.","The PSIQ closeness correlations suggest a gradient: LLMs approach human structure for modality-labeled sensory items, but not for scene-context items. A testable prediction is that questionnaires tapping autobiographical or episodic content will show the largest human-LLM divergence."],"forward_implications":["If human imagination networks reflect a shared internal world model, then individual differences in imagery vividness are structured: dependency patterns among imagined scenarios carry information that total scores miss.","Betweenness centrality proved unstable across human populations, so it is not a reliable marker of shared imagination structure; expected influence, strength, and closeness are the stable measures.","Because LLM failure was consistent across model scale (12B to 272B), architecture, and memory conditions, scaling or adding conversational memory is unlikely to reproduce human imagination network structure.","For PSIQ, closeness centrality in several LLMs did correlate with humans, suggesting LLMs come closer to human structure for sensory modalities than for environmental scene contexts."],"fun_headline_variants":["LLM imagination collapses to one cluster; humans stay structured","Human mind maps align across countries, LLMs don't","Six AI models fail to mimic human imagination networks","Imagination networks: humans universal, LLMs degenerate","Why LLMs can't capture human mental imagery structure"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that LLMs lack human-like imagination structure assumes that the network estimator has enough statistical power to detect edges in LLM rating data; if LLM responses vary too little across items, a single-cluster or empty graph is an artifact of the method, not evidence about internal world models.","fun_headline_variants_meta":{"raw":{"variants":["LLM imagination collapses to one cluster; humans stay structured","Human mind maps align across countries, LLMs don't","Six AI models fail to mimic human imagination networks","Imagination networks: humans universal, LLMs degenerate","Why LLMs can't capture human mental imagery structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00077,"raw_usage":{"total_tokens":3312,"prompt_tokens":875,"completion_tokens":2437,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":2360}},"tokens_in":619,"tokens_out":2437,"duration_ms":13635,"temperature":1.0,"reasoning_tokens":2360,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T11:29:30.865425+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the inter-item covariance (or average pairwise Spearman correlation) of each LLM's raw vividness ratings and run EBICglasso on synthetic data with the same low variance but a planted multi-cluster structure; if the estimator returns one cluster for the synthetic data, the median ARI = 0 in this study is explained by estimation power, not by the absence of structure in LLMs.","supporting_citations":[],"review_version":1}