Pith. sign in

REVIEW 3 major objections 5 minor 78 references

Human vividness ratings form shared imagination networks across populations, while six large language models largely fail to replicate this structure, collapsing to single clusters.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Human vividness-rating networks are correlated across populations and cluster by questionnaire context, whereas LLM-derived networks are mostly degenerate single-clusters, showing a human-LLM divergence in imagined-scene structure.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection A useful human-side benchmark with an unexamined artifact at the core: the LLM single-cluster result may reflect empty network estimation rather than absent internal world structure. the 3 major comments →

arxiv 2510.04391 v5 pith:P5ITWKNU submitted 2025-10-05 cs.AI cs.CLcs.SIq-bio.NC

Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate

classification cs.AI cs.CLcs.SIq-bio.NC
keywords imagination networksinternal world modelsvividness ratingspsychological network analysislarge language modelsregularized partial correlationscommunity detectionAdjusted Rand Index
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that imagination has a measurable relational structure, and that people across different populations share it while large language models do not. Using vividness ratings from two imagery questionnaires, the authors build 'imagination networks' in which each imagined scenario is a node and edges capture how vividness ratings move together. Human networks show correlated node importance and clusters that line up with the questionnaires' scene and sensory categories. LLM networks, across model family, scale, and conversational memory, mostly collapse into a single cluster, with near-zero alignment to human clusters. The authors read this as evidence that human imagination draws on a world model organized by embodied memory, which text-only training does not reproduce.

Core claim

The central discovery is a systematic divergence between human and LLM imagination networks. Across three human populations, regularized partial-correlation networks estimated from vividness ratings show substantial cross-population correlations for expected influence, strength, and closeness (r = 0.31 to 0.93), and community detection returns clusters that align with the questionnaire contexts (ARI up to 0.40 for VVIQ-2 scenes and 0.87 to 1.0 for PSIQ sensory modalities); betweenness was the one unstable measure. For six LLM variants, correlations with human centrality were weak and mostly non-significant after correction, and most LLM networks had a single cluster, giving median ARI = 0 ag

What carries the argument

The key instrument is the imagination network: a regularized partial-correlation graph (EBICglasso) in which each questionnaire item is a node and edges are partial associations between vividness ratings, pruned of weak and spurious connections. Two levels of comparison carry the argument. At the micro level, node centrality measures (expected influence, strength, closeness, betweenness) are correlated across networks to ask whether the same imagined scenarios are similarly important in different populations. At the meso level, the walktrap community-detection algorithm recovers clusters and the Adjusted Rand Index quantifies how aligned those clusters are between two networks. Together they

Load-bearing premise

The claim that LLMs lack human-like imagination structure assumes that the network estimator has enough statistical power to detect edges in LLM rating data; if LLM responses vary too little across items, a single-cluster or empty graph is an artifact of the method, not evidence about internal world models.

What would settle it

Compute the inter-item covariance (or average pairwise Spearman correlation) of each LLM's raw vividness ratings and run EBICglasso on synthetic data with the same low variance but a planted multi-cluster structure; if the estimator returns one cluster for the synthetic data, the median ARI = 0 in this study is explained by estimation power, not by the absence of structure in LLMs.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If human imagination networks reflect a shared internal world model, then individual differences in imagery vividness are structured: dependency patterns among imagined scenarios carry information that total scores miss.
  • Betweenness centrality proved unstable across human populations, so it is not a reliable marker of shared imagination structure; expected influence, strength, and closeness are the stable measures.
  • Because LLM failure was consistent across model scale (12B to 272B), architecture, and memory conditions, scaling or adding conversational memory is unlikely to reproduce human imagination network structure.
  • For PSIQ, closeness centrality in several LLMs did correlate with humans, suggesting LLMs come closer to human structure for sensory modalities than for environmental scene contexts.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • One thing the paper leaves open: the single-cluster LLM result may be an estimation artifact. If LLM vividness ratings have very low inter-item covariance, the regularized graph estimator will produce an empty or fully connected graph by construction, which automatically yields ARI = 0 against any multi-cluster human partition. A simulation with LLM-like response variance but planted human-like cl
  • The framework implies a concrete alignment metric for AI development: fine-tune an LLM to match human centrality vectors and cluster assignments, then test whether downstream imagination-dependent behavior (scene description, planning) becomes more human-like. That would test whether network similarity is causally relevant or merely decorative.
  • The PSIQ closeness correlations suggest a gradient: LLMs approach human structure for modality-labeled sensory items, but not for scene-context items. A testable prediction is that questionnaires tapping autobiographical or episodic content will show the largest human-LLM divergence.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces "imagination networks"—EBICglasso partial-correlation graphs of vividness ratings on the VVIQ-2 and PSIQ—and compares these networks across three human populations (Florida, Poland, London; N=2,743) and six LLM variants under independent and cumulative prompting. The central claims are: (1) human imagination networks show robust cross-population centrality correlations for expected influence, strength, and closeness, and community structure that is aligned across human groups; and (2) LLM networks largely fail to replicate this structure, with most LLM configurations producing degenerate single-cluster topologies (median ARI=0), which the authors interpret as evidence that LLMs lack a human-like internal world model. The paper also reports instability of betweenness centrality and task-dependent differences in LLM performance.

Significance. If the empirical claims were fully supported, this would be a valuable contribution to comparative cognitive modeling: it proposes a network-based, falsifiable operationalization of internal world models and applies it across multiple human samples and LLM families. Strengths include the use of multiple LLM architectures, two validated questionnaires, and the promise of public code and data. However, the two load-bearing claims—human cross-population robustness and LLM degeneracy—currently rest on partially non-independent comparisons and on an estimation procedure that may confound low covariance with absent structure. The manuscript's conclusions therefore outrun the evidence as presented.

major comments (3)
  1. [Results, Centrality correlations (Fig. 4 and Fig. 5)] The claim of robust human cross-population centrality correlations is overstated because the highest correlations come from overlapping composite groups. For instance, PSIQ expected influence Florida vs. London is r=0.425 with p_BH=0.12 (Fig. 4B-1), and VVIQ-2 closeness Florida vs. Poland2 is r=0.313 with p_BH=0.23 (Fig. 5A-1)—both non-significant after FDR. In contrast, comparisons such as Florida vs. Florida+London (r=0.914) and Poland1 vs. Poland All (r=0.931) are necessarily inflated by shared participants. The paper should report and interpret the non-overlapping comparisons separately and temper the abstract's 'robust cross-population' language accordingly.
  2. [Results, Clustering (Fig. 6; Tables S3–S6)] The abstract states that community detection 'recovered clusters aligned with VVIQ-2 scene contexts (ARI = 0.27–0.40)'. However, the reported ARI values are between data-driven clusterings of different human networks (e.g., Florida vs. Poland1), not between those clusterings and the questionnaire's a priori scene-context partition. The paper never computes ARI against the known eight VVIQ-2 contexts or seven PSIQ modalities. As a result, the claim of alignment with scene contexts is not supported by the analysis presented. Please either compute alignment against the true item-role partition or rephrase the claim.
  3. [Methods, Network Analysis; Results, Clustering (LLM single-cluster findings)] The conclusion that LLMs 'fail to replicate human network structure' rests primarily on the observation that most LLM networks are single-cluster (median ARI=0). This outcome is exactly what EBICglasso produces when no partial-correlation edges survive regularization. If LLM vividness ratings have low item covariance or are dominated by a single general factor, the estimated graph will be empty by construction, and any empty graph yields ARI=0 against any multi-cluster human partition. The manuscript does not provide raw item-covariance summaries for LLM responses, simulation-based power analyses, or synthetic null-network comparisons to rule out this estimator artifact. Supplementary Tables S1/S2 report CS-coefficients, which assess centrality stability under case-dropping, not edge-recovery accuracy. Thus the single-cluster finding is not yet diagnostic about whether LLMs lack a struct
minor comments (5)
  1. [Abstract / Fig. 4 caption] The correlation range 'r = 0.31–0.93' mixes non-significant and overlapping-group comparisons; consider reporting a range restricted to non-overlapping, FDR-significant tests.
  2. [Methods, Network Analysis] The sentence 'See Table S1, S2 for clustering assignment' appears to refer to clustering, but Tables S1/S2 are CS-coefficient tables; the clustering assignments are in Tables S3/S4. Please correct this cross-reference.
  3. [Fig. 2 and Fig. 3 captions] 'EbicGlasso' is misspelled; the standard spelling is 'EBICglasso' as used in the text.
  4. [Results, Overview of Network Estimation] The paper states that 'we inferred that the imagined scenario was similarly involved across the networks' from centrality correlations. This linking of network correlations to internal-world-model similarity is partly definitional and should be flagged as an interpretive assumption rather than an empirical inference.
  5. [Discussion] The claim that LLMs 'lack clear phenomenological structures' is stronger than the evidence supports, given the open possibility that the degeneracy reflects estimation limitations. Please temper this and related statements throughout the Discussion.

Circularity Check

0 steps flagged

No significant circularity: the network comparison is self-contained; only non-load-bearing self-citation found.

full rationale

The study does not derive a prediction from an input that is itself the target. Imagination networks are estimated directly from human and LLM vividness ratings via EBICglasso; the human-consistency and LLM-divergence claims are then based on pairwise centrality correlations and ARI computed from the estimated partitions. No fitted parameter is renamed as a prediction, and no uniqueness theorem or ansatz from the authors' prior work is invoked to force the conclusion. The only self-citation (ref. 43) concerns background theory on heterarchical imagination and is not load-bearing. The strongest potential concern is statistical: ARI=0 is mathematically forced whenever one partition is single-cluster, so the paper's 'median ARI=0' is a restatement of the single-cluster LLM partitions rather than an independent alignment measurement, and the EBICglasso could yield empty graphs under low partial correlations without a null-model check (the paper itself admits initial instability for LLMs). This is a validity threat, not a circular reduction, because the single-cluster partition is an empirical estimate and the failure claim also rests on centrality correlations. Therefore a low circularity score is appropriate.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 1 invented entities

The central comparison rests on the assumption that questionnaire vividness ratings—from both humans and LLMs—are valid indices of an underlying 'internal world model,' and that the network estimation procedure can recover its structure. The population-diversity sampling introduces a free reweighting scheme fitted to human quantiles. The IWM construct is not independently evidenced.

free parameters (3)
  • EBICglasso tuning parameter (gamma) = 0.5 (fixed)
    Standard default in psychological networks; not fitted, but affects edge sparsity and thus cluster structure.
  • Population diversity sampling bin parameters = 60 simulations per 10th-percentile bin; total 600
    Post-hoc selection of LLM simulations based on human total-score quantiles; affects LLM network topology and comparability.
  • Number of LLM simulations per network = 600 (downsampled from 1000)
    Choice of sample size for network estimation; affects stability and edge detection.
axioms (6)
  • domain assumption Vividness ratings are a valid measure of the subjective quality of internal imagination.
    Methods/Imagination Tasks: 'We assume that vividness ratings capture a psychological experience based on the perceived vividness of the internal world.'
  • domain assumption A partial-correlation network of item vividness ratings represents the structure of an internal world model.
    Introduction and Network Analysis: edges between nodes are interpreted as representational dependencies in an IWM.
  • domain assumption LLM textual vividness ratings can be treated as direct analogues of human subjective vividness ratings.
    Methods/AI Models: LLMs are treated as cognitive agents whose outputs map to the same psychological scale as human reports.
  • domain assumption The Polish translation of VVIQ-2 is invariance-equivalent to the English version for network structure.
    Participants: Polish data used the Polish translated version; other populations used English, with no measurement-invariance check.
  • standard math Walktrap community detection on regularized partial-correlation networks yields meaningful, comparable clusters.
    Methods/Network Analysis: community detection is applied to the estimated networks with no validation against ground-truth clusterings.
  • standard math EBICglasso with Spearman correlations and gamma=0.5 provides reliable network estimates for ordinal ratings.
    Methods/Network Analysis: the authors note polychoric correlations would be preferred but were abandoned due to instability; Spearman is used instead.
invented entities (1)
  • Internal World Model (IWM) as a latent representational structure no independent evidence
    purpose: Explanatory construct for why human networks cluster and align, and why LLMs do not
    Defined and measured through the same network metrics used for comparison; no external behavioral or neural validation is provided, and the construct is inferred from the very networks it is meant to explain.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate." pith.science (2026). https://pith.science/paper/P5ITWKNU

@misc{pith2026251004391,
  author       = {Pith},
  title        = {Pith review of: Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P5ITWKNU}},
  note         = {Machine review of arXiv:2510.04391}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic large language model (LLM) populations remains unknown. We applied psychological network analysis to vividness ratings from two validated questionnaires: the Vividness of Visual Imagery Questionnaire (VVIQ-2) and the Plymouth Sensory Imagery Questionnaire (PSIQ), across geographically and linguistically distinct human samples (Florida, Poland, and London; total N = 2,743) and six large language models (LLMs; Gemma3-12B/27B, their quantization-aware counterparts, Llama3.3-70B, and Llama4-16x17B). Imagination networks were constructed as regularized partial correlation graphs, with node centrality and community structure compared across populations using Pearson correlations and the Adjusted Rand Index (ARI). Human networks showed robust cross-population centrality correlations for expected influence, strength, and closeness (r = 0.31-0.93), and community detection recovered clusters aligned with VVIQ-2 scene contexts (ARI = 0.27-0.40) and PSIQ sensory modalities (ARI = 0.87-1.0). Betweenness centrality was unstable across all populations, consistent with its sensitivity to individual experiential history. LLMs failed to replicate human network structure: LLM-human centrality correlations were weak and largely non-significant after correction, and most LLM configurations produced degenerate single-cluster topologies (median ARI = 0). This failure was consistent across model architectures, parameter scales (12B-272B), and conversational conditions. We posit that these findings may be driven by human imagination networks reflecting memory organization accumulated through embodied experience, a representational structure that linguistic training alone does not reproduce regardless of model scale and conversational memory.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

78 extracted references · 4 canonical work pages

  1. [1]

    & Mateen, F

    Drubach, D., Benarroch, E. & Mateen, F. Imagination: its definition, purposes and neurobiology. Rev. Neurol. 45 , 353 (2007)

  2. [2]

    Moulton, S. T. & Kosslyn, S. M. Imagining predictions: mental imagery as mental emulation. Philos. Trans. R. Soc. B Biol. Sci. 364 , 1273–1280 (2009)

  3. [3]

    The human imagination: the cognitive neuroscience of visual mental imagery

    Pearson, J. The human imagination: the cognitive neuroscience of visual mental imagery. Nat. Rev. Neurosci. 20 , 624–634 (2019)

  4. [4]

    & Khacharem, A

    Mguidich, H., Zoudji, B. & Khacharem, A. Does imagination enhance learning? A systematic review and meta-analysis. Eur. J. Psychol. Educ. 39 , 1943–1978 (2024)

  5. [5]

    & Wilkinson, S

    Jones, M. & Wilkinson, S. From Prediction to Imagination. in The Cambridge Handbook of the Imagination (ed. Abraham, A.) 94–110 (Cambridge University Press, Cambridge, 2020). doi:10.1017/9781108580298.007

  6. [6]

    B., Robertson, T

    Klein, S. B., Robertson, T. E. & Delton, A. W. Facing the future: Memory as an evolved system for planning future acts. Mem. Cognit. 38 , 13–22 (2010)

  7. [7]

    Schacter, D. L. et al. The Future of Memory: Remembering, Imagining, and the Brain. Neuron 76 , 677–694 (2012)

  8. [8]

    Goddu, M. K. & Gopnik, A. The development of human causal learning and reasoning. Nat. Rev. Psychol. 3 , 319–339 (2024)

  9. [9]

    W., Sheskin, M

    Magid, R. W., Sheskin, M. & Schulz, L. E. Imagination and the generation of new ideas. Cogn. Dev. 34 , 99–110 (2015)

  10. [10]

    & Hayden, B

    Kidd, C. & Hayden, B. Y. The psychology and neuroscience of curiosity. Neuron 88 , 449–460 (2015). IMAGINE INTERNAL WORLD 43

  11. [11]

    & Tisserand, Y

    Sergeant-Perthuis, G., Ruet, N., Rudrauf, D., Ognibene, D. & Tisserand, Y. Influence of the Geometry of the world model on Curiosity Based Exploration. Preprint at https://doi.org/10.48550/arXiv.2304.00188 (2023)

  12. [12]

    Clement, C. A. & Falmagne, R. J. Logical reasoning, world knowledge, and mental imagery: Interconnections in cognitive processes. Mem. Cognit. 14 , 299–307 (1986)

  13. [13]

    & May, E

    Knauff, M. & May, E. Mental imagery, reasoning, and blindness. Q. J. Exp. Psychol. 59 , 161–177 (2006)

  14. [14]

    Reasoning with imagination

    Myers, J. Reasoning with imagination. in Epistemic uses of imagination 103–121 (Routledge, 2021)

  15. [15]

    & Pugalee, D

    Douville, P. & Pugalee, D. K. Investigating the relationship between mental imaging and mathematical problem solving. in Proceedings of the international conference of Mathematics Education into the 21stcentury project, September 2003 62–67 (2003)

  16. [16]

    Angell, J. R. Psychology: An Introductory Study of the Structure and Function of Human Consciousness . (H. Holt, 1908)

  17. [17]

    Silver, D. et al. Mastering the game of Go without human knowledge. Nature 550 , 354–359 (2017)

  18. [18]

    & Sutton, R

    Silver, D., Singh, S., Precup, D. & Sutton, R. S. Reward is enough. Artif. Intell. 299 , 103535 (2021)

  19. [19]

    & Veness, J

    Silver, D. & Veness, J. Monte-Carlo Planning in Large POMDPs. in Advances in Neural Information Processing Systems (eds Lafferty, J., Williams, C., Shawe-Taylor, J., Zemel, R. & Culotta, A.) vol. 23 (Curran Associates, Inc., 2010)

  20. [20]

    Eppe, M. et al. Intelligent problem-solving as integrated hierarchical reinforcement learning. Nat. Mach. Intell. 4 , 11–20 (2022). IMAGINE INTERNAL WORLD 44

  21. [21]

    Lin, Y.-C. et al. MIRA: Mental Imagery for Robotic Affordances. in Proceedings of The 6th Conference on Robot Learning 1916–1927 (PMLR, 2023)

  22. [22]

    Matsuo, Y. et al. Deep learning, reinforcement learning, and world models. Neural Netw. 152 , 267–275 (2022)

  23. [23]

    Racanière, S. et al. Imagination-Augmented Agents for Deep Reinforcement Learning. in Advances in Neural Information Processing Systems vol. 30 (Curran Associates, Inc., 2017)

  24. [24]

    Weber, T. et al. Imagination-Augmented Agents for Deep Reinforcement Learning. arXiv.org https://arxiv.org/abs/1707.06203v2 (2017)

  25. [25]

    Wu, W. et al. Mind’s Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models. in (2024)

  26. [26]

    J., Zhou, J

    Gershman, S. J., Zhou, J. & Kommers, C. Imaginative Reinforcement Learning: Computational Principles and Neural Mechanisms. J. Cogn. Neurosci. 29 , 2103–2113 (2017)

  27. [27]

    M., Zaleskiewicz, T

    Smieja, J. M., Zaleskiewicz, T. & Gasiorowska, A. Mental imagery shapes emotions in people’s decisions related to risk taking. Cognition 257 , 106082 (2025)

  28. [28]

    & Traczyk, J

    Zaleskiewicz, T., Bernady, A. & Traczyk, J. Entrepreneurial Risk Taking Is Related to Mental Imagery: A Fresh Look at the Old Issue of Entrepreneurship and Risk. Appl. Psychol. 69 , 1438–1469 (2020)

  29. [29]

    & Sobkow, A

    Zaleskiewicz, T., Traczyk, J. & Sobkow, A. Decision making and mental imagery: A conceptual synthesis and new research directions. J. Cogn. Psychol. 35 , 603–633 (2023)

  30. [30]

    & Ullman, T

    Balaban, H. & Ullman, T. D. The capacity limits of moving objects in the imagination. Nat. Commun. 16 , 5899 (2025)

  31. [31]

    The imaginative mind

    Abraham, A. The imaginative mind. Hum. Brain Mapp. 37 , 4197–4211 (2016). IMAGINE INTERNAL WORLD 45

  32. [32]

    & Bubic, A

    Abraham, A. & Bubic, A. Semantic memory as the root of imagination. Front. Psychol. 6 , (2015)

  33. [33]

    Imagination Machines: A New Challenge for Artificial Intelligence

    Mahadevan, S. Imagination Machines: A New Challenge for Artificial Intelligence. Proc. AAAI Conf. Artif. Intell. 32 , (2018)

  34. [34]

    R., Pan, L., Vu, M.-A., Laiser, N

    Addis, D. R., Pan, L., Vu, M.-A., Laiser, N. & Schacter, D. L. Constructive episodic simulation of the future and the past: Distinct subsystems of a core brain network mediate imagining and remembering. Neuropsychologia 47 , 2222–2238 (2009)

  35. [35]

    Schacter, D. L. & Addis, D. R. The cognitive neuroscience of constructive memory: remembering the past and imagining the future. Philos. Trans. R. Soc. B Biol. Sci. 362 , 773–786 (2007)

  36. [36]

    L., Addis, D

    Schacter, D. L., Addis, D. R. & Buckner, R. L. Episodic simulation of future events: Concepts, data, and applications. Ann. N. Y. Acad. Sci. 1124 , 39–60 (2008)

  37. [37]

    & Maguire, E

    Hassabis, D., Kumaran, D. & Maguire, E. A. Using imagination to understand the neural basis of episodic memory. J. Neurosci. 27 , 14365–14374 (2007)

  38. [38]

    & Maguire, E

    Hassabis, D. & Maguire, E. A. Deconstructing episodic memory with construction. Trends Cogn. Sci. 11 , 299–306 (2007)

  39. [39]

    Schacter, D. L. & Addis, D. R. On the constructive episodic simulation of past and future events. Behav. Brain Sci. 30 , 331–332 (2007)

  40. [40]

    L., Benoit, R

    Schacter, D. L., Benoit, R. G. & Szpunar, K. K. Episodic future thinking: mechanisms and functions. Curr. Opin. Behav. Sci. 17 , 41–50 (2017)

  41. [41]

    Spagna, A. et al. Visual mental imagery: Evidence for a heterarchical neural architecture. Phys. Life Rev. 48 , 113–131 (2024). IMAGINE INTERNAL WORLD 46

  42. [42]

    Spagna, A. et al. Competing models of visual mental imagery: Reverse hierarchy or heterarchy? Phys. Life Rev. 51 , 96–100 (2024)

  43. [43]

    & Odegaard, B

    Ranjan, S. & Odegaard, B. Heterarchy or hierarchy? Insights from a new model of visual imagination. Phys. Life Rev. 49 , 74–76 (2024)

  44. [44]

    & Kanwisher, N

    Casto, C., Ivanova, A., Fedorenko, E. & Kanwisher, N. What does it mean to understand language? ArXiv Prepr. ArXiv251119757 (2025)

  45. [45]

    Schwarzkopf, D. S. et al. Vividness of mental imagery is a broad trait measure of internally generated visual experiences. 2025.09.30.679629 Preprint at https://doi.org/10.1101/2025.09.30.679629 (2025)

  46. [46]

    Marks, D. F. New directions for mental imagery research. J. Ment. Imag. 19 , 153–167 (1995)

  47. [47]

    & Ganis, G

    Andrade, J., May, J., Deeprose, C., Baugh, S.-J. & Ganis, G. Assessing vividness of mental imagery: The Plymouth Sensory Imagery Questionnaire. Br. J. Psychol. 105 , 547–563 (2014)

  48. [48]

    R., Yao, S., Narasimhan, K

    Sumers, T. R., Yao, S., Narasimhan, K. & Griffiths, T. L. Cognitive Architectures for Language Agents. Preprint at https://doi.org/10.48550/arXiv.2309.02427 (2024)

  49. [49]

    Language Models as Agent Models

    Andreas, J. Language Models as Agent Models. Preprint at https://doi.org/10.48550/arXiv.2212.01681 (2022)

  50. [50]

    L., Isola, P

    Wang, S. L., Isola, P. & Cheung, B. Words That Make Language Models Perceive. Preprint at https://doi.org/10.48550/arXiv.2510.02425 (2025)

  51. [51]

    Jankowska, D. M. & Karwowski, M. How Vivid Is Your Mental Imagery? Eur. J. Psychol. Assess. 39 , 437–448 (2023). IMAGINE INTERNAL WORLD 47

  52. [52]

    Clark, I. A. & Maguire, E. A. Release of cognitive and multimodal MRI data including real-world tasks and hippocampal subfield segmentations. Sci. Data 10 , 540 (2023)

  53. [53]

    & Arabie, P

    Hubert, L. & Arabie, P. Comparing partitions. J. Classif. 2 , 193–218 (1985)

  54. [54]

    & Fried, E

    Epskamp, S. & Fried, E. I. A tutorial on regularized partial correlation networks. Psychol. Methods 23 , 617–634 (2018)

  55. [55]

    & Christensen, A

    Golino, H. & Christensen, A. EGAnet: Exploratory Graph Analysis – A Framework for Estimating the Number of Dimensions in Multivariate Data Using Network Psychometrics . (2025). doi:10.32614/CRAN.package.EGAnet

  56. [56]

    & Latapy, M

    Pons, P. & Latapy, M. Computing communities in large networks using random walks. in International symposium on computer and information sciences 284–293 (Springer, 2005)

  57. [57]

    Community detection in graphs

    Fortunato, S. Community detection in graphs. Phys. Rep. 486 , 75–174 (2010)

  58. [58]

    Bringmann, L. F. et al. What do centrality measures measure in psychological networks? J. Abnorm. Psychol. 128 , 892–903 (2019)

  59. [59]

    A Treatise of Human Nature

    Hume, D. A Treatise of Human Nature . (Oxford University Press, Oxford, 2009)

  60. [60]

    Two Kinds of Imaginative Vividness

    Langkau, J. Two Kinds of Imaginative Vividness. Can. J. Philos. 51 , 33–47 (2021)

  61. [61]

    Vividness and content

    Fazekas, P. Vividness and content. Mind Lang. 39 , 61–79 (2024)

  62. [62]

    Mental strength: A theory of experience intensity

    Morales, J. Mental strength: A theory of experience intensity. Philos. Perspect. 37 , 248–268 (2023)

  63. [63]

    & Oizumi, M

    Kawakita, G., Zeleznikow-Johnston, A., Tsuchiya, N. & Oizumi, M. Gromov–Wasserstein unsupervised alignment reveals structural correspondences between the color similarity structures of humans and large language models. Sci. Rep. 14 , 15917 (2024)

  64. [64]

    & Friston, K

    Pezzulo, G., Parr, T., Cisek, P., Clark, A. & Friston, K. Generating meaning: active inference and the scope and limits of passive AI. Trends Cogn. Sci. 28 , 97–112 (2024). IMAGINE INTERNAL WORLD 48

  65. [65]

    & Varley, R

    Fedorenko, E. & Varley, R. Language and thought are not the same thing: evidence from neuroimaging and neurological patients: Language versus thought. Ann. N. Y. Acad. Sci. 1369 , 132–153 (2016)

  66. [66]

    Mahowald, K. et al. Dissociating language and thought in large language models. Trends Cogn. Sci. 28 , 517–540 (2024)

  67. [67]

    Y., Rambachan, A., Kleinberg, J

    Vafa, K., Chen, J. Y., Rambachan, A., Kleinberg, J. & Mullainathan, S. Evaluating the world model implicit in a generative model. Adv. Neural Inf. Process. Syst. 37 , 26941–26975 (2024)

  68. [68]

    & Rinard, M

    Jin, C. & Rinard, M. Emergent representations of program semantics in language models trained on programs. in Forty-first International Conference on Machine Learning (2024)

  69. [69]

    & Everitt, T

    Richens, J., Abel, D., Bellot, A. & Everitt, T. General agents contain world models. Preprint at https://doi.org/10.48550/arXiv.2506.01622 (2025)

  70. [70]

    Gonsalves, B. et al. Neural Evidence That Vivid Imagining Can Lead to False Remembering. Psychol. Sci. 15 , 655–660 (2004)

  71. [71]

    & Fleming, S

    Dijkstra, N. & Fleming, S. M. Subjective signal strength distinguishes reality from imagination. Nat. Commun. 14 , 1627 (2023)

  72. [72]

    Zhang, S. et al. Personalizing Dialogue Agents: I have a dog, do you have pets too? in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Gurevych, I. & Miyao, Y.) 2204–2213 (Association for Computational Linguistics, Melbourne, Australia, 2018). doi:10.18653/v1/P18-1205

  73. [73]

    & Srivastava, N

    Singhal, I. & Srivastava, N. Dynamics of mental imagery. Conscious. Cogn. 131 , 103865 (2025). IMAGINE INTERNAL WORLD 49

  74. [74]

    Jacob, B. et al. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. Preprint at https://doi.org/10.48550/arXiv.1712.05877 (2017)

  75. [75]

    He, J. et al. Does Prompt Formatting Have Any Impact on LLM Performance? Preprint at https://doi.org/10.48550/arXiv.2411.10541 (2024)

  76. [76]

    Borsboom, D. et al. Network analysis of multivariate data in psychological science. Nat. Rev. Methods Primer 1 , 1–18 (2021)

  77. [77]

    & Fried, E

    Epskamp, S. & Fried, E. I. Package ‘bootnet’. (2025)

  78. [78]

    _i” and “_c

    Buitinck, L. et al. API design for machine learning software: experiences from the scikit-learn project. in ECML PKDD Workshop: Languages for Data Mining and Machine Learning 108–122 (2013). IMAGINE INTERNAL WORLD 50 Supplementary Material Internal World Models as Imagination Networks in Cognitive Agents Saurabh Ranjan and Brian Odegaard Department of Psy...

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.