Pith. sign in

REVIEW 3 major objections 4 minor 3 cited by

Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Large language models' internal pattern geometry correlates with human frontal EEG during abstract reasoning.

desk verdict Plausible and well-controlled abstract-reasoning LLM-EEG alignment study, but the Methods are unreadable in this version; send to review once a readable copy exists. read the letter →

arxiv 2508.10057 v1 pith:IKKPM7AX submitted 2025-08-12 q-bio.NC cs.AIcs.CL

classification q-bio.NCcs.AIcs.CL
keywords largelanguagemodelsabstractreasoningrepresentationalgeometryfixation-relatedpotentialsEEGpatterncompletionneuralalignmentintermediatelayers
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models, though trained on text, organize abstract visual patterns internally the way human brains do. On an abstract-pattern-completion task, the authors report that only the largest models tested, around 70 billion parameters, reach human-comparable accuracy, and that these models also reproduce the human pattern-by-pattern difficulty ordering. The key comparison is representational: every model clusters the abstract pattern categories in its intermediate layers, and for the task-optimal layers the geometry of those clusters correlates moderately with human frontal fixation-related potentials recorded with EEG during the same task. These correlations were not found against response-locked ERPs or resting EEG, which the authors read as evidence that the alignment is tied to task-relevant encoding rather than generic brain activity. If the claim holds, LLM internal representations offer a concrete, testable window into the neural organization of abstract reasoning.

What carries the argument

The central object is representational geometry: the arrangement of internal activation patterns (for LLM layers) or neural signals (for human EEG) in a high-dimensional space, captured by the similarity structure among pattern categories. The comparison aligns LLM intermediate-layer geometries with human frontal fixation-related potentials (FRPs, the brain's electrical responses time-locked to when a participant's gaze fixates on the pattern). This machinery does the work of measuring whether two very different systems encode the same abstract categories in the same relational arrangement, beyond whether they give the same answers.

What would settle it

Run the same models on the same patterns rendered only as grids of symbols with no rule names, and compare the layer geometries with the same frontal FRPs. If the correlations fall to zero, the alignment depends on explicit verbal descriptions rather than abstract pattern representations.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLMs might mirror human brain mechanisms in abstract reasoning, not just their outputs. Mechanistically, the authors find that the representational geometry of certain intermediate layers in every LLM tested separates the abstract pattern categories, and the strength of this separation tracks task accuracy. When they compare the similarity structure of the task-optimal layers with human frontal EEG signals collected on the same pattern-completion task, they observe moderate positive correlations. Because the same comparisons with response-locked ERPs and resting EEG diverge, the authors conclude that the shared geometry is specific to the neural signal tied

Load-bearing premise

The machine-readable version of each pattern must be mentally equivalent to the visual grid humans saw; if it named the rules in words, the measured brain-model similarity could just be shared language about the patterns.

Editorial extensions

If this is right

  • If the alignment is real, LLM internal states become candidate in-silico models of human abstract reasoning, usable to generate predictions about where and when frontal EEG should encode pattern categories.
  • The correlation between category-clustering strength and task accuracy predicts that larger or better-trained models should show stronger brain-geometry alignment, a quantitative claim that can be tested across model families.
  • The divergence from response-locked ERPs and resting EEG means the effect is not a generic brain-signal correlation; future studies can use the same control design to isolate task-specific neural codes.
  • Pattern-type difficulty profiles shared by humans and the largest models provide a behavioral signature that can be used to compare future LLMs without EEG.
  • Moderate positive correlations imply the shared representational space is partial, so the paper supports convergence of structure, not identity of mechanism.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the paper's input-format details are not visible here, the strongest test is to vary how patterns are verbalized: if explicit rule names drive the correlations, the shared geometry may reflect shared linguistic description rather than shared visual abstraction; the paper does not settle this.
  • A natural extension is to present the grids as images to a vision-language model: if the frontal-FRP correlations persist across modalities, the alignment would be more clearly about abstract structure than about text.
  • The moderate size of the correlations suggests the shared space is partial; a falsifiable refinement would be to predict FRP differences between pattern categories from LLM layer distances and compare those predictions at specific time windows.
  • The pattern-completion task's category structure could also be varied systematically to see whether the geometric alignment tracks relational complexity, which would connect the result to theories of relational composition in reasoning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports a comparison between eight open-source LLMs and human participants on an abstract-pattern-completion task. It claims that: (i) only the largest models (~70B parameters) reach human-comparable accuracy; (ii) Qwen-2.5-72B and DeepSeek-R1-70B also match the human pattern-specific difficulty profile; (iii) all tested LLMs form intermediate-layer representations that cluster the abstract pattern categories; and (iv) moderate positive correlations exist between task-optimal LLM layer geometries and human frontal fixation-related potentials, with negative-control EEG measures (response-locked ERPs, resting EEG) diverging from this pattern. The authors interpret these findings as preliminary evidence of shared representational principles between biological and artificial intelligence. The supplied full text is almost entirely corrupted by an encoding error, so the Methods, Results tables, and figures cannot be read; this report is accordingly based mainly on the abstract, the decipherable fragments, and the reader's stress-test note.

Significance. If the central claim holds, the paper would provide a notable cross-modal alignment result: LLM internal geometries and human frontal FRPs would share a common organization of abstract pattern categories, with specificity to task-relevant neural signals. The design has several strengths as visible in the abstract: the use of named open-source models, a real human EEG sample, and negative-control EEG measures. The central RSA correlation is not forced by construction in the sense that nothing guarantees LLM hidden geometries must match frontal FRP geometries. However, the format-equivalence question is load-bearing: if the text prompts verbalize the rules or category labels, the reported correlations could reflect shared linguistic description rather than shared neurocognitive reasoning. Because the methods are unreadable in the supplied text, the soundness of the central claim cannot currently be evaluated.

major comments (3)
  1. [Methods (entire section)] The full text as supplied is mojibake; I cannot read the methods, results tables, or figure captions. Every load-bearing technical detail is therefore unverifiable: the exact stimulus set, the prompt templates given to the LLMs, the layer-selection procedure, the RSA construction, the EEG preprocessing and frontal FRP definition, and the statistical inference procedure. The abstract's 'moderate positive correlations' cannot be checked against the actual numbers or against the analysis pipeline. This is not a local presentation issue; it blocks evaluation of the central claim. A clean, readable version of the manuscript is required before further review.
  2. [Abstract / LLM prompt format] The central claim requires that the LLM's text input be cognitively equivalent to the visuospatial format shown to humans. If the prompts enumerate elements, state rules, or name pattern categories, then LLM intermediate layers will cluster those categories partly because the categories are present as tokens, and the RSA correlation with human frontal FRPs could reflect shared linguistic descriptions rather than shared neurocognitive principles. The negative-control EEG measures (response-locked ERPs, resting EEG) do not address this stimulus-format confound. The manuscript should include the verbatim prompts and text stimuli for all pattern types, and ideally a control condition with minimally verbalized inputs or a human text-control experiment. This is the key uncertainty identified in the stress-test note, and it remains unresolved.
  3. [RSA construction / pattern-type partition] The manuscript appears to use the same pattern-type categorical grouping both to define LLM clustering and to construct the neural representational geometries. If the RDMs are computed after averaging over items within each pattern type, the effective number of distinct conditions is small and the correlation may be inflated by the coarse partition. The authors should report whether RDMs are computed at the individual-stimulus level or the pattern-type level, state the number of conditions and trials entering each RDM cell, and describe the permutation or bootstrap procedure used to assess significance. Without this information, the strength and specificity of the 'moderate positive correlations' cannot be interpreted.
minor comments (4)
  1. [Full text encoding] The PDF/LaTeX encoding is corrupt throughout; please resupply the manuscript in a readable format. This is not a scientific criticism of the authors, but it prevents any detailed review.
  2. [Task-optimal layer selection] The notion of 'task-optimal LLM layer' is used in the abstract and appears central to the RSA analysis, but the criterion for selecting this layer is not decipherable. Please specify the selection metric (e.g., best accuracy, best clustering, best RSA) and whether it was chosen per model and per dataset.
  3. [Statistical reporting] The abstract reports 'moderate positive correlations' without confidence intervals or multiple-comparison corrections. Since many LLM layers and many EEG time windows/electrodes are likely tested, the manuscript should state the correction procedure and report effect sizes with uncertainty.
  4. [Data and code availability] The final sections mention data/code availability in corrupted text. A clear statement with repository links is needed, especially because the RSA pipeline and prompt templates must be shared for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No demonstrated circularity: the central RSA correlation is an empirical comparison between independently measured representational geometries.

full rationale

The paper's central claim is that LLM hidden-layer representational geometries correlate with human frontal FRPs during an abstract-pattern-completion task. This is a falsifiable empirical comparison: nothing in the construction guarantees that LLM hidden activations must align with the human EEG dissimilarity structure. The shared 'pattern type' partition is the experimental independent variable used to build both the LLM and human RDMs; RSA then tests whether the two dissimilarity matrices match. That is the standard RSA logic and is not circular. Negative controls (response-locked ERPs and resting EEG) provide specificity. The main potential circularity would be if the LLM prompts explicitly verbalized pattern categories or if the 'task-optimal layer' were selected by maximizing human correlation; either would make the reported correlation expected by construction. However, the supplied full text is largely mojibake and the Methods section is unreadable, so I cannot quote any statement to that effect. Under the hard rule that circularity must be exhibited, not speculated, no circular step can be identified. The prompt-format confound is a validity concern, not a demonstrated circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

This is an empirical RSA study, so the ledger is short on numbers and long on design choices. The main degrees of freedom: which LLM layer is called 'task-optimal' and how it is chosen; which EEG channels, time window, and single-trial averaging define the human FRP RDM; and the categorical partition of items into 'pattern types', which is reused for the LLM clustering, the difficulty profile, and the EEG comparison. I could not audit any of these choices because the methods text is corrupted, so the ledger is provisional from the abstract.

free parameters (3)
  • Task-optimal LLM layer index (per model) = not reported in abstract
    The reported correlations use 'task-optimal LLM layers'; the selection criterion is unstated. If the layer was selected to maximize similarity to human data or task accuracy, the correlation is conditioned on the selection.
  • FRP time window and frontal electrode set = not reported in abstract
    Determines the human RDM compared against LLM layers; choosing the window and electrodes that show the effect is a researcher degree of freedom.
  • Pattern-type categorical grouping = number of categories not reported in abstract
    The same item partition drives the LLM clustering, the human difficulty profile, and the EEG category contrasts; the granularity of the partition sets the size of the RDMs and the stability of the correlations.
assumptions (4)
  • domain assumption Representational dissimilarity matrices from LLM activations and from EEG are directly comparable via a monotonic transform.
    Central RSA premise: implies that distance geometry in high-dimensional activation spaces and in averaged EEG topographies reflects the same latent structure. Stated nowhere, assumed by the comparison.
  • domain assumption The LLM input format and the human visual format instantiate the same abstract reasoning task.
    Humans saw visual matrix-style patterns with fixation-related EEG; open LLMs process text. If the text version spells out the rules, the LLM geometry tracks the verbal description rather than visuospatial reasoning. This is the weakest load-bearing premise.
  • domain assumption FRP averages per pattern type are reliable estimates of neural representation.
    With few trials per pattern type or few participants, RDM noise would attenuate or destabilize the reported correlations; trial counts and exclusion rules are unreadable.
  • domain assumption Clustering of pattern categories in intermediate layers reflects abstract category structure rather than trivial input differences.
    Because the text strings for different pattern types differ systematically in wording, hidden-layer separation by category is partly guaranteed by the inputs. The abstract reports the clustering as a finding without, as far as can be seen, an input-shuffling or word-level baseline.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning." pith.science (2026). https://pith.science/paper/IKKPM7AX

@misc{pith2026250810057,
  author       = {Pith},
  title        = {Pith review of: Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IKKPM7AX}},
  note         = {Machine review of arXiv:2508.10057}
}
read the original abstract

This study investigates whether large language models (LLMs) mirror human neurocognition during abstract reasoning. We compared the performance and neural representations of human participants with those of eight open-source LLMs on an abstract-pattern-completion task. We leveraged pattern type differences in task performance and in fixation-related potentials (FRPs) as recorded by electroencephalography (EEG) during the task. Our findings indicate that only the largest tested LLMs (~70 billion parameters) achieve human-comparable accuracy, with Qwen-2.5-72B and DeepSeek-R1-70B also showing similarities with the human pattern-specific difficulty profile. Critically, every LLM tested forms representations that distinctly cluster the abstract pattern categories within their intermediate layers, although the strength of this clustering scales with their performance on the task. Moderate positive correlations were observed between the representational geometries of task-optimal LLM layers and human frontal FRPs. These results consistently diverged from comparisons with other EEG measures (response-locked ERPs and resting EEG), suggesting a potential shared representational space for abstract patterns. This indicates that LLMs might mirror human brain mechanisms in abstract reasoning, offering preliminary evidence of shared principles between biological and artificial intelligence.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Large Language Models Align with the Human Brain during Creative Thinking

    q-bio.NC 2026-04 unverdicted novelty 7.0 of 10

    LLMs show scaling and training-dependent alignment with human brain responses in creativity-related networks during divergent thinking tasks, measured via RSA on fMRI data.

  2. Letting the neural code speak: Automated characterization of monkey visual neurons through human language

    q-bio.NC 2026-05 unverdicted novelty 6.0 of 10

    Natural-language descriptions generated and verified through generative models and digital twins capture the selectivity of most neurons in macaque V1 and V4.

  3. Letting the neural code speak: Automated characterization of monkey visual neurons through human language

    q-bio.NC 2026-05 unverdicted novelty 6.0 of 10

    Natural language descriptions generated via a closed-loop pipeline with digital twins capture the selectivity of most neurons in macaque V1 and V4, with synthesized images driving 96% of V4 neurons into the top or bot...

Reference graph

Works this paper leans on

72 extracted references · 46 canonical work pages · cited by 2 Pith papers

  1. [1]

    S., Malhotra, G., Dujmović, M., Montero, M

    Bowers, J. S., Malhotra, G., Dujmović, M., Montero, M. L., Tsvetkov, C., Biscione, V., Puebla, G., Adolfi, F., Hummel, J. E., Heaton, R. F., Evans, B. D., Mitchell, J., and Blything, R. (2022). Deep Problems with Neural Network Models of Human Vision . The Behavioral and Brain Sciences , pages 1--74

  2. [2]

    T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M

    Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y. (2023). Sparks of artificial general intelligence: early experiments with GPT -4. arXiv:2303.12712 [cs]

  3. [3]

    and King, J.-R

    Caucheteux, C. and King, J.-R. (2022). Brains and algorithms partially converge in natural language processing. Communications Biology , 5(1):1--10. Publisher: Nature Publishing Group

  4. [4]

    M., Spadoni, A

    Caudle, M. M., Spadoni, A. D., Schiehser, D. M., Simmons, A. N., and Bomyea, J. (2023). Neural activity and network analysis for understanding reasoning using the matrix reasoning task. Cognitive Processing , 24(4):585--594

  5. [5]

    Y., Shamosh, N

    Choi, Y. Y., Shamosh, N. A., Cho, S. H., DeYoung, C. G., Lee, M. J., Lee, J.-M., Kim, S. I., Cho, Z.-H., Kim, K., Gray, J. R., and Lee, K. H. (2008). Multiple bases of human intelligence revealed by cortical thickness and neural activation. Journal of Neuroscience , 28(41):10323--10329

  6. [6]

    Chollet, F. (2019). On the measure of intelligence. arXiv:1911.01547 [cs]

  7. [7]

    Chuderski, A. (2022). Fluid intelligence emerges from representing relations. Journal of Intelligence , 10(3):51

  8. [8]

    and Liversedge, S

    Degno, F. and Liversedge, S. P. (2020). Eye movements and fixation-related potentials in reading: a review. Vision , 4(1):11

Show all 72 references
  1. [9]

    C., Janarthanan, S., Culham, J

    Dima, D. C., Janarthanan, S., Culham, J. C., and Mohsenzadeh, Y. (2024). Shared representations of human actions across vision and language. Neuropsychologia , 202:108962

  2. [10]

    C., Allen, E., Wu, Y., Naselaris, T., Kay, K., and Charest, I

    Doerig, A., Kietzmann, T. C., Allen, E., Wu, Y., Naselaris, T., Kay, K., and Charest, I. (2024). Visual representations in the human brain are aligned with large language models. arXiv:2209.11737 version: 2

  3. [11]

    Duncan, J. (2010). The multiple-demand ( MD ) system of the primate brain: mental programs for intelligent behaviour. Trends in Cognitive Sciences , 14(4):172--179

  4. [12]

    Engbert, R. (2006). Microsaccades: a microcosm for research on oculomotor control, attention, and visual perception. In Martinez-Conde, S., Macknik, S. L., Martinez, L. M., Alonso, J. M., and Tse, P. U., editors, Progress in Brain Research , volume 154 of Visual Perception , p...

  5. [13]

    A., and Kao, J

    Feghhi, E., Hadidi, N., Song, B., Blank, I. A., and Kao, J. C. (2024). What Are Large Language Models Mapping to in the Brain ? A Case Against Over - Reliance on Brain Scores . arXiv:2406.01538 [cs] version: 1

  6. [14]

    D., and Bunge, S

    Ferrer, E., O'Hare, E. D., and Bunge, S. A. (2009). Fluid reasoning and the developing brain. Frontiers in Neuroscience , 3(1):46--51

  7. [15]

    Gawin, C., Sun, Y., and Kejriwal, M. (2025). Navigating semantic relations: challenges for language models in abstract common-sense reasoning. arXiv:2502.14086 [cs]

  8. [16]

    Gendron, G., Bao, Q., Witbrock, M., and Dobbie, G. (2024). Large language models are not strong abstract reasoners. arXiv:2305.19555 [cs]

  9. [17]

    Gluth, S., Rieskamp, J., and Büchel, C. (2013). Classic EEG motor potentials track the emergence of value-based decisions. Neuroimage , 79:394--403

  10. [18]

    Goldstein, A., Grinstein-Dabush, A., Schain, M., Wang, H., Hong, Z., Aubrey, B., Schain, M., Nastase, S. A., Zada, Z., Ham, E., Feder, A., Gazula, H., Buchnik, E., Doyle, W., Devore, S., Dugan, P., Reichart, R., Friedman, D., Brenner, M., Hassidim, A., Devinsky, O., Flinker, A...

  11. [19]

    A., Strohmeier, D., Brodbeck, C., Goj, R., Jas, M., Brooks, T., Parkkonen, L., and Hämäläinen, M

    Gramfort, A., Luessi, M., Larson, E., Engemann, D. A., Strohmeier, D., Brodbeck, C., Goj, R., Jas, M., Brooks, T., Parkkonen, L., and Hämäläinen, M. (2013). MEG and EEG data analysis with MNE -python. Frontiers in Neuroscience , 7:267

  12. [20]

    R., Chabris, C

    Gray, J. R., Chabris, C. F., and Braver, T. S. (2003). Neural mechanisms of general fluid intelligence. Nature Neuroscience , 6(3):316--322. Publisher: Nature Publishing Group

  13. [21]

    Hersche, M., Camposampiero, G., Wattenhofer, R., Sebastian, A., and Rahimi, A. (2024). Towards learning to reason: comparing LLMs with neuro-symbolic on arithmetic relations in abstract reasoning. arXiv:2412.05586 [cs] version: 1

  14. [22]

    Huh, M., Cheung, B., Wang, T., and Isola, P. (2024). The platonic representation hypothesis. arXiv:2405.07987 [cs]

  15. [23]

    Iaia, C., Choksi, B., Wiebers, E., Roig, G., and Fiebach, C. J. (2025). The representational alignment between humans and language models is implicitly driven by a concreteness effect. arXiv:2505.15682 [cs] version: 1

  16. [24]

    Ju, T., Sun, W., Du, W., Yuan, X., Ren, Z., and Liu, G. (2024). How large language models encode context knowledge? A layer-wise probing study. arXiv:2402.16061 [cs]

  17. [25]

    Kriegeskorte, N., Mur, M., and Bandettini, P. A. (2008). Representational similarity analysis - connecting the branches of systems neuroscience. Frontiers in Systems Neuroscience , 2. Publisher: Frontiers

  18. [26]

    Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012). ImageNet Classification with Deep Convolutional Neural Networks . In Advances in Neural Information Processing Systems , volume 25. Curran Associates, Inc

  19. [27]

    Lan, M., Torr, P., Meek, A., Khakzar, A., Krueger, D., and Barez, F. (2025). Quantifying feature space universality across large language models via sparse autoencoders. arXiv:2410.06981 [cs] version: 4

  20. [28]

    Langlois, D., Chartier, S., and Gosselin, D. (2010). An introduction to independent component analysis: InfoMax and FastICA algorithms. Tutorials in Quantitative Methods for Psychology , 6(1):31--38

  21. [29]

    LeCun, Y., Bengio, Y., and Hinton, G. (2015). Deep learning. Nature , 521(7553):436--444. Publisher: Nature Publishing Group

  22. [30]

    Lee, S., Sim, W., Shin, D., Seo, W., Park, J., Lee, S., Hwang, S., Kim, S., and Kim, S. (2025). Reasoning abilities of large language models: In -depth analysis on the abstraction and reasoning corpus. ACM Trans. Intell. Syst. Technol. , page 3712701. Just Accepted

  23. [31]

    and Cooper, S

    Lei, G. and Cooper, S. J. (2025). The representation and recall of interwoven structured knowledge in LLMs : a geometric and layered analysis. arXiv:2502.10871 [cs] version: 1

  24. [32]

    Lei, Y., Ge, X., Zhang, Y., Yang, Y., and Ma, B. (2025). Do large language models think like the brain? Sentence -level evidence from fMRI and hierarchical embeddings. arXiv:2505.22563 [cs]

  25. [33]

    and Mitchell, M

    Lewis, M. and Mitchell, M. (2024). Evaluating the robustness of analogical reasoning in large language models. arXiv:2411.14215 [cs]

  26. [34]

    P., Höchenberger, R., and Scheltienne, M

    Li, A., Feitelberg, J., Saini, A. P., Höchenberger, R., and Scheltienne, M. (2022). MNE - ICALabel : automatically annotating ICA components with ICLabel in python. Journal of Open Source Software , 7(76):4484

  27. [35]

    D., Vasconcelos, N., Golan, T., Luo, D., and Deng, H

    Li, Y., Gao, Q., Zhao, T., Wang, B., Sun, H., Lyu, H., Hawkins, R. D., Vasconcelos, N., Golan, T., Luo, D., and Deng, H. (2025). Core knowledge deficits in multi-modal language models. arXiv:2410.10855 [cs]

  28. [36]

    Liang, S., Garg, S., and Moghaddam, R. Z. (2025). The SWE -bench illusion: when state-of-the-art LLMs remember instead of reason. arXiv:2506.12286 [cs]

  29. [37]

    K., Nunez, M

    Lui, K. K., Nunez, M. D., Cassidy, J. M., Vandekerckhove, J., Cramer, S. C., and Srinivasan, R. (2021). Timing of readiness potentials reflect a decision-making process in the human brain. Computational Brain & Behavior , 4(3):264--283

  30. [38]

    Marjieh, R., Sucholutsky, I., van Rijn, P., Jacoby, N., and Griffiths, T. L. (2024). Large language models predict human sensory judgments across six modalities. Scientific Reports , 14(1):21445. Publisher: Nature Publishing Group

  31. [39]

    T., Yao, S., Friedman, D., Hardy, M

    McCoy, R. T., Yao, S., Friedman, D., Hardy, M. D., and Griffiths, T. L. (2024). Embers of autoregression show how large language models are shaped by the problem they are trained to solve. Proceedings of the National Academy of Sciences , 121(41):e2322420121. Publisher: Procee...

  32. [40]

    Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2023). Locating and editing factual associations in GPT . arXiv:2202.05262 [cs]

  33. [41]

    A., Bickel, S., Mehta, A

    Mischler, G., Li, Y. A., Bickel, S., Mehta, A. D., and Mesgarani, N. (2024). Contextual Feature Extraction Hierarchies Converge in Large Language Models and the Brain

  34. [42]

    B., and Moskvichev, A

    Mitchell, M., Palmarini, A. B., and Moskvichev, A. (2023). Comparing humans, GPT -4, and GPT - 4V on abstraction and reasoning tasks. arXiv:2311.09247 [cs]

  35. [43]

    Musker, S., Duchnowski, A., Millière, R., and Pavlick, E. (2025). LLMs as models for analogical reasoning. arXiv:2406.13803 [cs] version: 2

  36. [44]

    Newell, A. (1955). The chess machine: an example of dealing with a complex task by adaptation. In Proceedings of the March 1-3, 1955, Western Joint Computer Conference , AFIPS '55 ( Western ), pages 101--108, New York, NY, USA. Association for Computing Machinery

  37. [45]

    and Simon, H

    Newell, A. and Simon, H. (1956). The logic theory machine–a complex information processing system. IRE Transactions on Information Theory , 2(3):61--79

  38. [46]

    D., Watts, D

    Nguyen, T. D., Watts, D. J., and Whiting, M. E. (2025). Empirically evaluating commonsense intelligence in large language models with large-scale human judgments. arXiv:2505.10309 [cs]

  39. [47]

    Nili, H., Wingfield, C., Walther, A., Su, L., Marslen-Wilson, W., and Kriegeskorte, N. (2014). A toolbox for representational similarity analysis. PLoS computational biology , 10(4):e1003553

  40. [48]

    Palmarini, A. B. and Mitchell, M. (2024). Abstract understanding of core-knowledge concepts: Humans vs. llms. In ICML 2024 Workshop on LLMs and Cognition

  41. [49]

    and Kim, G

    Park, H. and Kim, G. (2025). Where do LLMs encode the knowledge to assess the ambiguity? In Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert, S., Darwish, K., and Agarwal, A., editors, Proceedings of the 31st International Conference on Comput...

  42. [50]

    L., and Onofrj, M

    Perfetti, B., Saggino, A., Ferretti, A., Caulo, M., Romani, G. L., and Onofrj, M. (2007). Differential patterns of cortical activation as a function of fluid reasoning complexity. Human Brain Mapping , 30(2):497--510

  43. [51]

    Perrin, F., Pernier, J., Bertrand, O., and Echallier, J. F. (1989). Spherical splines for scalp potential and current density mapping. Electroencephalography and Clinical Neurophysiology , 72(2):184--187

  44. [52]

    Santarnecchi, E., Emmendorfer, A., and Pascual-Leone, A. (2017). Dissecting the parieto-frontal correlates of fluid intelligence: a comprehensive ALE meta-analysis study. Intelligence , 63:9--28

  45. [53]

    A., Tuckute, G., Kauf, C., Hosseini, E

    Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., and Fedorenko, E. (2021). The neural architecture of language: integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences of the U...

  46. [54]

    Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., a...

  47. [55]

    Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D. (2018). A general reinforcement learning algorithm that masters chess, shogi, and go through self-play....

  48. [56]

    Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D. (2017). Mastering the game of go without human knowle...

  49. [57]

    R., LeCun, Y., and Shwartz-Ziv, R

    Skean, O., Arefin, M. R., LeCun, Y., and Shwartz-Ziv, R. (2024). Does representation matter? Exploring intermediate layers in large language models. arXiv:2412.09563 [cs]

  50. [58]

    R., Zhao, D., Patel, N., Naghiyev, J., LeCun, Y., and Shwartz-Ziv, R

    Skean, O., Arefin, M. R., Zhao, D., Patel, N., Naghiyev, J., LeCun, Y., and Shwartz-Ziv, R. (2025). Layer by layer: uncovering hidden representations in language models. arXiv:2502.02013 [cs]

  51. [59]

    Sourati, Z., Ilievski, F., Sommerauer, P., and Jiang, Y. (2024). ARN : analogical reasoning on narratives. arXiv:2310.00996 [cs]

  52. [60]

    E., Pafford, A., Maas, H

    Stevenson, C. E., Pafford, A., Maas, H. L. J. v. d., and Mitchell, M. (2025). Can large language models generalize analogy solving like people can? arXiv:2411.02348 [cs]

  53. [61]

    Sun, W., Song, X., Li, P., Yin, L., Zheng, Y., and Liu, S. (2025). The curse of depth in large language models. arXiv:2502.05795 [cs]

  54. [62]

    Tschentscher, N., Mitchell, D., and Duncan, J. (2017). Fluid intelligence predicts novel rule implementation in a distributed frontoparietal control network. Journal of Neuroscience , 37(18):4841--4847

  55. [63]

    Wang, Y., Chen, W., Han, X., Lin, X., Zhao, H., Liu, Y., Zhai, B., Yuan, J., You, Q., and Yang, H. (2024). Exploring the reasoning abilities of multimodal large language models ( MLLMs ): a comprehensive survey on emerging trends in multimodal reasoning. arXiv:2401.06805 [cs]

  56. [64]

    J., and Lu, H

    Webb, T., Holyoak, K. J., and Lu, H. (2023). Emergent Analogical Reasoning in Large Language Models

  57. [65]

    W., Holyoak, K

    Webb, T. W., Holyoak, K. J., and Lu, H. (2025). Evidence from counterfactual tasks supports emergent analogical reasoning in large language models. PNAS Nexus , 4(5):pgaf135

  58. [66]

    and Huckle, J

    Williams, S. and Huckle, J. (2024). Easy problems that LLMs get wrong. arXiv:2405.19616 [cs]

  59. [67]

    and Schein, A

    Wolfram, C. and Schein, A. (2025). Layers at similar depths generate similar activations across LLM architectures. arXiv:2504.08775 [cs]

  60. [68]

    Yang, Y., Chen, M., Liu, Q., Hu, M., Chen, Q., Zhang, G., Hu, S., Zhai, G., Qiao, Y., Wang, Y., Shao, W., and Luo, P. (2025). Truly assessing fluid intelligence of large language models through dynamic reasoning evaluation. arXiv:2506.02648 [cs] version: 1

  61. [69]

    Yax, N., Anlló, H., and Palminteri, S. (2024). Studying and improving reasoning in humans and machines. Communications Psychology , 2(1):51. Publisher: Nature Publishing Group

  62. [70]

    Zhang, Y., Dong, Y., and Kawaguchi, K. (2024). Investigating layer importance in large language models. In Proceedings of the 7th Blackboxnlp Workshop : Analyzing and Interpreting Neural Networks for NLP , pages 469--479, Miami, Florida, US. Association for Computational Linguistics

  63. [71]

    Zhang, Z., Guo, S., Zhou, W., Luo, Y., Zhu, Y., Zhang, L., and Li, L. (2025). Brain-model neural similarity reveals abstractive summarization performance. Scientific Reports , 15(1):370. Publisher: Nature Publishing Group

  64. [72]

    Zurrin, R., Wong, S. T. S., Roes, M. M., Percival, C. M., Chinchani, A., Arreaza, L., Kusi, M., Momeni, A., Rasheed, M., Mo, Z., Goghari, V. M., and Woodward, T. S. (2024). Functional brain networks involved in the raven's standard progressive matrices task and their relation ...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.