Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

Large Language Models and Emergence: A Complex Systems Perspective

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Genuine emergence in LLMs would require internal coarse-grained representations that are causally responsible for behavior; the paper argues current evidence falls short.

desk verdict A useful, well-written position paper arguing that LLM emergence claims need internal coarse-grained causal evidence, but the central definition is stipulative and one key citation is unverifiable. read the letter →

arxiv 2506.11135 v1 pith:6IPX4PWS submitted 2025-06-10 cs.CL cs.AIcs.LGcs.NE

classification cs.CLcs.AIcs.LGcs.NE
keywords emergencelargelanguagemodelscoarse-grainingeffectivetheoriesscalinglawsemergentintelligencemechanisticinterpretabilitycomplexsystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that 'emergence' has a precise meaning in complex systems: new coarse-grained variables and compressed descriptions that screen off microscopic details. The authors review three kinds of LLM emergence claims—sudden benchmark jumps with scale, capabilities that were not explicitly trained for, and internal world models—and conclude that none has conclusively shown the necessary internal coarse-grained representation that is causally responsible for the behavior. They further distinguish emergent capability from emergent intelligence, arguing that intelligence is efficient, compact, analogy-based problem solving, or 'less is more,' and that LLMs have not demonstrated this. If the paper is right, the burden of proof shifts: claims of LLM emergence, and especially emergent intelligence, require evidence from inside the network, not just external performance.

What carries the argument

The central object is a definitional criterion: coarse-graining coupled with compression as the minimum necessary signature of emergence, operationalized through five principles (scaling, criticality, compression, novel bases and manifolds, and generalization). The argument works by applying this criterion to LLMs and by adding the knowledge-out/knowledge-in distinction, which states that in systems whose complexity comes from complex inputs rather than simple components, macroscopic behavior alone cannot justify an emergence claim. This criterion carries the paper's conclusion: because no LLM study has yet identified causally responsible coarse-grained internal variables, all current emergence claims are incomplete.

What would settle it

One concrete check is to pick a widely cited emergent benchmark jump, then use causal interventions such as ablation or activation patching to test whether a compact low-dimensional internal representation exists that is responsible for the jump and screens off the remaining weights; a single confirmed case would show LLM emergence in the paper's sense, while repeated failures across many tasks would support the paper's position. A second check would measure the minimal description length of the network's effective internal variables before and after the jump, with no reduction indicating no emergence.

Watch

Extended reading notes

Core claim

The paper's central claim is that the term 'emergence' should be reserved for systems in which successful task performance is accompanied by new coarse-grained internal representations—reduced internal degrees of freedom—that are causally sufficient to explain or predict behavior and that screen off details of weights and activations. On this definition, surprising accuracy jumps, untrained abilities, and even internal world models are at best incomplete evidence unless the relevant coarse-grained variables are identified and shown to do causal work. The paper proposes five mechanisms from complexity science—scaling, criticality, compression, novel bases and manifolds, and generalization—as quantitative signatures of genuine emergence, and it argues that LLMs are knowledge-in systems where external behavior alone cannot establish emergence. It concludes that LLMs at present demonstrate emergent capability at most, not emergent intelligence, because intelligence is defined as the efficient internal use of coarse-graining to solve a broad range of problems through analogy and minimal modification, a low-bandwidth property that LLMs give no evidence of possessing.

Load-bearing premise

The load-bearing premise is that genuine emergence necessarily requires internal coarse-grained representations that screen off microscopic details; if external behavioral discontinuities were accepted as sufficient evidence, the paper's critique of LLM emergence claims would lose much of its force.

Editorial extensions

If this is right

  • Benchmark discontinuities or untrained abilities will not count as emergence unless accompanied by evidence of new coarse-grained internal representations with causal force.
  • Scaling laws and phase-transition analogies for LLMs are weakened, because the control variable called scale is high-dimensional and no internal phase has been demonstrated.
  • World-model claims such as the OthelloGPT case remain unproven until the internal model is shown to be both parsimonious and causally responsible for predictions.
  • The distinction between emergent capability and emergent intelligence implies that efficiency, compactness, and analogy-making, not raw capability, are the relevant evidence for intelligence claims.
  • If language encodes a complete representation of the world, emergence claims become even weaker because the model's internal degrees of freedom would simply converge on external degrees of freedom through training.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper's criterion is adopted, a testable prediction follows: candidate emergent abilities in LLMs will not be traceable to compact, low-dimensional internal mechanisms; where such mechanisms are found, they will correspond to a small set of features detectable by causal intervention.
  • A natural extension would measure energy or parameter efficiency per unit of generalization, turning the paper's 'less is more' definition of intelligence into a quantitative metric that could be compared across models and humans.
  • The same framework could be applied to other large AI models, such as vision transformers and reinforcement-learning agents, where emergence claims are also made on behavioral evidence alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This perspective/position paper argues that the term 'emergence' as used in the LLM literature has been too loose. The authors propose that genuine emergence requires more than surprising or discontinuous external performance: it requires identifying internal coarse-grained variables that form effective mechanisms and screen off microscopic details (Section 6.1). They review existing LLM emergence claims, articulate five conditions for emergence (scaling, criticality, compression, novel bases, generalization), introduce a distinction between knowledge-out and knowledge-in emergence, and evaluate candidate evidence such as OthelloGPT. They conclude that current evidence for emergent LLM capabilities is incomplete and that emergent intelligence in LLMs is not demonstrated.

Significance. The paper's main value is conceptual: it makes explicit that external benchmark jumps are not by themselves evidence of emergence, and it offers a framework for seeking internal mechanistic evidence. The authors are appropriately cautious about asserting LLM emergence and give credit to prior critical work. However, the central definition is stipulated rather than defended, and the paper contains no new quantitative analysis, so its critical force is conditional on accepting a specific philosophical stance. If the definitional issues are addressed, the framework could be a useful reference for researchers studying claimed emergent abilities in large language models.

major comments (4)
  1. [Section 3] The claim that coarse-graining coupled with predictive compression is 'a minimum and necessary signature of any emergence claim' is asserted, not argued. Because the paper itself concedes that 'There is no single accepted definition of emergence', the reader needs a justification for privileging the effective-theory account over canonical alternatives such as weak emergence, strong/weak emergence as formulated by Chalmers, or computational emergence, under which macro-level regularities can be emergent without the system implementing internal coarse-grained variables. This is load-bearing: the critique of Wei et al. and of benchmark discontinuities in Sections 5.1 and 6.1 depends on rejecting external-behavior evidence as insufficient. Reframing the conclusion as conditional on the effective-theory notion, or defending the definition, would remove the circularity concern.
  2. [Section 5.1] The paper's most positive example of scale-induced emergence rests on 'Guth & Ménard (2025, in prep)', which is not verifiable. The specific claim that covariance spectra of weights transition from exponential to scale-free at the peak of double descent is presented as 'more persuasive' evidence, yet no preprint, data, or reference is supplied. The authors should either replace this with a publicly available or peer-reviewed source, or explicitly flag it as anecdotal and remove it from the evidence chain.
  3. [Section 4] In the discussion of Table 1, the assertion that 'In none of the examples can emergence be claimed based purely on macroscopic properties' is not supported by argument or citation. This is exactly the contested point: weak emergence and similar positions permit macroscopic regularities to be emergent even when they are not internally represented as coarse-grained models. Without supporting this assertion, the knowledge-out/knowledge-in distinction does not do the work the paper assigns to it.
  4. [Section 6.1] The paper states that, for cases (1) and (2), 'None of these assumptions has been conclusively verified.' This is appropriately cautious, but the manuscript sometimes slips from 'not verified' to 'therefore not emergent', particularly in the OthelloGPT discussion in Section 5.2. The conclusion should be framed consistently as insufficient evidence rather than as a demonstrated absence of the phenomenon.
minor comments (5)
  1. [Throughout] There are numerous typos and wording errors, e.g., 'has lead to' (Section 1), 'disucss' (Section 1), 'pizoelectic' (Section 3), 'casually sufficient' (Section 4), 'the the conditions' (Section 6.1), and inconsistent capitalization of 'LLMS' (Section 4).
  2. [Section 1] 'Sterling's formula' should be 'Stirling's formula'.
  3. [Section 5.1] In the sentence about double descent, 'double descent provide no evidence for emergence' should be 'double descent provides no evidence for emergence'.
  4. [Section 5.4] Reference [50] has a typo in its title: 'Data Nisualization' should be 'Data Visualization'.
  5. [References] Some substantive points rely on non-archival web sources, such as the LessWrong post [47] used to qualify the OthelloGPT discussion; please cite peer-reviewed or archival versions where available.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's emergence criterion is an explicitly stated and externally cited definitional premise, not a hidden self-referential derivation.

full rationale

This is a conceptual/position paper with no fitted parameters, numerical predictions, or equations whose outputs are encoded in their inputs. The central claim—that LLM emergence claims are incomplete without evidence of causally responsible internal coarse-grained representations—is presented as a stipulated definitional framework, not as an empirically derived result. Section 3 explicitly concedes 'There is no single accepted definition of emergence' and then proposes a 'minimum and necessary signature' with citations to external literature (Israeli & Goldenfeld 2006; Chalmers 2006; Carroll & Parola 2024; Anderson 1972). The critique of Wei et al. and of benchmark discontinuities is a transparent conditional argument: if one adopts this operationalization, then external performance jumps alone are insufficient. That is an explicit and contestable premise, not a disguised tautology, and the paper's own conclusion is hedged ('None of these assumptions has been conclusively verified'). First-author prior works [16, 22, 23, 31] are cited for conceptual scaffolding and the KO/KI taxonomy, but the evaluation of LLM evidence is independently argued using external studies (Wei et al., Schaeffer et al., Li et al., Nanda et al., Lin et al., Vafa et al.), and no uniqueness theorem or ansatz is imported from those self-citations to forbid alternatives. The only questionable citation is the unpublished 'Guth & Ménard (2025, in prep)' example, but that is an illustrative example rather than a load-bearing derivation. The contested status of the coarse-graining definition is a philosophical correctness risk, not circularity. Therefore no circular step is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's argument rests on definitional axioms about emergence and intelligence rather than on fitted parameters or new entities. No free parameters are used. The axioms are domain assumptions drawn from the authors' prior conceptual framework and are not independently justified in the paper.

assumptions (3)
  • domain assumption Genuine emergence requires a coarse-grained internal description that screens off microscopic details.
    Asserted in Section 3 as a minimum and necessary signature of any emergence claim; contested in the philosophy and complexity science literature.
  • domain assumption Human intelligence is an emergent property characterized by efficient, parsimonious, analogy-based problem solving.
    Assumed in Section 6.2 without a formal definition or empirical proof; underpins the separation of capability from intelligence.
  • domain assumption Performance discontinuities in LLM benchmarks do not by themselves constitute emergence.
    Central to the argument, but presented as a definitional stance rather than an empirical finding.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Language Models and Emergence: A Complex Systems Perspective." pith.science (2026). https://pith.science/paper/6IPX4PWS

@misc{pith2026250611135,
  author       = {Pith},
  title        = {Pith review of: Large Language Models and Emergence: A Complex Systems Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6IPX4PWS}},
  note         = {Machine review of arXiv:2506.11135}
}
read the original abstract

Emergence is a concept in complexity science that describes how many-body systems manifest novel higher-level properties, properties that can be described by replacing high-dimensional mechanisms with lower-dimensional effective variables and theories. This is captured by the idea "more is different". Intelligence is a consummate emergent property manifesting increasingly efficient -- cheaper and faster -- uses of emergent capabilities to solve problems. This is captured by the idea "less is more". In this paper, we first examine claims that Large Language Models exhibit emergent capabilities, reviewing several approaches to quantifying emergence, and secondly ask whether LLMs possess emergent intelligence.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    MMM packages knowledge as typed vertices, edges and pens with free-text labels under a small set of normative rules, enabling structural interoperability and decentralisation without semantic convergence.

  2. Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

    cs.CY 2026-07 accept novelty 4.0 of 10

    AI safety is a systems-governance problem: six recurring organizational failure patterns from past disasters remain unlearned in AI development, so component-level fixes like benchmarks and alignment cannot deliver safety.

Reference graph

Works this paper leans on

77 extracted references · 64 canonical work pages · cited by 2 Pith papers

  1. [1]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models.Transac- tions on Machine Learning Research, 2022

  2. [2]

    Emergent abilities in large language models: A survey.arXiv:2503.05788, 2025

    Leonardo Berti, Flavio Giorgi, and Gjergji Kasneci. Emergent abilities in large language models: A survey.arXiv:2503.05788, 2025

  3. [3]

    Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Tren- ton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thomp- son, Sam Zimmerman, Kelley ...

  4. [4]

    Preventing the immense increase in the life-cycle energy and carbon foot- prints of llm-powered intelligent chatbots.Engineering, 40:202–210, 2024

    Peng Jiang, Christian Sonne, Wangliang Li, Fengqi You, and Siming You. Preventing the immense increase in the life-cycle energy and carbon foot- prints of llm-powered intelligent chatbots.Engineering, 40:202–210, 2024

  5. [5]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv:2001.08361, 2020

  6. [6]

    Are emergent abil- ities of large language models a mirage?Advances in Neural Information Processing Systems, 36:55565–55581, 2023

    Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abil- ities of large language models a mirage?Advances in Neural Information Processing Systems, 36:55565–55581, 2023

  7. [7]

    Are emergent abilities in large language models just in-context learning? InProceedings of The 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), 2024

    Sheng Lu, Irina Bigoulaeva, Rachneet Sachdeva, Harish Tayyar Madabushi, and Iryna Gurevych. Are emergent abilities in large language models just in-context learning? InProceedings of The 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), 2024

  8. [8]

    Numeracy from literacy: Data science as an emergent skill from large language models.arXiv:2301.13382, 2023

    David Noever and Forrest McKee. Numeracy from literacy: Data science as an emergent skill from large language models.arXiv:2301.13382, 2023. 14

Show all 77 references
  1. [9]

    Holyoak, and Hongjing Lu

    Taylor Webb, Keith J. Holyoak, and Hongjing Lu. Emergent analogical reasoning in large language models.Nature Human Behaviour, 7(9):1526– 1541, 2023

  2. [10]

    Nay, DavidKaramardian, Sarah B.Lawsky, WentingTao, Meghana Bhat, Raghav Jain, Aaron Travis Lee, Jonathan H

    John J. Nay, DavidKaramardian, Sarah B.Lawsky, WentingTao, Meghana Bhat, Raghav Jain, Aaron Travis Lee, Jonathan H. Choi, and Jungo Ka- sai. Large language models as tax attorneys: A case study in legal ca- pabilities emergence.Philosophical Transactions of the Royal Society A...

  3. [11]

    Emergent world representations: Explor- ing a sequence model trained on a synthetic task

    Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Explor- ing a sequence model trained on a synthetic task. InProceedings of the International Conference on Learning Representations (ICLR-2023), 2023

  4. [12]

    Luccioni

    Anna Rogers and Alexandra S. Luccioni. Position: Key claims in LLM re- search have a long tail of footnotes. InProceedings of the 41st International Conference on Machine Learning (ICML 2024), 2024

  5. [13]

    Bedau and Paul Humphreys.Emergence: Contemporary Readings in Philosophy and Science

    Mark A. Bedau and Paul Humphreys.Emergence: Contemporary Readings in Philosophy and Science. MIT Press, March 2008

  6. [14]

    Hendry, and Tom Lancaster.The Routledge Hand- book of Emergence

    Sophie Gibb, Robin F. Hendry, and Tom Lancaster.The Routledge Hand- book of Emergence. Routledge, March 2019

  7. [15]

    Chalmers

    David J. Chalmers. Strong and weak emergence. In Phillip Claton and Paul Davies, editors,The Re-Emergence of Emergence. Oxford University Press, New York, 2006

  8. [16]

    Flack, and Nihat Ay

    David Krakauer, Nils Bertschinger, Eckehard Olbrich, Jessica C. Flack, and Nihat Ay. The information theory of individuality.Theory in Biosciences, 139:209–223, 2020

  9. [17]

    Carroll and Achyuth Parola

    Sean M. Carroll and Achyuth Parola. What emergence can possibly mean. arXiv:2410.15468, October 2024

  10. [18]

    Butterworth-Heinemann, 2015

    Jiri Blazek.Computational Fluid Dynamics: Principles and Applications. Butterworth-Heinemann, 2015

  11. [19]

    Coarse-graining of cellular automata, emergence, and the predictability of complex systems.Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, 73(2):026203, 2006

    Navot Israeli and Nigel Goldenfeld. Coarse-graining of cellular automata, emergence, and the predictability of complex systems.Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, 73(2):026203, 2006

  12. [20]

    P. W. Anderson. More is different.Science, 177(4047):393–396, August 1972

  13. [21]

    Salmon.Statistical Explanation and Statistical Relevance, vol- ume 69

    Wesley C. Salmon.Statistical Explanation and Statistical Relevance, vol- ume 69. University of Pittsburgh Press, 2010

  14. [22]

    Krakauer

    David C. Krakauer. Symmetry—simplicity, broken symmetry—complexity. Interface Focus, 13(3):20220075, 2023

  15. [23]

    Krakauer.The Complex World

    David C. Krakauer.The Complex World. Santa Fe Institute Press, Santa Fe, NM, 2024

  16. [24]

    Feynman, Robert B

    Richard P. Feynman, Robert B. Leighton, and Matthew L. Sands.The Feynman Lectures on Physics. Addison-Wesley Publishing Company, 1963. 15

  17. [25]

    Stuart A. Newman. Development and evolution: The physics connection. In Alan C. Love, editor,Conceptual Change in Biology: Scientific and Philosophical Perspectives on Evolution and Development, pages 421–440. Springer Netherlands, Dordrecht, 2015

  18. [26]

    Stigmergy as a universal coordination mechanism I: Definition and components.Cognitive Systems Research, 38:4–13, 2016

    Francis Heylighen. Stigmergy as a universal coordination mechanism I: Definition and components.Cognitive Systems Research, 38:4–13, 2016

  19. [27]

    A. M. Turing. The chemical basis of morphogenesis.Philosophical Trans- actions of the Royal Society, 237(641):37–72, 1952

  20. [28]

    Reynolds

    Craig W. Reynolds. Flocks, herds and schools: A distributed behavioral model. InProceedings of the 14th Annual Conference on Computer Graph- ics and Interactive Techniques, 1987

  21. [29]

    R. Linsker. Perceptual neural organization: Some approaches based on network models and information theory.Annual Reviews of Neuroscience, 13(1):257–281, 1990

  22. [30]

    J. J. Hopfield. Neural networks and physical systems with emergent collec- tive computational abilities.Proc. Natl. Acad. Sci. U. S. A., 79(8):2554– 2558, April 1982

  23. [31]

    Krakauer

    David C. Krakauer. Unifying complexity science and machine learning. Frontiers in Complex Systems, 1:1235202, 2023

  24. [32]

    Solé.Phase Transitions

    Ricard V. Solé.Phase Transitions. Primers in Complex Systems. Princeton University Press, Princeton, NJ, 2011

  25. [33]

    Flack, Doug Erwin, Tanya Elliot, and David C

    Jessica C. Flack, Doug Erwin, Tanya Elliot, and David C. Krakauer. Timescales, symmetry, and uncertainty reduction in the origins of hier- archy in biological systems.Evolution Cooperation and Complexity, pages 45–74, 2013

  26. [34]

    Schuster.Criticality in Neural Systems

    Heinz G. Schuster.Criticality in Neural Systems. John Wiley & Sons, 2014

  27. [35]

    O’Dwyer, Steven W

    James P. O’Dwyer, Steven W. Kembel, and Thomas J. Sharpton. Back- bones of evolutionary history test biodiversity theory for microbes.Pro- ceedings of the National Academy of Sciences, 112(27):8356–8361, 2015

  28. [36]

    Brown and Geoffrey B

    James H. Brown and Geoffrey B. West.Scaling in Biology. Oxford Univer- sity Press, 2000

  29. [37]

    Oxford University Press, 2018

    Stefan Thurner, Rudolf Hanel, and Peter Klimek.Introduction to the The- ory of Complex Systems. Oxford University Press, 2018

  30. [38]

    Kempes, M

    Christopher P. Kempes, M. A. R. Koehl, and Geoffrey B. West. The scales that limit: The physical boundaries of evolution.Frontiers in Ecology and Evolution, 7:242, 2019

  31. [39]

    Deep double descent: Where bigger models and more data hurt.Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124003, 2021

    Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep double descent: Where bigger models and more data hurt.Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124003, 2021. 16

  32. [40]

    A u-turn on double descent: Rethinking parameter counting in statistical learning.Ad- vances in Neural Information Processing Systems, 36:55932–55962, 2023

    Alicia Curth, Alan Jeffares, and Mihaela van der Schaar. A u-turn on double descent: Rethinking parameter counting in statistical learning.Ad- vances in Neural Information Processing Systems, 36:55932–55962, 2023

  33. [41]

    Rocks, Ila Rani Fiete, and Oluwasanmi Koyejo

    Rylan Schaeffer, Mikail Khona, Zachary Robertson, Akhilan Boopathy, Kateryna Pistunova, Jason W. Rocks, Ila Rani Fiete, and Oluwasanmi Koyejo. Double descent demystified: Identifying, interpreting & ablating the sources of a deep learning puzzle.arXiv preprint arXiv:2303.14151, 2023

  34. [42]

    Leavitt, and Naomi Saphra

    Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, and Naomi Saphra. Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in mlms.arXiv preprint arXiv:2309.07311, 2023

  35. [43]

    Critical phase transitioninlargelanguagemodels.arXiv preprint arXiv:2406.05335, 2024

    Kai Nakaishi, Yoshihiko Nishikawa, and Koji Hukushima. Critical phase transitioninlargelanguagemodels.arXiv preprint arXiv:2406.05335, 2024

  36. [44]

    Olshausen and David J

    Bruno A. Olshausen and David J. Field. Sparse coding of sensory inputs. Current Opinion in Neurobiology, 14(4):481–487, 2004

  37. [45]

    Charles F. Stevens. What the fly’s nose tells the fly’s brain.Proceedings of the National Academy of Sciences, 112(30):9460–9465, 2015

  38. [46]

    Emergent lin- ear representations in world models of self-supervised sequence models

    Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent lin- ear representations in world models of self-supervised sequence models. arXiv:2309.00941, 2023

  39. [47]

    J. Y. Lin, Jack S., Adam Karvonen, and Can Rager. OthelloGPT learned a bag of heuristics, 2024.https://www.lesswrong.com/posts/ gcpNuEZnxAPayaKBY/othellogpt-learned-a-bag-of-heuristics-1

  40. [48]

    Self-organization and the selection of pinwheel density in visual cortical development.New Journal of Physics, 10(1):015009, 2008

    Matthias Kaschube, Michael Schnabel, and Fred Wolf. Self-organization and the selection of pinwheel density in visual cortical development.New Journal of Physics, 10(1):015009, 2008

  41. [49]

    Foundations and formalizations of self-organization.Ad- vances in Applied Self-Organizing Systems, pages 19–37, 2008

    Daniel Polani. Foundations and formalizations of self-organization.Ad- vances in Applied Self-Organizing Systems, pages 19–37, 2008

  42. [50]

    Learning nonlinear principal manifolds by self-organising maps

    Hujun Yin. Learning nonlinear principal manifolds by self-organising maps. InPrincipal Manifolds for Data Nisualization and Dimension Reduction, pages 68–95. Springer, 2008

  43. [51]

    Crutchfield

    Adam Rupe and James P. Crutchfield. On principles of emergent organi- zation.Physics Reports, 1071:1–47, 2024

  44. [52]

    Ferry, Joshua Ching, and Takashi Kawai

    Quentin RV. Ferry, Joshua Ching, and Takashi Kawai. Emergence and function of abstract representations in self-supervised transformers.arXiv preprint arXiv:2312.05361, 2023

  45. [53]

    A rainbow in deep network black boxes.Journal of Machine Learning Research, 25(350):1–59, 2024

    Florentin Guth, Brice Ménard, Gaspar Rochette, and Stéphane Mallat. A rainbow in deep network black boxes.Journal of Machine Learning Research, 25(350):1–59, 2024

  46. [54]

    Measuring the vc- dimension of a learning machine.Neural Computation, 6(5):851–876, 1994

    Vladimir Vapnik, Esther Levin, and Yann LeCun. Measuring the vc- dimension of a learning machine.Neural Computation, 6(5):851–876, 1994. 17

  47. [55]

    Basic Books, 2013

    Leslie Valiant.Probably Approximately Correct: Nature’s Algorithms for Learning and Prospering in a Complex World. Basic Books, 2013

  48. [56]

    How do we know how smart AI systems are?Science, 381(6654):eadj5957, 2023

    Melanie Mitchell. How do we know how smart AI systems are?Science, 381(6654):eadj5957, 2023

  49. [57]

    GSM-Symbolic: Understand- ing the limitations of mathematical reasoning in large language models

    Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. GSM-Symbolic: Understand- ing the limitations of mathematical reasoning in large language models. arXiv:2410.05229, 2024

  50. [58]

    Evaluatingtherobustnessofanalogical reasoning in GPT models.Transactions on Machine Learning Research, 2024

    MarthaLewisandMelanieMitchell. Evaluatingtherobustnessofanalogical reasoning in GPT models.Transactions on Machine Learning Research, 2024

  51. [59]

    Toni J. B. Liu, Nicolas Boullé, Raphaël Sarfati, and Christopher J. Earls. Llms learn governing principles of dynamical systems, revealing an in- context neural scaling law.arXiv preprint arXiv:2402.00795, 2024

  52. [60]

    What is a macrostate? sub- jective observations and objective dynamics.Found

    Cosma Rohilla Shalizi and Cristopher Moore. What is a macrostate? sub- jective observations and objective dynamics.Found. Phys., 55(1), February 2025

  53. [61]

    Evaluating the world model implicit in a generative model

    Keyon Vafa, Justin Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. Evaluating the world model implicit in a generative model. Advances in Neural Information Processing Systems, 37:26941–26975, 2024

  54. [62]

    Will digital intelligence replace biological intelligence?,

    Geoffrey Hinton. Will digital intelligence replace biological intelligence?,

  55. [63]

    On the measure of intelligence.arXiv:1911.01547, 2019

    François Chollet. On the measure of intelligence.arXiv:1911.01547, 2019

  56. [64]

    Princeton University Press, 2014

    Joseph Mazur.Enlightening symbols: A short history of mathematical no- tation and its hidden powers. Princeton University Press, 2014

  57. [65]

    Princeton University Press, 2025

    RaúlRojas.The Language of Mathematics: The Stories behind the Symbols. Princeton University Press, 2025

  58. [66]

    Folk psychology as a theory, 2021

    Daniel Hutto and Ian Ravenscroft. Folk psychology as a theory, 2021

  59. [67]

    Abstract representations emerge in human hippocampal neurons during inference.Nature, 632(8026):841–849, 2024

    Hristos S Courellis, Juri Minxha, Araceli R Cardenas, Daniel L Kimmel, Chrystal M Reed, Taufik A Valiante, C Daniel Salzman, Adam N Mamelak, Stefano Fusi, and Ueli Rutishauser. Abstract representations emerge in human hippocampal neurons during inference.Nature, 632(8026):841–...

  60. [68]

    An implicit plan overrides an explicit strategy during visuomotor adaptation.Journal of neuroscience, 26(14):3642–3645, 2006

    Pietro Mazzoni and John W Krakauer. An implicit plan overrides an explicit strategy during visuomotor adaptation.Journal of neuroscience, 26(14):3642–3645, 2006

  61. [69]

    Learning of a sequential motor skill comprises explicit and implicit components that consolidate differently.Journal of neurophys- iology, 101(5):2218–2229, 2009

    M Felice Ghilardi, Clara Moisello, Giulia Silvestri, Claude Ghez, and John W Krakauer. Learning of a sequential motor skill comprises explicit and implicit components that consolidate differently.Journal of neurophys- iology, 101(5):2218–2229, 2009. 18

  62. [70]

    When money is not enough: awareness, success, and variability in motor learning.PLoS One, 9(1):e86580, 2014

    Harry Manley, Peter Dayan, and Jörn Diedrichsen. When money is not enough: awareness, success, and variability in motor learning.PLoS One, 9(1):e86580, 2014

  63. [71]

    Representationincognitivesciencebynicholasshea: but is it thinking? the philosophy of representation meets systems neuroscience, 2022

    JohnWKrakauer. Representationincognitivesciencebynicholasshea: but is it thinking? the philosophy of representation meets systems neuroscience, 2022

  64. [72]

    Intelligent data analysis of intelligent systems

    David C Krakauer, Jessica C Flack, Simon DeDeo, Doyne Farmer, and Daniel Rockmore. Intelligent data analysis of intelligent systems. InIn- ternational Symposium on Intelligent Data Analysis, pages 8–17. Springer, 2010

  65. [73]

    Einstein

    Lincoln Barnett and Albert Einstein.The Universe and Dr. Einstein. Courier Corporation, 2005

  66. [74]

    The science of brute force.Com- munications of the ACM, 60(8):70–79, 2017

    Marijn JH Heule and Oliver Kullmann. The science of brute force.Com- munications of the ACM, 60(8):70–79, 2017

  67. [75]

    The MIT Press, 1996

    Harold Abelson and Gerald Jay Sussman.Structure and interpretation of computer programs. The MIT Press, 1996. 19

  68. [2024]

    Romanes Lecture, Oxford, UK,https://www.youtube.com/watch? v=N1TEjTeQeg0

  69. [2025]

    pub/2025/attribution-graphs/biology.html

    Transformer Circuits Threadhttps://transformer-circuits. pub/2025/attribution-graphs/biology.html

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.