Pith. sign in

REVIEW 3 major objections 5 minor 4 references

The brain versus AI: World-model-based versatile circuit computation underlying diverse functions in the neocortex and cerebellum

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper argues that the neocortex and cerebellum are general-purpose world-model circuits that predict the future from past inputs, learn from prediction errors, and reuse those models for understanding and generation.

desk verdict A good, broad synthesis that repackages existing theories as P-U-G, but its one load-bearing premise about cerebellar nucleocortical feedback being 'essential' is asserted, not demonstrated. read the letter →

arxiv 2411.16075 v2 pith:TJKPFOQN submitted 2024-11-25 q-bio.NC cs.AI

classification q-bio.NCcs.AI
keywords neocortexcerebellumprediction-errorlearningworldmodelspredictivecodingrecurrentneuralnetworksinternalmirrorneuronsystem
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review paper argues that the neocortex and cerebellum perform one underlying computation despite their outwardly different jobs. The proposed computation is prediction-error-based world modeling: each structure predicts future states of the external world from past inputs, learns by minimizing prediction errors, and thus builds compact internal models of the world. Those models are used in three ways: to predict what comes next, to understand sensory input through compressed abstract representations, and to generate outputs such as actions, sentences, or imitations by reusing the prediction machinery. The paper claims this single circuit computation explains the remarkable diversity of cortical and cerebellar functions across sensory, cognitive, and motor domains, and that recent general-purpose AI trained on next-word or next-frame prediction has converged on the same learning principle. A sympathetic reader would care because the theory turns a mystery—how uniform circuits support diverse functions—into a concrete, testable account.

What carries the argument

The central object is the prediction-error-learning recurrent neural network (RNN), a circuit that takes past inputs, predicts the next input, and updates its weights to reduce the prediction error; the paper treats this as the universal motif of circuit computation. In the neocortex the concrete form is the deep predictive-coding circuit, in which each layer transmits only unpredicted 'newsworthy' error to the higher layer and receives top-down compressed predictions, with error signals generated and consumed locally. In the cerebellum the concrete form is a three-layer RNN that mirrors granule–Purkinje–nucleus connectivity, including nucleocortical feedback loops, with the inferior olive supplying the prediction-error teaching signal. The paper also uses a three-element decomposition of circuit computation—circuit structure, input/output signals, and learning algorithm—as its comparison scheme for aligning brain circuits with AI circuits across sensory, cognitive, and motor domains.

What would settle it

Silence or block the nucleocortical feedback pathway in the right lateral cerebellum during a next-word-prediction task: if next-word prediction and syntactic processing remain intact, the claim that recurrent nucleocortical feedback is essential to the unified circuit computation is false.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that both the neocortex and the cerebellum are world-model learners: they predict future world states from past information and learn from prediction errors, and the resulting internal models support Prediction, Understanding, and Generation. The paper identifies these three processes as the universal modes of circuit computation: generating future information, interpreting the external world via compressed and abstracted sensory information, and repurposing the future-information generation mechanism to produce other outputs such as action plans, language plans, and imitation. In the cerebellum, the supporting evidence is a brain-imitating three-layer recurrent circuit—input (granule) cells, Purkinje cells, and output (nucleus) cells with nucleocortical feedback—that, when trained only to predict the next word, spontaneously developed syntactic processing. In the neocortex, the supporting evidence is the convergence of predictive-coding circuits, recurrent neural networks trained on video prediction, and transformer language models trained by next-word prediction whose internal signals align with human cortical language signals. The paper concludes that the diverse functions of the neocortex and cerebellum are not separate modules but three uses of one world-model computation.

Load-bearing premise

The account stands or falls on the claim that the cerebellum's nucleocortical feedback projections are functionally essential, making the cerebellum a true three-layer recurrent network; if those loops are merely incidental, the unified cerebellar world-model theory collapses.

Editorial extensions

If this is right

  • If the theory is correct, visual recognition, language comprehension and production, and motor control in the neocortex and cerebellum are not separate modules but three modes—Prediction, Understanding, Generation—of a single world-model computation.
  • Cerebellar language functions (next-word prediction and syntactic processing) should be unified in one circuit computation, and the same three-layer recurrent framework should also account for forward and inverse internal models in motor control.
  • Reward-related signals observed in neocortex and cerebellum during conditioning should be reinterpreted, at least in part, as world-model prediction errors or event representations rather than exclusively as model-free reinforcement learning signals.
  • Prediction-error learning, treated in the paper as unsupervised learning, should be sufficient in principle to explain acquisition of language and action models without innate task-specific structures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: The theory predicts that disrupting next-word prediction in the cerebellum should impair not only comprehension but also sentence generation, because Generation reuses the same prediction machinery; this is a direct consequence the authors largely leave for future work.
  • Editorial inference: If the unified account is right, one should find transformer-like self-attention signatures within local neocortical circuits, not only at a coarse area level; the paper raises this as an open question rather than asserting it.
  • Editorial inference: The three-element comparison scheme implies a sharper criterion for 'brain-like' AI: two circuits should be judged similar only when structure, input/output, and learning align, so future models that match on one element but not others should be treated as partial analogues.
  • Editorial inference: A testable extension would be to train a cerebellar-architecture network on both word sequences and motor sequences and ask whether the same intermediate representation supports syntax and movement prediction; the paper's framework suggests it should.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This review-style paper proposes a unified theory of neocortical and cerebellar computation: both structures predict future world events from past inputs, learn from prediction errors, and thereby acquire world models that support Prediction, Understanding, and Generation. The argument is organized around a three-element comparison of circuit structure, input/output signals, and learning algorithms between the brain and modern AI, with emphasis on prediction-error-learning RNNs, transformers, and world-model-based reinforcement learning. The paper claims that this P-U-G framework explains how uniform cortical and cerebellar circuits achieve diverse sensory, cognitive, and motor functions.

Significance. If the central claim is correct, the paper would provide a genuinely unifying perspective on neocortical and cerebellar function, connecting internal-model theory, predictive coding, mirror-neuron theory, and modern large-scale AI in a single framework. The three-element rubric (circuit structure, input/output, learning) is a useful organizing device, and the review is broad and current, covering vision, language, motor control, and reinforcement learning. The authors are also explicit that their theory should be testable, and they point to concrete empirical anchors such as next-word prediction in language areas and feedback pathways in the cerebellum. However, the paper is not a derivation or a test of the theory; its support is largely analogical and selective, and several load-bearing premises are asserted rather than established. The main value is as a synthetic perspective that generates specific hypotheses, not as a demonstrated account.

major comments (3)
  1. [§1.2, Fig. 3a] The sentence 'These feedback projections are essential for the predictive functions of the cerebellum' (citing refs. 137, 142, 143) overstates what the cited studies show. Gao et al. 2016 and Xiao et al. 2023 demonstrate that nucleocortical and pontine feedback amplifies or facilitates associative learning, and Ohmae et al. 2021 is a preprint proposing a recurrent-circuit mechanism; none of these studies selectively eliminates the feedback pathway and shows a loss of predictive function. Because the unified P-U-G account for the cerebellum depends on the three-layer RNN being the actual circuit rather than a convenient model, the authors should either provide loss-of-function evidence or soften the claim to 'contributes to' and explicitly state that the recurrent architecture is a modeling hypothesis. A concrete test would be to ablate the recurrent (nucleocortical) connections in the authors' own next-word-prediction simulation (ref. 145) and determine whether syntactic processing disappears; without such a test, the 'essential' claim is unsupported.
  2. [Discussion, 'Proposal for a new theory'] The theory's three categories are not operationally defined, which makes the universal claim that all diverse neocortical and cerebellar functions arise from Prediction, Understanding, and Generation difficult to falsify. For instance, any output that is not a literal future-state prediction can be labeled 'Generation' because it 'repurposes' the prediction mechanism, and any compressed representation can be labeled 'Understanding.' The paper does not specify independent neural or behavioral criteria for assigning a function to one category, nor does it state which observations would count against the theory. Please provide explicit operational definitions (e.g., in terms of predicted variables, error signals, and output modalities) and at least one disconfirmable prediction, such as a specific circuit manipulation that should abolish Understanding but not Prediction.
  3. [§1.1–1.2, Discussion] The 'convergent evolution' argument partly relies on AI systems that were explicitly designed to mimic brain theories, so the inference from AI success to brain computation is partially circular. For example, PredNet is an implementation of predictive coding, and transformer attention was designed in reference to neocortical attention; their performance cannot independently confirm that the neocortex performs these computations. The paper should explicitly separate 'brain-inspired by design' from 'convergent without design' and identify at least one major AI success that was not engineered around a brain theory (next-word prediction in transformers is a plausible candidate) to support the convergence claim. This is a methodological concern about the evidence, not a rejection of the theory.
minor comments (5)
  1. [§1.1, 'Direct comparison of object-recognition processing'] 'Yanis and colleagues' should read 'Yamins and colleagues' (ref. 66).
  2. [§2.1, 'Sentence generation in motor language processing'] 'GPT reopposes the word-prediction circuit to word generation' should read 'repurposes' rather than 'reopposes.'
  3. [Discussion, 'Significance of scaling up the circuit size'] The parenthetical 'Sutton, "The Bitter Success"' should be 'The Bitter Lesson' (see ref. 277).
  4. [§1.2, 'A theory of circuit computation in the cerebellum'] The description of the authors' artificial network (3000 input neurons, Purkinje cells, recurrent pathway) is presented in a review without stating that the full implementation and analysis are in ref. 145; please add an explicit pointer so readers can verify the simulation and its limitations.
  5. [§1.2, Discussion] The relabeling of cerebellar prediction-error learning as 'unsupervised learning' conflicts with standard machine-learning terminology, since next-word prediction uses a target signal and the inferior olive provides a teaching signal; a short definition or table would clarify the intended meaning.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the P-U-G theory is an integrative proposal grounded in external evidence and published simulations, not a derivation that reduces to its own inputs.

full rationale

The paper is a review and theory proposal rather than a derivation with fitted parameters or equations. Its central Prediction-Understanding-Generation theory is explicitly presented as an integrative extension of established internal-model, predictive-coding, and mirror-neuron theories, and it draws on a broad external literature for the component claims, including predictive coding experiments, fMRI and ECoG comparisons between language areas and next-word-prediction AI, lesion and stimulation studies of cerebellar prediction, and motor-control studies of forward and inverse models. The only load-bearing self-references are the authors' own cerebellar RNN simulation (ref. 145) and a bioRxiv study (ref. 142). These are used to support two claims: that recurrent nucleocortical feedback is 'essential' for cerebellar predictive functions, and that training a cerebellar-like three-layer RNN on next-word prediction yields emergent syntactic processing. Neither use is circular by construction: the simulation is a published, independent computational result with stated architectural assumptions that do not include the syntactic-classification outcome, and the 'essential' claim, while overstated relative to the cited experiments, is an evidentiary weakness rather than a definitional equivalence. No prediction in the paper is a renamed fitted parameter, and no result is forced by self-citation alone. The paper is therefore not significantly circular; the minor self-citations raise at most a small evidentiary caution, not a circularity score above 1.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The theory adds no free parameters or new entities; its main burden is a set of domain assumptions about brain learning rules and the transferability of AI analogies.

assumptions (4)
  • domain assumption The neocortex and cerebellum use prediction-error learning as their primary learning rule.
    Core to the P-U-G theory; while supported by predictive coding and cerebellar learning literature (Sections 1.1, 2.3), it is not proven for all functions and is debated.
  • domain assumption The cerebellum can be described as a three-layer RNN with functionally important nucleocortical feedback.
    Invoked in Section 1.2 and Figure 3a; depends on recent anatomical evidence for feedback projections, and the functional significance is still under investigation.
  • ad hoc to paper Convergence between large-scale AI trained by prediction-error and brain signals implies shared underlying computation.
    The paper treats AI convergence as evidence for brain mechanisms (Sections 1.2, 2.1, Discussion), but this is an analogical inference, not direct brain evidence.
  • domain assumption The neocortex has a uniform circuit structure across areas.
    Standard motif in the paper's framing (from Douglas and Martin 2004), but area-specific variation exists.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The brain versus AI: World-model-based versatile circuit computation underlying diverse functions in the neocortex and cerebellum." pith.science (2026). https://pith.science/paper/TJKPFOQN

@misc{pith2026241116075,
  author       = {Pith},
  title        = {Pith review of: The brain versus AI: World-model-based versatile circuit computation underlying diverse functions in the neocortex and cerebellum},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TJKPFOQN}},
  note         = {Machine review of arXiv:2411.16075}
}
read the original abstract

AI's significant recent advances using general-purpose circuit computations offer a potential window into how the neocortex and cerebellum of the brain are able to achieve a diverse range of functions across sensory, cognitive, and motor domains, despite their uniform circuit structures. However, comparing the brain and AI is challenging unless clear similarities exist, and past reviews have been limited to comparison of brain-inspired vision AI and the visual neocortex. Here, to enable comparisons across diverse functional domains, we subdivide circuit computation into three elements -- circuit structure, input/outputs, and the learning algorithm -- and evaluate the similarities for each element. With this novel approach, we identify wide-ranging similarities and convergent evolution in the brain and AI, providing new insights into key concepts in neuroscience. Furthermore, inspired by processing mechanisms of AI, we propose a new theory that integrates established neuroscience theories, particularly the theories of internal models and the mirror neuron system. Both the neocortex and cerebellum predict future world events from past information and learn from prediction errors, thereby acquiring models of the world. These models enable three core processes: (1) Prediction -- generating future information, (2) Understanding -- interpreting the external world via compressed and abstracted sensory information, and (3) Generation -- repurposing the future-information generation mechanism to produce other types of outputs. The universal application of these processes underlies the ability of the neocortex and cerebellum to accomplish diverse functions with uniform circuits. Our systematic approach, insights, and theory promise groundbreaking advances in understanding the brain.

Figures

Figures reproduced from arXiv: 2411.16075 by the authors.

Figure 1
Figure 1. Deep-layered autoencoders and CNNs as theories of visual information processing in the neocortex. a, Autoencoder circuit. In a representative autoencoder circuit, the input neurons receive a 2D visual image (e.g., the pixel information of a photograph). The restoration neurons (downstream via disynaptic projections) are trained to encode the same information as the input neurons (In 1996, Olshausen & Field trained t… view at source ↗
Figure 2
Figure 2. Prediction-error-learning RNNs as a processing theory for dynamic visual information in the neocortex. a, b, Predictive coding circuit. a, The error signal (from the prediction-error neurons) is central not only for learning but also for information flow in the circuit. a, To establish a deep-layered structure, each layer (corresponding to the entire circuit in a) sends a copy of the error signal to the next higher … view at source ↗
Figure 7
Figure 7. Lottery ticket hypothesis. This hypothesis about neural network learning suggests that a pre-learning subnetwork that happens to have appropriate initial values suitable for learning continues to be reinforced and eventually becomes responsible for the information processing of the entire circuit. For example, in a three-layer 3-6-3 network the number of 3-3-3 subnetworks is Combination(6, 3) = 20; as the circuit sc… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

4 extracted references · 4 canonical work pages

  1. [1]

    AI gijyutsu no saizensen

    1 Saxe, A., Nelli, S. & Summerfield, C. If deep learning is the answer, what is the question? Nat Rev Neurosci 22, 55-67, doi:10.1038/s41583-020-00395-8 (2021). 2 Diedrichsen, J., King, M., Hernandez-Castillo, C., Sereno, M. & Ivry, R. B. Universal Transform or Multiple Functionality? Understanding the Contribution of the Human Cerebellum across Task Doma...

  2. [2]

    The dramatic success of CNN in object recognition from visual images. From 2010 to 2012, traditional image recognition methods that did not use neural networks (blue dots; each dot represents the performance of each team in a competition) struggled to surpass the 75% accuracy barrier (thick gray line). In 2012, AlexNet of Hinton group overwhelmingly won t...

  3. [3]

    Comparison of intermediate processing stages between CNNs trained with unsupervised vs supervised learning. a. A CNN trained with unsupervised autoencoder learning. Greedy Layer-wise Training was conducted at each layer of a three- layer CNN for natural image input. Intermediate processing was examined by visualizing the optimal stimulus (image input that...

  4. [4]

    looking ahead

    The design philosophy and network structure of Alpha GO. After DQN (2015), DeepMind developed Alpha GO (2016) and Alpha GO Zero (2017) to conquer GO, a game thought to be impossible to evaluate using reinforcement learning due to its vast search space. DeepMind achieved this by deepening the Deep Q Network and combining it with Monte Carlo Tree Search, wh...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.