Pith. sign in

REVIEW 4 major objections 6 minor 66 references

Beyond Attention: Toward Machines with Intrinsic Higher Mental States

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Triadic Q-K-V loops before attention let a Transformer learn faster with near-linear cost.

desk verdict A plausible latent-query attention variant whose headline claim about the triadic modulation loop is not actually tested, because every result varies the loop and the latent queries together. read the letter →

arxiv 2505.06257 v1 pith:ZKKDXYWO submitted 2025-05-02 cs.LG cs.AIcs.NE

classification cs.LGcs.AIcs.NE
keywords Co4triadicmodulationtransformerattentionlatentqueriespyramidalneuronmodelmentalstateslinearcomplexityreinforcementlearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a Transformer can learn much faster and with far fewer heads, layers, and tokens if it first runs triadic modulation loops among latent questions, clues, and hypotheses before applying attention. The proposed Co4 mechanism is inspired by two integration sites in neocortical pyramidal neurons, where contextual input amplifies only what is coherent with the current signal. If the claim holds, attention-based models could approach or exceed standard Transformer accuracy while cutting the asymptotic cost from $O(N^2)$ to roughly $O(N)$. The paper reports quick proof-of-concept results on reinforcement learning, CIFAR-10, and a synthetic bAbI question-answering task.

What carries the argument

The central object is the Cooperation Equation, $Cooperation(R,C)=ReLU6(R^2+2R+C(1+|R|))$, an asynchronous modulatory transfer function used by three populations of two-point-neuron-inspired units for $Q$, $K$, and $V$. In this function the contextual field $C$ acts as the driving force, amplifying coherent responses and attenuating incoherent ones regardless of the raw signal $R$'s strength. Triadic loops apply this function cyclically: $K$ and $V$ provide contextual input to $Q$; $Q$ and $V$ shape $K$; $Q$ and $K$ together drive $V$. This pre-attention co-adaptation is what embeds reasoning into a single lightweight attention step and removes the need for a feedforward residual block.

What would settle it

Run Co4 with fixed $L=4$ or $L=8$ on sequences of increasing length, for example synthetic bAbI-style stories with 60, 600, and 6000 tokens, and record accuracy and wall-clock cost; if accuracy drops when $L$ is held fixed and restoring accuracy requires $L$ to grow with $N$, the paper's approximate $O(N)$ claim is false.

Watch

Extended reading notes

Core claim

The paper's central claim is that pre-attention triadic modulation among latent queries ($Q$), evolving keys ($K$), and evolving values ($V$) enables deep, parallel, representation-level reasoning that standard attention only approximates after many layers. Using a two-point-neuron-inspired Cooperation Equation, Co4 lets the latent question tokens adapt to the clues and hypotheses, lets the keys adapt to who is asking, and lets the values emerge from the coherent interaction of questions and clues. The final attention matrix is only $L \times N$, where $L$ is a small fixed number of latent tokens (4 or 8 in the experiments), so the dominant cost becomes about $O(N)$. Across the reported tasks, Co4 reaches higher scores or accuracies in fewer epochs, often with one layer and one head, while the standard Transformer needs more depth and width to approach comparable performance.

Load-bearing premise

The speed and near-linear cost claims rest on the assumption that a small fixed number of latent question tokens (4 or 8 in the experiments) remains sufficient as the input gets longer and tasks get harder; the paper reports no scaling experiment showing that $L$ can stay small as $N$ grows.

Editorial extensions

If this is right

  • With a small fixed $L$, the attention cost becomes about $O(N)$ instead of $O(N^2)$, so longer sequences cost far less at the same parameter count.
  • Co4 removes the feedforward residual block from the Transformer block, eliminating one of the most expensive per-layer components.
  • Shallow one- or two-layer Co4 models reach accuracies that standard Transformers need more heads and layers to approach, implying the same reasoning can be compressed into fewer resources.
  • On the synthetic bAbI-style task, Co4 reaches 96-98% accuracy with one or two layers and one or two heads, while the standard Transformer tops out near 77% in the reported setting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: a direct scaling test that fixes $L$ and grows $N$ across several orders of magnitude would decide whether the approximate $O(N)$ cost claim holds in practice.
  • Beyond the paper: the Cooperation Equation could be inserted as a context-gating layer in CNNs or RNNs, testing whether the speed-up comes from modulation in general or from its specific interaction with attention.
  • Beyond the paper: ablating the triadic loop to a one-shot conditioning step would reveal how much of the gain comes from bidirectional Q-K-V co-adaptation rather than from the extra nonlinearity alone.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes Co4, a transformer variant that replaces the standard Q-K-V attention computation with triadic modulation loops among latent query tokens (Q), keys (K), and values (V), using a nonlinear "Cooperation Equation" (Eq. 1) inspired by context-sensitive pyramidal-cell physiology. The claimed benefits are faster learning, fewer required heads/layers/tokens, and near-linear O(N) complexity. Experiments are reported on three RL tasks (CartPole, PyBullet Ant, CarRacing), CIFAR-10 image classification, and a synthetic bAbI-style QA benchmark, always comparing Co4 against standard Transformer baselines. The paper concludes that Co4 embeds representation-level reasoning and may be a step toward machines with intrinsic higher mental states, while acknowledging that the current evidence is from quick proof-of-concept tests.

Significance. If the central claim were established, Co4 would be a practically relevant contribution: a transformer variant that achieves competitive accuracy with substantially fewer heads, layers, and tokens, at near-linear attention cost. The paper is also valuable for explicitly connecting cellular neuroscience (apical amplification, two-point neurons) to a concrete architectural mechanism, and for testing the idea across RL, vision, and QA settings. Credit is due for the explicit complexity analysis, for comparing against established baselines, and for clearly stating that the experiments are quick-to-run proof-of-concept tests rather than a full evaluation. However, the current evidence does not yet justify the headline claims: the experimental design confounds the latent-query mechanism with the Eq. (1) modulation, the linear-complexity claim depends on an untested assumption, and the empirical results lack repeated trials, error bars, and code. The contribution is therefore promising but preliminary.

major comments (4)
  1. [Section 4, Tables 3-4] The central causal claim that triadic modulation loops 'enable ... faster learning' is confounded, because every Co4 configuration changes two ingredients at once: (i) the latent query mechanism (L learnable query tokens interacting with N input tokens) and (ii) the Eq. (1) nonlinear modulation applied to Q, K, and V. The Transformer baselines use input-derived queries, and no control is run that sets C=0 or replaces Eq. (1) by the identity while holding L, heads, layers, parameter count, epochs, and seeds fixed. Since Perceiver ([2]) already shows that a latent-query bottleneck alone gives O(LN) attention and can accelerate learning on small tasks, the reported gains are equally consistent with 'latent queries plus ordinary attention.' An ablation is required to attribute the results to the triadic modulation.
  2. [Section 4.1] The O(N) complexity claim rests on the assumption that the number of latent query tokens L (or L_q) is a small, fixed fraction of the input length, but no experiment varies N while holding L fixed, and no argument is given that L does not need to grow with task complexity or sequence length. If L grows with N, the attention cost O(L*N) is no longer linear. The MAC comparison also needs clarification: the standard-Transformer formula includes the FF-network term PE^2, yet Section 5.1 states that all results are presented without the second residual (FF) connection; it should be stated explicitly whether the Transformer baselines include or exclude the FF block, and the complexity comparison should be made on the same architectural footing.
  3. [Section 5, Tables 3-4, Figure 6] The empirical support is too thin to substantiate 'orders-of-magnitude faster learning with significantly reduced computational demand.' Tables 3 and 4 report single validation accuracies without error bars or seeds; Figure 6 shows training curves without wall-clock time or FLOPs measurements; and no code or configuration files are released. The paper itself acknowledges that these are 'quick-to-run tests' with a more detailed analysis to follow, so the quantitative claims in the abstract should be either softened or supported by multiple seeds and direct timing/complexity measurements.
  4. [Section 5.3, Table 3] The statement that Co4 uses 'the same architecture and number of parameters' as the Transformer baseline is not substantiated. In Table 3 the parameter counts are identical for Transformer and Co4 (e.g., 0.215M, 0.61M, 1.20M), but Co4 adds learnable latent query vectors and Eq. (1) operations; if the counts are indeed identical, the baseline must contain compensating parameters somewhere. The paper should specify the exact module composition of both models, including the presence or absence of the feedforward residual block, so that the comparison is interpretable.
minor comments (6)
  1. [Section 5.2, Eq. (8)] Equation (8) contains a parenthesis error: 'C(1 + |R(Ot,a t−1))|)' should read 'C(1 + |R(Ot,a t−1)|)'.
  2. [Section 4, Eq. (1)] Reference [66], 'Cooperation is All You Need', is cited only in the Acknowledgments but appears to be the source of the Cooperation equation; it should be cited at Eq. (1) and the relation to that work should be discussed in the text.
  3. [Section 4, Figure 3 caption] The phrase 'representation-level Chain-of-Thought' and 'silent yet deep form of Chain-of-Thought' is presented as a factual property, but the paper provides no analysis of reasoning chains; the claim should be labeled as an analogy or supported by evidence.
  4. [Section 5.4, Tables and text] The dataset name is spelled inconsistently as 'bAbI' and 'bABi'; please use one consistent spelling throughout.
  5. [Section 4.1] The symbol L is used both for the number of layers and for the number of latent query tokens, which makes the complexity formulas confusing; use one symbol for layers and a distinct symbol (e.g., L_q) for latent tokens consistently.
  6. [Figure 6] The text refers to 'Figure 6 (bottom right)' for the CarRacing result, but the panel is not labeled in the caption; please label each subplot.

Circularity Check

0 steps flagged · score 2.0 of 10

No construction-level circularity; Eq. (1) is fully specified and baselines are external, though benchmark-guided selection and a latent-query confound weaken the causal claim.

full rationale

This is an empirical architecture paper rather than a formal derivation, and I find no step in which a claimed result is equivalent by construction to an input or fitted quantity. The central nonlinearity, Eq. (1), is explicitly defined in the manuscript, so the cited prior work (e.g., refs. [45], [66]) serves as provenance rather than load-bearing support: a reader could implement Co4 without consulting those papers. The empirical comparisons are made against external baselines (standard Transformer, CCT, and the PI-agent baselines of [60]), not against quantities constructed from Co4's own fitted values. Two concerns are real but are not circularity. First, the paper states that Eq. (1) was 'guided by empirical experiments' with alternative nonlinearities; this is benchmark-guided model selection, and it weakens the strength of the reported advantage but does not make the later benchmark results a formal restatement of a fitted prediction. Second, all Co4 experiments vary the latent-query mechanism together with the triadic modulation loop, so the causal role of the loop is not isolated; this is an omitted-control problem, not a circular reduction. The self-citations are minor and non-load-bearing. Therefore the circularity score is low.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claims rest on a small number of hand-chosen design elements: the empirical selection of eq (1), the assumption of a fixed small latent-query count, and the belief that triadic modulation preserves task-relevant information. No new physical entities are postulated.

free parameters (1)
  • Latent query count Lq = 4 for bAbI, 8 for CIFAR-10; unspecified for RL
    The O(N*L) to O(N) reduction depends on Lq being small and fixed; the paper chooses it per task (Tables 3 and 4) without a scaling study.
assumptions (4)
  • ad hoc to paper The Cooperation equation (eq. 1), ReLU6(R^2 + 2R + C(1 + |R|)), is an appropriate model of context-sensitive pyramidal-cell modulation.
    The paper states the function was chosen after empirical experiments with several nonlinear RF functions (Section 4, after eq. 5); it is not derived from the cellular theory or from first principles.
  • domain assumption A small fixed number of latent query tokens Lq suffices to represent the questions needed for arbitrary input lengths.
    The complexity section assumes Lq is small relative to N; no scaling experiments support this (Section 4.1).
  • domain assumption The triadic Q-K-V modulation loops preserve or improve the information needed for downstream attention relative to standard Q-K similarity.
    The paper asserts that modulation pre-selects relevant information, but provides no information-theoretic or formal argument; this is the mechanism behind the claimed gains (Section 4).
  • domain assumption Neurobiological mapping of mental states to RF/CF strength ranges is a valid design guide.
    The paper relies on [4] and related literature for the claim that pyramidal-cell states correspond to perceptual and imaginative processing; this is a background assumption for the narrative, not for the benchmarks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Attention: Toward Machines with Intrinsic Higher Mental States." pith.science (2026). https://pith.science/paper/ZKKDXYWO

@misc{pith2026250506257,
  author       = {Pith},
  title        = {Pith review of: Beyond Attention: Toward Machines with Intrinsic Higher Mental States},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZKKDXYWO}},
  note         = {Machine review of arXiv:2505.06257}
}
abstract

Attending to what is relevant is fundamental to both the mammalian brain and modern machine learning models such as Transformers. Yet, determining relevance remains a core challenge, traditionally offloaded to learning algorithms like backpropagation. Inspired by recent cellular neurobiological evidence linking neocortical pyramidal cells to distinct mental states, this work shows how models (e.g., Transformers) can emulate high-level perceptual processing and awake thought (imagination) states to pre-select relevant information before applying attention. Triadic neuronal-level modulation loops among questions ($Q$), clues (keys, $K$), and hypotheses (values, $V$) enable diverse, deep, parallel reasoning chains at the representation level and allow a rapid shift from initial biases to refined understanding. This leads to orders-of-magnitude faster learning with significantly reduced computational demand (e.g., fewer heads, layers, and tokens), at an approximate cost of $\mathcal{O}(N)$, where $N$ is the number of input tokens. Results span reinforcement learning (e.g., CarRacing in a high-dimensional visual setup), computer vision, and natural language question answering.

Figures

Figures reproduced from arXiv: 2505.06257 by the authors.

Figure 1
Figure 1. A pyramidal two-point neuron in the mammalian neocor [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. An example of “thinking fast and slow” as discussed in [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Co4 architecture: N denotes the number of input tokens, and each token has an embedding dimension of E. Q1, Q2,...,QL represent the latent query tokens input to the associated Q-TPNs. K1, K2,...,KN represent the Key tokens input to the associated K-TPNs. V1, V2,...,VN represent the Value tokens input to the associated V-TPNs. This configuration forms part of the “seeing” state (i.e., sensory processing). In the “see… view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Vector field visualization and contour plots for deeper [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 6
Figure 6. Figure 6: Training Results: In Cart-Pole, PyBullet Ant, and CarRacing (with high dimensional visual environment [96 x 96 x 4]) tasks, [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: CIFAR-10 Training Results: In all tested configurations, [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Synthetic dataset similar to Facebook’s bAbI dataset [62, [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Synthetic Facebook bAbI Training Results: In all tested configurations, [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 58 canonical work pages

  1. [60]

    The sensory neuron as a trans- former: Permutation-invariant neural networks for re- inforcement learning,

    Y . Tang and D. Ha, “The sensory neuron as a trans- former: Permutation-invariant neural networks for re- inforcement learning,” Advances in Neural Informa- tion Processing Systems , vol. 34, pp. 22 574–22 587, 2021

  2. [2]

    Perceiver: General perception with iterative attention,

    A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisser- man, and J. Carreira, “Perceiver: General perception with iterative attention,” in International conference on machine learning. PMLR, 2021, pp. 4651–4664

  3. [1]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural infor- mation processing systems, vol. 30, 2017

  4. [3]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in neural in- formation processing systems , vol. 35, pp. 23 716– 23 736, 2022

  5. [4]

    Cellular psychology: relating cognition to context-sensitive pyramidal cells,

    W. A. Phillips, T. Bachmann, W. Spratling, L. Muckli, L. Petro, and T. Zolnik, “Cellular psychology: relating cognition to context-sensitive pyramidal cells,”Trends in Cognitive Sciences, 2024

  6. [5]

    A new cellular mechanism for coupling inputs arriving at dif- ferent cortical layers,

    M. E. Larkum, J. J. Zhu, and B. Sakmann, “A new cellular mechanism for coupling inputs arriving at dif- ferent cortical layers,” Nature, vol. 398, no. 6725, pp. 338–341, 1999

  7. [6]

    W. A. Phillips, The Cooperative Neuron: Cellular Foundations of Mental Life. Oxford University Press, 2023

  8. [7]

    A cellular mechanism for cortical asso- ciations: an organizing principle for the cerebral cor- tex,

    M. Larkum, “A cellular mechanism for cortical asso- ciations: an organizing principle for the cerebral cor- tex,” Trends in neurosciences, vol. 36, no. 3, pp. 141– 151, 2013

Show all 66 references
  1. [8]

    Are dendrites conceptually useful?

    M. E. Larkum, “Are dendrites conceptually useful?” Neuroscience, vol. 489, pp. 4–14, 2022

  2. [9]

    Neocortical layer 1: An elegant solution to 11 top-down and bottom-up integration,

    B. Schuman, S. Dellal, A. Pr ¨onneke, R. Machold, and B. Rudy, “Neocortical layer 1: An elegant solution to 11 top-down and bottom-up integration,” Annual Review of Neuroscience, vol. 44, no. 1, pp. 221–252, 2021, pMID: 33730511

  3. [10]

    Illuminating dendritic function with computational models,

    P. Poirazi and A. Papoutsi, “Illuminating dendritic function with computational models,” Nature Reviews Neuroscience, vol. 21, pp. 1–19, 05 2020

  4. [11]

    A perspective on cortical layering and layer-spanning neuronal elements,

    M. E. Larkum, L. S. Petro, R. N. Sachdev, and L. Muckli, “A perspective on cortical layering and layer-spanning neuronal elements,” Frontiers in neu- roanatomy, vol. 12, p. 56, 2018

  5. [12]

    The dynamics of functional brain networks: integrated network states during cognitive task performance,

    J. M. Shine, P. G. Bissett, P. T. Bell, O. Koyejo, J. H. Balsters, K. J. Gorgolewski, C. A. Moodie, and R. A. Poldrack, “The dynamics of functional brain networks: integrated network states during cognitive task performance,” Neuron, vol. 92, no. 2, pp. 544– 554, 2016

  6. [13]

    Human cognition involves the dy- namic integration of neural activity and neuromodula- tory systems,

    J. M. Shine, M. Breakspear, P. T. Bell, K. A. Eh- goetz Martens, R. Shine, O. Koyejo, O. Sporns, and R. A. Poldrack, “Human cognition involves the dy- namic integration of neural activity and neuromodula- tory systems,” Nature neuroscience, vol. 22, no. 2, pp. 289–296, 2019

  7. [14]

    Neuromodulatory influences on integra- tion and segregation in the brain,

    J. M. Shine, “Neuromodulatory influences on integra- tion and segregation in the brain,” Trends in cognitive sciences, vol. 23, no. 7, pp. 572–583, 2019

  8. [15]

    Computational models link cellular mechanisms of neuromodulation to large- scale neural dynamics,

    J. M. Shine, E. J. M ¨uller, B. Munn, J. Cabral, R. J. Moran, and M. Breakspear, “Computational models link cellular mechanisms of neuromodulation to large- scale neural dynamics,” Nature neuroscience, vol. 24, no. 6, pp. 765–776, 2021

  9. [16]

    Gaba b receptor-mediated regulation of dendro-somatic synergy in layer 5 pyramidal neu- rons,

    J. M. Schulz, J. W. Kay, J. Bischofberger, and M. E. Larkum, “Gaba b receptor-mediated regulation of dendro-somatic synergy in layer 5 pyramidal neu- rons,” Frontiers in cellular neuroscience , vol. 15, p. 718413, 2021

  10. [17]

    Cellular mech- anisms of conscious processing,

    J. Aru, M. Suzuki, and M. Larkum, “Cellular mech- anisms of conscious processing,” Trends in Cognitive Sciences, vol. 25, 10 2021

  11. [18]

    An integrative, multiscale view on neural theories of consciousness,

    J. F. Storm, P. C. Klink, J. Aru, W. Senn, R. Goebel, A. Pigorini, P. Avanzini, W. Vanduffel, P. R. Roelf- sema, M. Massimini et al., “An integrative, multiscale view on neural theories of consciousness,” Neuron, vol. 112, no. 10, pp. 1531–1552, 2024

  12. [19]

    Apical amplification—a cellular mechanism of conscious perception?

    T. Marvan, M. Pol ´ak, T. Bachmann, and W. A. Phillips, “Apical amplification—a cellular mechanism of conscious perception?” Neuroscience of conscious- ness, vol. 2021, no. 2, p. niab036, 2021

  13. [20]

    Dendritic integration dysfunction in neurodevelopmental disorders,

    A. D. Nelson and K. J. Bender, “Dendritic integration dysfunction in neurodevelopmental disorders,”Devel- opmental Neuroscience, vol. 43, no. 3-4, pp. 201–221, 2021

  14. [21]

    Dysfunctions of cellular context- sensitivity in neurodevelopmental learning disabil- ities,

    A. Granato, W. A. Phillips, J. M. Schulz, M. Suzuki, and M. E. Larkum, “Dysfunctions of cellular context- sensitivity in neurodevelopmental learning disabil- ities,” Neuroscience & Biobehavioral Reviews , p. 105688, 2024

  15. [22]

    Apical drive—a cellular mechanism of dreaming?

    J. Aru, F. Siclari, W. A. Phillips, and J. F. Storm, “Apical drive—a cellular mechanism of dreaming?” Neuroscience & Biobehavioral Reviews, vol. 119, pp. 440–455, 2020

  16. [23]

    Dendritic inte- gration theory: a thalamo-cortical theory of state and content of consciousness,

    T. Bachmann, M. Suzuki, and J. Aru, “Dendritic inte- gration theory: a thalamo-cortical theory of state and content of consciousness,” Philosophy and the Mind Sciences, vol. 1, no. II, 2020

  17. [24]

    The free-energy principle: a unified brain theory?

    K. Friston, “The free-energy principle: a unified brain theory?” Nature reviews neuroscience, vol. 11, no. 2, pp. 127–138, 2010

  18. [25]

    A review of predictive coding al- gorithms,

    M. W. Spratling, “A review of predictive coding al- gorithms,” Brain and cognition, vol. 112, pp. 92–97, 2017

  19. [26]

    Canonical microcir- cuits for predictive coding,

    A. M. Bastos, W. M. Usrey, R. A. Adams, G. R. Man- gun, P. Fries, and K. J. Friston, “Canonical microcir- cuits for predictive coding,”Neuron, vol. 76, no. 4, pp. 695–711, 2012

  20. [27]

    A theory of cortical responses,

    K. Friston, “A theory of cortical responses,” Philo- sophical transactions of the Royal Society B: Biologi- cal sciences, vol. 360, no. 1456, pp. 815–836, 2005

  21. [28]

    Deep temporal models and active inference,

    K. J. Friston, R. Rosch, T. Parr, C. Price, and H. Bow- man, “Deep temporal models and active inference,” Neuroscience & Biobehavioral Reviews , vol. 90, pp. 486–501, 2018

  22. [29]

    The graphical brain: Belief propagation and active inference,

    K. J. Friston, T. Parr, and B. de Vries, “The graphical brain: Belief propagation and active inference,” Net- work neuroscience, vol. 1, no. 4, pp. 381–414, 2017

  23. [30]

    Cellular mechanisms of cooperative context-sensitive predictive inference,

    T. Marvan and W. A. Phillips, “Cellular mechanisms of cooperative context-sensitive predictive inference,” Current Research in Neurobiology, p. 100129, 2024

  24. [31]

    Burst-dependent synaptic plasticity can co- ordinate learning in hierarchical circuits,

    A. Payeur, J. Guerguiev, F. Zenke, B. A. Richards, and R. Naud, “Burst-dependent synaptic plasticity can co- ordinate learning in hierarchical circuits,”Nature neu- roscience, vol. 24, no. 7, pp. 1010–1019, 2021

  25. [32]

    Single-phase deep learning in cortico-cortical networks,

    e. a. Greedy, Will, “Single-phase deep learning in cortico-cortical networks,” Advances in Neural Infor- mation Processing Systems, 2022. 12

  26. [33]

    Towards deep learning with segregated dendrites,

    J. Guerguiev, T. Lillicrap, and B. Richards, “Towards deep learning with segregated dendrites,”eLife, vol. 6, p. e22901, 12 2017

  27. [34]

    Dendritic cortical microcircuits approx- imate the backpropagation algorithm,

    J. Sacramento, R. Ponte Costa, Y . Bengio, and W. Senn, “Dendritic cortical microcircuits approx- imate the backpropagation algorithm,” Advances in neural information processing systems, vol. 31, 2018

  28. [35]

    Lo- cal plasticity rules can learn deep representations us- ing self-supervised contrastive predictions,

    B. Illing, J. Ventura, G. Bellec, and W. Gerstner, “Lo- cal plasticity rules can learn deep representations us- ing self-supervised contrastive predictions,”Advances in Neural Information Processing Systems, 2022

  29. [36]

    Continual learn- ing through synaptic intelligence,

    F. Zenke, B. Poole, and S. Ganguli, “Continual learn- ing through synaptic intelligence,” in International conference on machine learning. PMLR, 2017, pp. 3987–3995

  30. [37]

    Overcoming catastrophic forgetting in neural networks,

    J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ra- malho, A. Grabska-Barwinska et al. , “Overcoming catastrophic forgetting in neural networks,” Proceed- ings of the national academy of sciences , vol. 114, no. 13, pp. 35...

  31. [38]

    Linking memories across time via neuronal and dendritic over- laps in model neurons with active dendrites,

    G. Kastellakis, A. J. Silva, and P. Poirazi, “Linking memories across time via neuronal and dendritic over- laps in model neurons with active dendrites,” Cell re- ports, vol. 17, no. 6, pp. 1491–1504, 2016

  32. [39]

    Modeling somatic and den- dritic spike mediated plasticity at the single neuron and network level,

    J. Bono and C. Clopath, “Modeling somatic and den- dritic spike mediated plasticity at the single neuron and network level,” Nature communications, vol. 8, no. 1, p. 706, 2017

  33. [40]

    Emergence of sta- ble synaptic clusters on dendrites through synaptic rewiring,

    T. Limbacher and R. Legenstein, “Emergence of sta- ble synaptic clusters on dendrites through synaptic rewiring,” Frontiers in computational neuroscience , vol. 14, p. 57, 2020

  34. [41]

    Contextual modulation in mammalian neocortex is asymmetric,

    J. W. Kay and W. A. Phillips, “Contextual modulation in mammalian neocortex is asymmetric,” Symmetry, vol. 12, no. 5, p. 815, 2020

  35. [42]

    A com- parison of partial information decompositions using data from real and simulated layer 5b pyramidal cells,

    J. W. Kay, J. M. Schulz, and W. A. Phillips, “A com- parison of partial information decompositions using data from real and simulated layer 5b pyramidal cells,” Entropy, vol. 24, no. 8, p. 1021, 2022

  36. [43]

    Two-compartment neuronal spiking model express- ing brain-state specific apical-amplification,-isolation and-drive regimes,

    E. Pastorelli, A. Yegenoglu, N. Kolodziej, W. Wybo, F. Simula, S. Diaz, J. F. Storm, and P. S. Paolucci, “Two-compartment neuronal spiking model express- ing brain-state specific apical-amplification,-isolation and-drive regimes,”arXiv preprint arXiv:2311.06074, 2023

  37. [44]

    Transfer functions for burst firing probability in a model neo- cortical pyramidal cell,

    B. P. Graham, J. W. Kay, and W. A. Phillips, “Transfer functions for burst firing probability in a model neo- cortical pyramidal cell,” bioRxiv, pp. 2024–01, 2024

  38. [45]

    Unlocking the potential of two-point cells for energy-efficient and resilient train- ing of deep nets,

    A. Adeel, A. Adetomi, K. Ahmed, A. Hussain, T. Ar- slan, and W. A. Phillips, “Unlocking the potential of two-point cells for energy-efficient and resilient train- ing of deep nets,” IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 7, no. 3, pp. 818–828, 2023

  39. [46]

    The cortical microcircuitry of predictions and context: a multi-scale perspective,

    L. e. a. Muckli*, “The cortical microcircuitry of predictions and context: a multi-scale perspective,” Sep. 2023. [Online]. Available: https://doi.org/10. 5281/zenodo.8380094

  40. [47]

    An overlooked role of context-sensitive dendrites,

    M. Raza and A. Adeel, “An overlooked role of context-sensitive dendrites,” arXiv preprint arXiv:2408.11019, 2024

  41. [48]

    Memories off the top of your head,

    J. Shin, G. Doron, and M. Larkum, “Memories off the top of your head,” Science, vol. 374, pp. 538–539, 10 2021

  42. [49]

    Thinking, fast and slow,

    D. Kahneman, “Thinking, fast and slow,” Farrar, Straus and Giroux, 2011

  43. [50]

    Pattern recognition and modern computers,

    O. G. Selfridge, “Pattern recognition and modern computers,” in Proceedings of the March 1-3, 1955, western joint computer conference, 1955, pp. 91–93

  44. [51]

    D. E. Rumelhart, J. L. McClelland, P. R. Group et al., Parallel distributed processing, volume 1: Explo- rations in the microstructure of cognition: Founda- tions. The MIT press, 1986

  45. [52]

    Dou- ble bookkeeping and schizophrenia spectrum: di- vided unified phenomenal consciousness,

    J. Parnas, A. Urfer-Parnas, and H. Stephensen, “Dou- ble bookkeeping and schizophrenia spectrum: di- vided unified phenomenal consciousness,” European Archives of Psychiatry and Clinical Neuroscience, vol. 271, pp. 1513–1523, 2021

  46. [53]

    The phenomenologi- cal nature of schizophrenia and disorder of selfhood,

    J. Parnas and K. E. Sandsten, “The phenomenologi- cal nature of schizophrenia and disorder of selfhood,” Schizophrenia Research, vol. 270, pp. 197–201, 2024

  47. [54]

    First steps to- ward a psychopathology of

    W. Blankenburg and A. L. Mishara, “First steps to- ward a psychopathology of” common sense”,”Philos- ophy, psychiatry, & psychology, vol. 8, no. 4, pp. 303– 315, 2001

  48. [55]

    Local choices: Rationality and the contex- tuality of decision-making,

    I. Vlaev, “Local choices: Rationality and the contex- tuality of decision-making,” Brain sciences , vol. 8, no. 1, p. 8, 2018

  49. [56]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhou et al., “Chain-of-thought prompting elicits reasoning in large language models,” 13 Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022

  50. [57]

    Contextually guided unsupervised learning using local multivariate binary processors,

    J. Kay, D. Floreano, and W. A. Phillips, “Contextually guided unsupervised learning using local multivariate binary processors,” Neural Networks, vol. 11, no. 1, pp. 117–140, 1998

  51. [58]

    Context- sensitive processing in a model neocortical pyramidal cell with two sites of input integration,

    B. P. Graham, J. W. Kay, and W. A. Phillips, “Context- sensitive processing in a model neocortical pyramidal cell with two sites of input integration,” Neural Com- putation, vol. 37, no. 4, pp. 588–634, 2025

  52. [59]

    Deepseek llm: Scaling open-source language models with longter- mism,

    X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu et al., “Deepseek llm: Scaling open-source language models with longter- mism,” arXiv preprint arXiv:2401.02954, 2024

  53. [61]

    Escaping the big data paradigm with compact transformers,

    A. Hassani, S. Walton, N. Shah, A. Abuduweili, J. Li, and H. Shi, “Escaping the big data paradigm with compact transformers,” arXiv preprint arXiv:2104.05704, 2021

  54. [62]

    To- wards ai-complete question answering: A set of pre- requisite toy tasks,

    J. Weston, A. Bordes, S. Chopra, A. M. Rush, B. Van Merri¨enboer, A. Joulin, and T. Mikolov, “To- wards ai-complete question answering: A set of pre- requisite toy tasks,” arXiv preprint arXiv:1502.05698, 2015

  55. [63]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, “Long short-term memory,”Neural computation, vol. 9, no. 8, pp. 1735– 1780, 1997

  56. [64]

    Memory net- works,

    J. Weston, S. Chopra, and A. Bordes, “Memory net- works,” arXiv preprint arXiv:1410.3916, 2014

  57. [65]

    Bregman, Humankind: A hopeful history

    R. Bregman, Humankind: A hopeful history . Bloomsbury Publishing, 2020

  58. [66]

    Cooperation is all you need,

    A. Adeel, J. Muzaffar, K. Ahmed, and M. Raza, “Cooperation is all you need,” arXiv preprint arXiv:2305.10449, 2023. 14

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.