REVIEW 4 major objections 6 minor 37 references
An explicit low-dimensional bottleneck is necessary for out-of-distribution generalisation, and the causal emergence of that representation tracks both artificial and biological learning.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 11:47 UTC pith:SMEDPXTC
load-bearing objection Clean ablation shows a bottleneck is needed for rotational/OOD generalisation; the Ψ–generalisation link is suggestive but rests on a fixed PCA + Gaussian MI choice that the paper never stress-tests. the 4 major comments →
Emergent Generalization by Representation Learning in Artificial Neural Networks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
An explicit information bottleneck that forces a recurrent network to form a low-dimensional latent representation is necessary for rotational and out-of-distribution generalization on a next-step prediction task. The causal-emergence measure Ψ of that representation follows a non-monotonic trajectory (drop, minimum, rise) that coincides with the memorization-to-generalization transition, scales with task complexity, and its magnitude reliably predicts generalization performance; analogous dynamics appear in mouse CA1 and mPFC during spatial learning.
What carries the argument
The Ψ measure of causal emergence: the mutual information of the three-dimensional PCA projection of the bottleneck (or linear encoder) with its own future, minus the sum of the marginal mutual informations of the individual reservoir (or recorded) units with their futures. It quantifies how much more predictive the coarse-grained latent is than the sum of its parts, and thereby tracks the quality of the learned representation for generalization.
Load-bearing premise
That the three-dimensional PCA projection of the bottleneck (or of the linear encoder fitted to neural rates) is the right macro-variable for measuring causal emergence, and that Gaussian mutual-information estimates capture the relevant predictive structure.
What would settle it
Train the same reservoir architecture without any low-dimensional bottleneck (or with a deliberately non-emergent high-dimensional readout) and show that it still achieves comparable zero-shot generalization to held-out dynamical systems and random rotations; or show that Ψ computed on an alternative macro-variable fails to predict generalization while the original Ψ still does.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that an explicit low-dimensional bottleneck is necessary for rotational and out-of-distribution generalisation in a reservoir-computing next-step prediction task on high-dimensional projections of low-dimensional chaotic attractors. It further claims that the learned latent representation becomes causally emergent (Ψ of Rosas et al.), that Ψ follows a non-monotonic trajectory across the memorisation-to-generalisation (grokking) transition even while MSE falls monotonically, that post-training Ψ (especially its macro-predictability component) predicts generalisation loss after mixed-effects correction for attractor identity, and that analogous non-monotonic Ψ dynamics appear in mouse CA1 and mPFC during W-maze alternation learning. The architecture is a fixed reservoir plus a trainable linear bottleneck trained by backprop on next-step MSE, with ridge readout; task difficulty is controlled by autocorrelation-time subsampling (N_tau).
Significance. If the results hold under more robust measurement of emergence, the work would give a concrete functional argument for neural manifolds: compact, distributed, causally emergent representations are not merely descriptive but enable OOD generalisation. The synthetic design cleanly ablates the bottleneck, normalises task complexity across attractors, and shows zero-shot transfer to held-out dynamical systems. Linking grokking to a rise in Ψ, and reporting a parallel non-monotonic trajectory in CA1/mPFC that precedes behavioural decoder improvement, is a useful bridge between information-bottleneck theory, causal emergence, and systems neuroscience. The training objective does not optimise Ψ, so the reported Ψ–generalisation correlation is not forced by construction—an important strength.
major comments (4)
- Methods, Causal Emergence and Eq. (5)–(6): The central claim that “the magnitude of emergent structure reliably predicts generalisation performance” rests on Ψ computed from a fixed three-dimensional PCA projection of the bottleneck (or of the linear encoder on CA1/PFC rates) under a multivariate-Gaussian mutual-information estimator. PCA maximises variance, not predictive power; if the true predictive macro lives on a different linear or nonlinear subspace, both the non-monotonic trajectory and the Ψ–OOD-loss correlation can be artefacts of the projection. The Gaussian estimator captures only second-order linear dependencies, while the reservoir and chaotic attractors generate non-Gaussian, nonlinear structure. The manuscript never shows that the same Ψ–loss relationship survives under an alternative macro definition (e.g., CCA to the true latent, predictive ICA, or a bottleneck-sized V
- Results, “Learned Latent Representations are Causally Emergent” and Fig. 4H: The mixed-effects decomposition of Ψ into macro and micro predictability is the load-bearing statistical link to generalisation. After conditioning on macro, micro becomes positively associated with loss (distributed representations help). This is interesting but fragile: (i) both components are first normalised by the true latent attractor’s predictability, which is unavailable in the biological setting and couples the metric to ground-truth knowledge; (ii) the paper reports coefficients without effect sizes, confidence intervals, or leave-one-attractor-out stability; (iii) the claim that Ψ “reliably predicts” generalisation is stronger than a single mixed-effects fit on a modest number of attractors and N_tau values. A clearer predictive analysis (e.g., out-of-sample R² of L_G from post-training Ψ, or partial
- Results, “Emergent Representations in spatial learning tasks in mice” and Fig. 5: The biological analysis fits a linear bottleneck encoder + ridge decoder to predict next-step position/velocity from binned CA1 or mPFC rates, then computes Ψ on that bottleneck across sessions. This is a useful descriptive parallel, but it does not test the paper’s functional claim (that emergent low-dimensional representations confer a generalisation advantage). There is no OOD or rotational generalisation test in the animals, no ablation of the bottleneck, and no demonstration that sessions with higher Ψ support better transfer. The text itself notes that “further work is needed to establish the link from emergence to generalisation in biological neural networks.” The abstract and Discussion currently overstate the biological result as supporting a “causal role for learned representations in cognition.”
- Results, “Information Bottleneck enables Rotation Invariant OOD Generalisation” and Fig. 2A: The necessity claim for the bottleneck is central and currently rests on a single with/without comparison at fixed N_tau=20. Without a bottleneck the network can still fit training data; the failure mode on rotations and held-out attractors is therefore the key evidence. It would strengthen the paper to show that this failure is not an artefact of capacity or regularisation (e.g., match parameter count, add weight decay or early stopping on the no-bottleneck model, or replace the hard bottleneck with a soft IB penalty). If the no-bottleneck model still fails under those controls, the necessity claim is solid; if not, the result is more about regularisation than about low-dimensional representation per se.
minor comments (6)
- Eq. (1) in the main text writes the micro term as I(Y^i_t ; Y^i_{t+1}), while Eq. (6) in Methods and the Fig. 4 schematic write I(Y^i_t ; V_{t+1}). These are different quantities; the text should use a single consistent definition and state which one is plotted.
- Fig. 3 caption and panel labels: panel order and the reported Pearson r=0.336 / partial r=0.028 should be cross-checked against the figure; the ratio histogram (panel D) would benefit from a vertical line at 1 and a statement of the number of models.
- Methods, Model architecture: reservoir size (500), washout (1000), batch construction, and Optuna hyperparameter ranges should be fully listed (or pointed to a supplement/code) so the experiments are reproducible.
- Introduction and Discussion cite grokking and IB literature appropriately, but the claim that Ψ “precedes” the IB transition (Fig. 4C–D) is only visual; a quantitative lag or epoch-of-minimum comparison would make the ordering claim precise.
- Table 1 (Appendix): Lyapunov exponents and fractal dimensions for Chen and Sprott A are given approximately; state the source or computation method for consistency with Lorenz/Rössler.
- Typos / wording: “Y et” → “Yet” (abstract); “sparesely” → “sparsely”; “preceeds” → “precedes”; “supervinient” → “supervenient” (Methods).
Circularity Check
No significant circularity: bottleneck is trained only on next-step MSE; Ψ (Rosas) and OOD loss are independent post-hoc measures whose correlation is empirical, not definitional.
full rationale
The derivation chain is self-contained. Networks are trained solely by back-propagation of next-step MSE through a linear bottleneck (plus ridge readout); the training objective never includes Ψ, IB, CCA similarity, or OOD loss. After training, a fixed 3-D PCA of the bottleneck is scored with the externally defined Ψ measure of Rosas et al. (2020) and with Gaussian mutual information; generalisation is measured on held-out random rotations and entirely unseen attractors. The observed non-monotonic Ψ trajectory, its scaling with N au, and its correlation with LG (via linear mixed-effects) are therefore empirical associations, not identities forced by construction or by a fitted parameter being re-labelled as a prediction. The biological analysis likewise fits a linear encoder per session only for behavioural next-step prediction, then independently evaluates Ψ on that encoder; the claim that Ψ is non-monotonic while MSE falls is a comparison of two separately computed quantities, not a tautology. Self-citations (Rajpal et al. bioRxiv on prior applications of Ψ) supply background, not a load-bearing uniqueness theorem or ansatz. No equation reduces to its own input, no uniqueness is imported from the authors, and no known empirical pattern is merely renamed. Score 0 is therefore required.
Axiom & Free-Parameter Ledger
free parameters (5)
- N_tau (autocorrelation subsampling horizon)
- bottleneck dimensionality
- reservoir hyper-parameters (leak, spectral radius, sparsity, input scaling, bias)
- beta in information-bottleneck loss
- PCA rank for macro-variable V_t
axioms (4)
- standard math Mutual information between multivariate Gaussians is given by the closed-form determinant formula (Eq. 5).
- domain assumption Ψ = I(V_t; V_{t+1}) − Σ_i I(Y^i_t; Y^i_{t+1}) quantifies causal emergence (Rosas et al. 2020).
- ad hoc to paper The three-dimensional PCA projection of the bottleneck (or of the linear encoder) is a valid supervenient macro-variable for Ψ.
- domain assumption Next-step MSE after ridge regression is a sufficient training objective for learning generalisable latent dynamics.
read the original abstract
Dimensionality reduction has proven powerful for identifying neural manifolds, which are low-dimensional structures underlying high-dimensional neural activity. These low-dimensional representations have improved the interpretability of population-level coding. Yet whether such low-dimensional representations are biologically relevant and confer functional advantages in learning systems, or merely reflect neuron-level activity, remains contested in neuroscience. We show that an explicit information bottleneck forcing a recurrent neural network to learn a low-dimensional representation is necessary for rotational and out-of-distribution generalisation in a time-series prediction task. Using information-theoretic measures of causal emergence, we characterise the dynamics of this representation across the memorisation-to-generalisation transition, finding a non-monotonic trajectory which shows an initial decrease, a minimum, and a subsequent rise to a maximum, even as prediction loss falls monotonically. This trajectory scales with task complexity, and the magnitude of emergent structure reliably predicts generalisation performance. Analysis of CA1 hippocampal activity in mice learning an alternating maze task reveals analogous non-monotonic emergence dynamics that track behavioural performance. Together, these findings indicate that the ability of neural networks to learn compact, distributed and emergent representations confers a functional advantage for generalisation, supporting a causal role for learned representations in cognition.
Figures
Reference graph
Works this paper leans on
-
[1]
Cunningham, J. P. & Yu, B. M. Dimensionality reduction for large-scale neural recordings.Nat. neuroscience17, 1500–1509 (2014). 2.Perich, M. G., Narain, D. & Gallego, J. A. A neural manifold view of the brain.Nat. Neurosci.28, 1582–1597 (2025)
2014
-
[2]
& Laurent, G
Stopfer, M., Jayaraman, V . & Laurent, G. Intensity versus identity coding in an olfactory system.Neuron39, 991–1004 (2003)
2003
-
[3]
& Laurent, G
Mazor, O. & Laurent, G. Transient dynamics versus fixed points in odor representations by locust antennal lobe projection neurons.Neuron48, 661–673 (2005). 5.Golub, M. D.et al.Learning by neural reassociation.Nat. neuroscience21, 607–616 (2018). 6.Churchland, M. M.et al.Neural population dynamics during reaching.Nature487, 51–56 (2012)
2005
-
[4]
S., Rouault, H., Druckmann, S
Kim, S. S., Rouault, H., Druckmann, S. & Jayaraman, V . Ring attractor dynamics in the drosophila central brain.Science 356, 849–853 (2017)
2017
-
[5]
& Fiete, I
Chaudhuri, R., Gerçek, B., Pandey, B., Peyrache, A. & Fiete, I. The intrinsic attractor manifold and population dynamics of a canonical cognitive circuit across waking and sleep.Nat. neuroscience22, 1512–1520 (2019)
2019
-
[6]
S., Hermundstad, A
Kim, S. S., Hermundstad, A. M., Romani, S., Abbott, L. & Jayaraman, V . Generation of stable heading representations in diverse visual scenes.Nature576, 126–131 (2019). 10.Yang, Z., Inagaki, M., Gerfen, C. R., Fontolan, L. & Inagaki, H. K. Integrator dynamics in the cortico-basal ganglia loop for flexible motor timing.Nature649, 1244–1253 (2026)
2019
-
[7]
K., Shi, J., Phensy, A
Cho, K. K., Shi, J., Phensy, A. J., Turner, M. L. & Sohal, V . S. Long-range inhibition synchronizes and updates prefrontal task activity.Nature617, 548–554 (2023)
2023
-
[8]
A., Perich, M
Gallego, J. A., Perich, M. G., Chowdhury, R. H., Solla, S. A. & Miller, L. E. Long-term stability of cortical population dynamics underlying consistent behavior.Nat. neuroscience23, 260–270 (2020)
2020
-
[9]
& Proekt, A
Brennan, C. & Proekt, A. A quantitative model of conserved macroscopic dynamics predicts future motor commands. Elife8, e46814 (2019). 14.Proix, T., Perich, M. & Milekovic, T. Misinterpreting the horseshoe effect in neuroscience.bioRxiv(2022)
2019
-
[10]
Elsayed, G. F. & Cunningham, J. P. Structure in neural population recordings: an expected byproduct of simpler phenomena?Nat. neuroscience20, 1310–1318 (2017)
2017
-
[11]
D., Caballero, J
Humphries, M. D., Caballero, J. A., Evans, M., Maggi, S. & Singh, A. Spectral estimation for detecting low-dimensional structure in networks using arbitrary null models.Plos one16, e0254057 (2021)
2021
-
[12]
& Chaudhuri, R
De, A. & Chaudhuri, R. Common population codes produce extremely nonlinear neural manifolds.Proc. Natl. Acad. Sci. 120, e2305853120 (2023). 12/14
2023
-
[13]
Bertram, J., Dyballa, L., Keller, T. A., Kinger, S. & Zucker, S. W. Decoding alignment without encoding alignment: A critique of similarity analysis in neuroscience.arXiv preprint arXiv:2605.05907(2026). 19.Tishby, N., Pereira, F. C. & Bialek, W. The information bottleneck method.arXiv preprint physics/0004057(2000)
Pith/arXiv arXiv 2026
-
[14]
Shwartz-Ziv, R. & Tishby, N. Opening the black box of deep neural networks via information.arXiv preprint arXiv:1703.00810(2017)
Pith/arXiv arXiv 2017
-
[15]
& Soatto, S
Achille, A. & Soatto, S. Emergence of invariance and disentanglement in deep representations.J. Mach. Learn. Res.19, 1–34 (2018)
2018
-
[16]
Alemi, A. A., Fischer, I., Dillon, J. V . & Murphy, K. Deep variational information bottleneck.arXiv preprint arXiv:1612.00410(2016)
Pith/arXiv arXiv 2016
-
[17]
Arjovsky, M., Bottou, L., Gulrajani, I. & Lopez-Paz, D. Invariant risk minimization.arXiv preprint arXiv:1907.02893 (2019)
Pith/arXiv arXiv 1907
-
[18]
Intell.(2025)
Li, J.et al.Contrastive learning via variational information bottleneck.IEEE Transactions on Pattern Analysis Mach. Intell.(2025)
2025
-
[19]
M.et al.On the information bottleneck theory of deep learning.J
Saxe, A. M.et al.On the information bottleneck theory of deep learning.J. Stat. Mech. Theory Exp.2019, 124020 (2019)
2019
-
[20]
Hafez-Kolahi, H., Kasaei, S. & Soleymani-Baghshah, M. Do compressed representations generalize better?arXiv preprint arXiv:1909.09706(2019). 27.Goldfeld, Z.et al.Estimating information flow in deep neural networks.arXiv preprint arXiv:1810.05728(2018)
Pith/arXiv arXiv 1909
-
[21]
Humayun, A. I., Balestriero, R. & Baraniuk, R. Deep networks always grok and here is why.arXiv preprint arXiv:2402.15555(2024)
Pith/arXiv arXiv 2024
-
[22]
Power, A., Burda, Y ., Edwards, H., Babuschkin, I. & Misra, V . Grokking: Generalization beyond overfitting on small algorithmic datasets.arXiv preprint arXiv:2201.02177(2022)
Pith/arXiv arXiv 2022
-
[23]
Neural Inf
Liu, Z.et al.Towards understanding grokking: An effective theory of representation learning.Adv. Neural Inf. Process. Syst.35, 34651–34663 (2022)
2022
-
[24]
InEuropean Conference on Computer Vision, 20–37 (Springer, 2022)
Cui, Q.et al.Discriminability-transferability trade-off: An information-theoretic perspective. InEuropean Conference on Computer Vision, 20–37 (Springer, 2022)
2022
-
[25]
& Levine, S
Finn, C., Abbeel, P. & Levine, S. Model-agnostic meta-learning for fast adaptation of deep networks. InInternational conference on machine learning, 1126–1135 (PMLR, 2017)
2017
-
[26]
M., Ullman, T
Lake, B. M., Ullman, T. D., Tenenbaum, J. B. & Gershman, S. J. Building machines that learn and think like people.Behav. brain sciences40, e253 (2017)
2017
-
[27]
P., Albantakis, L
Hoel, E. P., Albantakis, L. & Tononi, G. Quantifying causal emergence shows that macro can beat micro.Proc. Natl. Acad. Sci.110, 19790–19795 (2013)
2013
-
[28]
Flack, J. C. Coarse-graining as a downward causation mechanism.Philos. transactions. Ser. A, Math. physical, engineering sciences375, 20160338 (2017)
2017
-
[29]
E.et al.Reconciling emergences: An information-theoretic approach to identify causal emergence in multivariate data.PLoS computational biology16, e1008289 (2020)
Rosas, F. E.et al.Reconciling emergences: An information-theoretic approach to identify causal emergence in multivariate data.PLoS computational biology16, e1008289 (2020)
2020
-
[30]
& Hakim, V
Brunel, N. & Hakim, V . Sparsely synchronized neuronal oscillations.Chaos: An Interdiscip. J. Nonlinear Sci.18(2008)
2008
-
[31]
A., Sas, M
Rajpal, H., Mediano, P. A., Sas, M. I., Jensen, H. J. & Rosas, F. E. Quantifying the emergence of population-level activity in neuronal systems.bioRxiv2026–02 (2026)
2026
-
[32]
Clauw, K., Stramaglia, S. & Marinazzo, D. Information-theoretic progress measures reveal grokking is an emergent phase transition.arXiv preprint arXiv:2408.08944(2024)
Pith/arXiv arXiv 2024
-
[33]
Tang, W., Shin, J. D. & Jadhav, S. P. Geometric transformation of cognitive maps for generalization across hippocampal- prefrontal circuits.Cell reports42(2023)
2023
-
[34]
J., Slatton, W
Wakhloo, A. J., Slatton, W. & Chung, S. Neural population geometry and optimal coding of tasks with shared latent structure.Nat. Neurosci.1–11 (2026)
2026
-
[35]
M., Luppi, A
Tolle, H. M., Luppi, A. I., Seth, A. K. & Mediano, P. A. Evolving reservoir computers reveal bidirectional coupling between predictive power and emergent dynamics.Patterns7(2026)
2026
-
[36]
Fountas, Z., Oomerjee, A., Bou-Ammar, H., Wang, J. & Burgess, N. Why the brain consolidates: Predictive forgetting for optimal generalisation.arXiv preprint arXiv:2603.04688(2026). 13/14
arXiv 2026
-
[37]
& Abbott, L
Chung, S. & Abbott, L. F. Neural population geometry: An approach for understanding biological and artificial neural networks.Curr. opinion neurobiology70, 137–144 (2021). Acknowledgements H.R. was funded by the Eric and Wendy Schmidt AI in Science Postdoctoral Fellowship, supported by Schmidt Sciences, LLC. A Training and Test Attractors In this study we...
2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.