Pith. sign in

REVIEW 4 major objections 6 minor 37 references

An explicit low-dimensional bottleneck is necessary for out-of-distribution generalisation, and the causal emergence of that representation tracks both artificial and biological learning.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 11:47 UTC pith:SMEDPXTC

load-bearing objection Clean ablation shows a bottleneck is needed for rotational/OOD generalisation; the Ψ–generalisation link is suggestive but rests on a fixed PCA + Gaussian MI choice that the paper never stress-tests. the 4 major comments →

arxiv 2607.10430 v1 pith:SMEDPXTC submitted 2026-07-11 q-bio.NC cs.ITcs.LGcs.NEmath.ITnlin.CD

Emergent Generalization by Representation Learning in Artificial Neural Networks

classification q-bio.NC cs.ITcs.LGcs.NEmath.ITnlin.CD
keywords neural manifoldsinformation bottleneckcausal emergenceout-of-distribution generalisationreservoir computinggrokkinghippocampal CA1representation learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether low-dimensional neural manifolds are more than a convenient description of high-dimensional activity. It answers by training a reservoir network with a forced bottleneck on next-step prediction of high-dimensional time series generated from several low-dimensional dynamical systems. Without the bottleneck the network memorizes the training projections but fails on random rotations and on held-out attractors; with the bottleneck it recovers a rotation-invariant latent basis and generalizes. Across training, a measure of causal emergence of the latent representation first falls then rises, even while prediction loss falls steadily; the final amount of emergence scales with task difficulty and predicts generalization. Parallel analysis of mouse CA1 and medial prefrontal activity during W-maze learning shows the same non-monotonic emergence trajectory, which precedes the behavioral performance minimum. The authors conclude that compact, distributed, emergent representations confer a functional advantage for generalization and therefore support a causal role for learned manifolds in cognition.

Core claim

An explicit information bottleneck that forces a recurrent network to form a low-dimensional latent representation is necessary for rotational and out-of-distribution generalization on a next-step prediction task. The causal-emergence measure Ψ of that representation follows a non-monotonic trajectory (drop, minimum, rise) that coincides with the memorization-to-generalization transition, scales with task complexity, and its magnitude reliably predicts generalization performance; analogous dynamics appear in mouse CA1 and mPFC during spatial learning.

What carries the argument

The Ψ measure of causal emergence: the mutual information of the three-dimensional PCA projection of the bottleneck (or linear encoder) with its own future, minus the sum of the marginal mutual informations of the individual reservoir (or recorded) units with their futures. It quantifies how much more predictive the coarse-grained latent is than the sum of its parts, and thereby tracks the quality of the learned representation for generalization.

Load-bearing premise

That the three-dimensional PCA projection of the bottleneck (or of the linear encoder fitted to neural rates) is the right macro-variable for measuring causal emergence, and that Gaussian mutual-information estimates capture the relevant predictive structure.

What would settle it

Train the same reservoir architecture without any low-dimensional bottleneck (or with a deliberately non-emergent high-dimensional readout) and show that it still achieves comparable zero-shot generalization to held-out dynamical systems and random rotations; or show that Ψ computed on an alternative macro-variable fails to predict generalization while the original Ψ still does.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper argues that an explicit low-dimensional bottleneck is necessary for rotational and out-of-distribution generalisation in a reservoir-computing next-step prediction task on high-dimensional projections of low-dimensional chaotic attractors. It further claims that the learned latent representation becomes causally emergent (Ψ of Rosas et al.), that Ψ follows a non-monotonic trajectory across the memorisation-to-generalisation (grokking) transition even while MSE falls monotonically, that post-training Ψ (especially its macro-predictability component) predicts generalisation loss after mixed-effects correction for attractor identity, and that analogous non-monotonic Ψ dynamics appear in mouse CA1 and mPFC during W-maze alternation learning. The architecture is a fixed reservoir plus a trainable linear bottleneck trained by backprop on next-step MSE, with ridge readout; task difficulty is controlled by autocorrelation-time subsampling (N_tau).

Significance. If the results hold under more robust measurement of emergence, the work would give a concrete functional argument for neural manifolds: compact, distributed, causally emergent representations are not merely descriptive but enable OOD generalisation. The synthetic design cleanly ablates the bottleneck, normalises task complexity across attractors, and shows zero-shot transfer to held-out dynamical systems. Linking grokking to a rise in Ψ, and reporting a parallel non-monotonic trajectory in CA1/mPFC that precedes behavioural decoder improvement, is a useful bridge between information-bottleneck theory, causal emergence, and systems neuroscience. The training objective does not optimise Ψ, so the reported Ψ–generalisation correlation is not forced by construction—an important strength.

major comments (4)
  1. Methods, Causal Emergence and Eq. (5)–(6): The central claim that “the magnitude of emergent structure reliably predicts generalisation performance” rests on Ψ computed from a fixed three-dimensional PCA projection of the bottleneck (or of the linear encoder on CA1/PFC rates) under a multivariate-Gaussian mutual-information estimator. PCA maximises variance, not predictive power; if the true predictive macro lives on a different linear or nonlinear subspace, both the non-monotonic trajectory and the Ψ–OOD-loss correlation can be artefacts of the projection. The Gaussian estimator captures only second-order linear dependencies, while the reservoir and chaotic attractors generate non-Gaussian, nonlinear structure. The manuscript never shows that the same Ψ–loss relationship survives under an alternative macro definition (e.g., CCA to the true latent, predictive ICA, or a bottleneck-sized V
  2. Results, “Learned Latent Representations are Causally Emergent” and Fig. 4H: The mixed-effects decomposition of Ψ into macro and micro predictability is the load-bearing statistical link to generalisation. After conditioning on macro, micro becomes positively associated with loss (distributed representations help). This is interesting but fragile: (i) both components are first normalised by the true latent attractor’s predictability, which is unavailable in the biological setting and couples the metric to ground-truth knowledge; (ii) the paper reports coefficients without effect sizes, confidence intervals, or leave-one-attractor-out stability; (iii) the claim that Ψ “reliably predicts” generalisation is stronger than a single mixed-effects fit on a modest number of attractors and N_tau values. A clearer predictive analysis (e.g., out-of-sample R² of L_G from post-training Ψ, or partial
  3. Results, “Emergent Representations in spatial learning tasks in mice” and Fig. 5: The biological analysis fits a linear bottleneck encoder + ridge decoder to predict next-step position/velocity from binned CA1 or mPFC rates, then computes Ψ on that bottleneck across sessions. This is a useful descriptive parallel, but it does not test the paper’s functional claim (that emergent low-dimensional representations confer a generalisation advantage). There is no OOD or rotational generalisation test in the animals, no ablation of the bottleneck, and no demonstration that sessions with higher Ψ support better transfer. The text itself notes that “further work is needed to establish the link from emergence to generalisation in biological neural networks.” The abstract and Discussion currently overstate the biological result as supporting a “causal role for learned representations in cognition.”
  4. Results, “Information Bottleneck enables Rotation Invariant OOD Generalisation” and Fig. 2A: The necessity claim for the bottleneck is central and currently rests on a single with/without comparison at fixed N_tau=20. Without a bottleneck the network can still fit training data; the failure mode on rotations and held-out attractors is therefore the key evidence. It would strengthen the paper to show that this failure is not an artefact of capacity or regularisation (e.g., match parameter count, add weight decay or early stopping on the no-bottleneck model, or replace the hard bottleneck with a soft IB penalty). If the no-bottleneck model still fails under those controls, the necessity claim is solid; if not, the result is more about regularisation than about low-dimensional representation per se.
minor comments (6)
  1. Eq. (1) in the main text writes the micro term as I(Y^i_t ; Y^i_{t+1}), while Eq. (6) in Methods and the Fig. 4 schematic write I(Y^i_t ; V_{t+1}). These are different quantities; the text should use a single consistent definition and state which one is plotted.
  2. Fig. 3 caption and panel labels: panel order and the reported Pearson r=0.336 / partial r=0.028 should be cross-checked against the figure; the ratio histogram (panel D) would benefit from a vertical line at 1 and a statement of the number of models.
  3. Methods, Model architecture: reservoir size (500), washout (1000), batch construction, and Optuna hyperparameter ranges should be fully listed (or pointed to a supplement/code) so the experiments are reproducible.
  4. Introduction and Discussion cite grokking and IB literature appropriately, but the claim that Ψ “precedes” the IB transition (Fig. 4C–D) is only visual; a quantitative lag or epoch-of-minimum comparison would make the ordering claim precise.
  5. Table 1 (Appendix): Lyapunov exponents and fractal dimensions for Chen and Sprott A are given approximately; state the source or computation method for consistency with Lorenz/Rössler.
  6. Typos / wording: “Y et” → “Yet” (abstract); “sparesely” → “sparsely”; “preceeds” → “precedes”; “supervinient” → “supervenient” (Methods).

Circularity Check

0 steps flagged

No significant circularity: bottleneck is trained only on next-step MSE; Ψ (Rosas) and OOD loss are independent post-hoc measures whose correlation is empirical, not definitional.

full rationale

The derivation chain is self-contained. Networks are trained solely by back-propagation of next-step MSE through a linear bottleneck (plus ridge readout); the training objective never includes Ψ, IB, CCA similarity, or OOD loss. After training, a fixed 3-D PCA of the bottleneck is scored with the externally defined Ψ measure of Rosas et al. (2020) and with Gaussian mutual information; generalisation is measured on held-out random rotations and entirely unseen attractors. The observed non-monotonic Ψ trajectory, its scaling with N au, and its correlation with LG (via linear mixed-effects) are therefore empirical associations, not identities forced by construction or by a fitted parameter being re-labelled as a prediction. The biological analysis likewise fits a linear encoder per session only for behavioural next-step prediction, then independently evaluates Ψ on that encoder; the claim that Ψ is non-monotonic while MSE falls is a comparison of two separately computed quantities, not a tautology. Self-citations (Rajpal et al. bioRxiv on prior applications of Ψ) supply background, not a load-bearing uniqueness theorem or ansatz. No equation reduces to its own input, no uniqueness is imported from the authors, and no known empirical pattern is merely renamed. Score 0 is therefore required.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

The central claim rests on the Rosas Ψ definition, the Gaussian MI estimator, the choice of 3-D PCA as macro-variable, the particular set of six attractors, the Nτ complexity normalisation, and the post-hoc linear encoder applied to biological rates. No new physical entities are postulated; free parameters are architectural and optimisation choices rather than fitted constants that define the result.

free parameters (5)
  • N_tau (autocorrelation subsampling horizon)
    Controls task complexity; varied systematically but chosen by hand as the number of samples before ACF < 1/e.
  • bottleneck dimensionality
    Swept; performance peaks near input dimension (10); treated as a free architectural hyper-parameter.
  • reservoir hyper-parameters (leak, spectral radius, sparsity, input scaling, bias)
    Optimised via TPE/Optuna on validation loss; not derived from first principles.
  • beta in information-bottleneck loss
    Fixed to 1 (equal trade-off); not optimised or derived.
  • PCA rank for macro-variable V_t
    Fixed to 3 to match latent attractor dimension; not learned.
axioms (4)
  • standard math Mutual information between multivariate Gaussians is given by the closed-form determinant formula (Eq. 5).
    Standard result used throughout for Ψ and IB; assumes linear Gaussian statistics.
  • domain assumption Ψ = I(V_t; V_{t+1}) − Σ_i I(Y^i_t; Y^i_{t+1}) quantifies causal emergence (Rosas et al. 2020).
    Adopted without re-derivation; positive Ψ is treated as sufficient for emergence.
  • ad hoc to paper The three-dimensional PCA projection of the bottleneck (or of the linear encoder) is a valid supervenient macro-variable for Ψ.
    Chosen for convenience and to match attractor dimension; not shown to be optimal or unique.
  • domain assumption Next-step MSE after ridge regression is a sufficient training objective for learning generalisable latent dynamics.
    Standard in reservoir computing; the paper shows it is insufficient without the bottleneck.

pith-pipeline@v1.1.0-grok45 · 18777 in / 3086 out tokens · 42511 ms · 2026-07-14T11:47:12.038260+00:00 · methodology

0 comments
read the original abstract

Dimensionality reduction has proven powerful for identifying neural manifolds, which are low-dimensional structures underlying high-dimensional neural activity. These low-dimensional representations have improved the interpretability of population-level coding. Yet whether such low-dimensional representations are biologically relevant and confer functional advantages in learning systems, or merely reflect neuron-level activity, remains contested in neuroscience. We show that an explicit information bottleneck forcing a recurrent neural network to learn a low-dimensional representation is necessary for rotational and out-of-distribution generalisation in a time-series prediction task. Using information-theoretic measures of causal emergence, we characterise the dynamics of this representation across the memorisation-to-generalisation transition, finding a non-monotonic trajectory which shows an initial decrease, a minimum, and a subsequent rise to a maximum, even as prediction loss falls monotonically. This trajectory scales with task complexity, and the magnitude of emergent structure reliably predicts generalisation performance. Analysis of CA1 hippocampal activity in mice learning an alternating maze task reveals analogous non-monotonic emergence dynamics that track behavioural performance. Together, these findings indicate that the ability of neural networks to learn compact, distributed and emergent representations confers a functional advantage for generalisation, supporting a causal role for learned representations in cognition.

Figures

Figures reproduced from arXiv: 2607.10430 by Dan Goodman, Hardik Rajpal.

Figure 1
Figure 1. Figure 1: D). For instance, the Lorenz attractor shows a faster decay of autocorrelation as compared to the Rössler attractor, which makes it more difficult to predict the next time step of the input data. We normalise the complexity of the task across the different dynamical systems by subsampling and keeping exactly Ntau timesteps of the generated data before the autocorrelation of the data drops below 1/e (see Fi… view at source ↗
Figure 2
Figure 2. Figure 2: A Comparing the next step prediction loss of the network with and without the bottleneck layer, for a fixed task complexity (Ntau = 20). The network without the bottleneck layer minimises training error but is not able to generalise to random rotations of the training attractors or to the held-out dynamical systems. B Task complexity modulates the training, validation and generalisation loss. Prediction er… view at source ↗
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: A Schematic representation of the micro states Yt of the reservoir network and the macro states Vt , which are the three-dimensional PCA projection of the bottleneck states. The Ψ measure quantifies the degree to which the macro states are more predictive of their own future state than the sum of the marginal predictability of the micro states. B The MSE decreases monotonically during training across the e… view at source ↗
Figure 5
Figure 5. Figure 5: A Schematic representation of the W maze spatial learning task. The mice are trained to learn a continuous alternation task in a W maze. B The sagittal view of the mouse brain showing the location of the CA1 and medial PFC regions where the neural activity is recorded during the task. C The setup for representation learning in the neural data. The preprocessed firing rates are passed through a simple linea… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

37 extracted references · 8 linked inside Pith

  1. [1]

    Cunningham, J. P. & Yu, B. M. Dimensionality reduction for large-scale neural recordings.Nat. neuroscience17, 1500–1509 (2014). 2.Perich, M. G., Narain, D. & Gallego, J. A. A neural manifold view of the brain.Nat. Neurosci.28, 1582–1597 (2025)

  2. [2]

    & Laurent, G

    Stopfer, M., Jayaraman, V . & Laurent, G. Intensity versus identity coding in an olfactory system.Neuron39, 991–1004 (2003)

  3. [3]

    & Laurent, G

    Mazor, O. & Laurent, G. Transient dynamics versus fixed points in odor representations by locust antennal lobe projection neurons.Neuron48, 661–673 (2005). 5.Golub, M. D.et al.Learning by neural reassociation.Nat. neuroscience21, 607–616 (2018). 6.Churchland, M. M.et al.Neural population dynamics during reaching.Nature487, 51–56 (2012)

  4. [4]

    S., Rouault, H., Druckmann, S

    Kim, S. S., Rouault, H., Druckmann, S. & Jayaraman, V . Ring attractor dynamics in the drosophila central brain.Science 356, 849–853 (2017)

  5. [5]

    & Fiete, I

    Chaudhuri, R., Gerçek, B., Pandey, B., Peyrache, A. & Fiete, I. The intrinsic attractor manifold and population dynamics of a canonical cognitive circuit across waking and sleep.Nat. neuroscience22, 1512–1520 (2019)

  6. [6]

    S., Hermundstad, A

    Kim, S. S., Hermundstad, A. M., Romani, S., Abbott, L. & Jayaraman, V . Generation of stable heading representations in diverse visual scenes.Nature576, 126–131 (2019). 10.Yang, Z., Inagaki, M., Gerfen, C. R., Fontolan, L. & Inagaki, H. K. Integrator dynamics in the cortico-basal ganglia loop for flexible motor timing.Nature649, 1244–1253 (2026)

  7. [7]

    K., Shi, J., Phensy, A

    Cho, K. K., Shi, J., Phensy, A. J., Turner, M. L. & Sohal, V . S. Long-range inhibition synchronizes and updates prefrontal task activity.Nature617, 548–554 (2023)

  8. [8]

    A., Perich, M

    Gallego, J. A., Perich, M. G., Chowdhury, R. H., Solla, S. A. & Miller, L. E. Long-term stability of cortical population dynamics underlying consistent behavior.Nat. neuroscience23, 260–270 (2020)

  9. [9]

    & Proekt, A

    Brennan, C. & Proekt, A. A quantitative model of conserved macroscopic dynamics predicts future motor commands. Elife8, e46814 (2019). 14.Proix, T., Perich, M. & Milekovic, T. Misinterpreting the horseshoe effect in neuroscience.bioRxiv(2022)

  10. [10]

    Elsayed, G. F. & Cunningham, J. P. Structure in neural population recordings: an expected byproduct of simpler phenomena?Nat. neuroscience20, 1310–1318 (2017)

  11. [11]

    D., Caballero, J

    Humphries, M. D., Caballero, J. A., Evans, M., Maggi, S. & Singh, A. Spectral estimation for detecting low-dimensional structure in networks using arbitrary null models.Plos one16, e0254057 (2021)

  12. [12]

    & Chaudhuri, R

    De, A. & Chaudhuri, R. Common population codes produce extremely nonlinear neural manifolds.Proc. Natl. Acad. Sci. 120, e2305853120 (2023). 12/14

  13. [13]

    A., Kinger, S

    Bertram, J., Dyballa, L., Keller, T. A., Kinger, S. & Zucker, S. W. Decoding alignment without encoding alignment: A critique of similarity analysis in neuroscience.arXiv preprint arXiv:2605.05907(2026). 19.Tishby, N., Pereira, F. C. & Bialek, W. The information bottleneck method.arXiv preprint physics/0004057(2000)

  14. [14]

    & Tishby, N

    Shwartz-Ziv, R. & Tishby, N. Opening the black box of deep neural networks via information.arXiv preprint arXiv:1703.00810(2017)

  15. [15]

    & Soatto, S

    Achille, A. & Soatto, S. Emergence of invariance and disentanglement in deep representations.J. Mach. Learn. Res.19, 1–34 (2018)

  16. [16]

    A., Fischer, I., Dillon, J

    Alemi, A. A., Fischer, I., Dillon, J. V . & Murphy, K. Deep variational information bottleneck.arXiv preprint arXiv:1612.00410(2016)

  17. [17]

    & Lopez-Paz, D

    Arjovsky, M., Bottou, L., Gulrajani, I. & Lopez-Paz, D. Invariant risk minimization.arXiv preprint arXiv:1907.02893 (2019)

  18. [18]

    Intell.(2025)

    Li, J.et al.Contrastive learning via variational information bottleneck.IEEE Transactions on Pattern Analysis Mach. Intell.(2025)

  19. [19]

    M.et al.On the information bottleneck theory of deep learning.J

    Saxe, A. M.et al.On the information bottleneck theory of deep learning.J. Stat. Mech. Theory Exp.2019, 124020 (2019)

  20. [20]

    & Soleymani-Baghshah, M

    Hafez-Kolahi, H., Kasaei, S. & Soleymani-Baghshah, M. Do compressed representations generalize better?arXiv preprint arXiv:1909.09706(2019). 27.Goldfeld, Z.et al.Estimating information flow in deep neural networks.arXiv preprint arXiv:1810.05728(2018)

  21. [21]

    I., Balestriero, R

    Humayun, A. I., Balestriero, R. & Baraniuk, R. Deep networks always grok and here is why.arXiv preprint arXiv:2402.15555(2024)

  22. [22]

    & Misra, V

    Power, A., Burda, Y ., Edwards, H., Babuschkin, I. & Misra, V . Grokking: Generalization beyond overfitting on small algorithmic datasets.arXiv preprint arXiv:2201.02177(2022)

  23. [23]

    Neural Inf

    Liu, Z.et al.Towards understanding grokking: An effective theory of representation learning.Adv. Neural Inf. Process. Syst.35, 34651–34663 (2022)

  24. [24]

    InEuropean Conference on Computer Vision, 20–37 (Springer, 2022)

    Cui, Q.et al.Discriminability-transferability trade-off: An information-theoretic perspective. InEuropean Conference on Computer Vision, 20–37 (Springer, 2022)

  25. [25]

    & Levine, S

    Finn, C., Abbeel, P. & Levine, S. Model-agnostic meta-learning for fast adaptation of deep networks. InInternational conference on machine learning, 1126–1135 (PMLR, 2017)

  26. [26]

    M., Ullman, T

    Lake, B. M., Ullman, T. D., Tenenbaum, J. B. & Gershman, S. J. Building machines that learn and think like people.Behav. brain sciences40, e253 (2017)

  27. [27]

    P., Albantakis, L

    Hoel, E. P., Albantakis, L. & Tononi, G. Quantifying causal emergence shows that macro can beat micro.Proc. Natl. Acad. Sci.110, 19790–19795 (2013)

  28. [28]

    Flack, J. C. Coarse-graining as a downward causation mechanism.Philos. transactions. Ser. A, Math. physical, engineering sciences375, 20160338 (2017)

  29. [29]

    E.et al.Reconciling emergences: An information-theoretic approach to identify causal emergence in multivariate data.PLoS computational biology16, e1008289 (2020)

    Rosas, F. E.et al.Reconciling emergences: An information-theoretic approach to identify causal emergence in multivariate data.PLoS computational biology16, e1008289 (2020)

  30. [30]

    & Hakim, V

    Brunel, N. & Hakim, V . Sparsely synchronized neuronal oscillations.Chaos: An Interdiscip. J. Nonlinear Sci.18(2008)

  31. [31]

    A., Sas, M

    Rajpal, H., Mediano, P. A., Sas, M. I., Jensen, H. J. & Rosas, F. E. Quantifying the emergence of population-level activity in neuronal systems.bioRxiv2026–02 (2026)

  32. [32]

    & Marinazzo, D

    Clauw, K., Stramaglia, S. & Marinazzo, D. Information-theoretic progress measures reveal grokking is an emergent phase transition.arXiv preprint arXiv:2408.08944(2024)

  33. [33]

    Tang, W., Shin, J. D. & Jadhav, S. P. Geometric transformation of cognitive maps for generalization across hippocampal- prefrontal circuits.Cell reports42(2023)

  34. [34]

    J., Slatton, W

    Wakhloo, A. J., Slatton, W. & Chung, S. Neural population geometry and optimal coding of tasks with shared latent structure.Nat. Neurosci.1–11 (2026)

  35. [35]

    M., Luppi, A

    Tolle, H. M., Luppi, A. I., Seth, A. K. & Mediano, P. A. Evolving reservoir computers reveal bidirectional coupling between predictive power and emergent dynamics.Patterns7(2026)

  36. [36]

    & Burgess, N

    Fountas, Z., Oomerjee, A., Bou-Ammar, H., Wang, J. & Burgess, N. Why the brain consolidates: Predictive forgetting for optimal generalisation.arXiv preprint arXiv:2603.04688(2026). 13/14

  37. [37]

    & Abbott, L

    Chung, S. & Abbott, L. F. Neural population geometry: An approach for understanding biological and artificial neural networks.Curr. opinion neurobiology70, 137–144 (2021). Acknowledgements H.R. was funded by the Eric and Wendy Schmidt AI in Science Postdoctoral Fellowship, supported by Schmidt Sciences, LLC. A Training and Test Attractors In this study we...