Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Multiscale nonlinear integration drives accurate encoding of input information

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that in multiscale stochastic processing networks, nonlinear integration (weighted sum first, then a tanh nonlinearity) yields higher input-output mutual information than nonlinear summation (tanh first, then sum), with…

desk verdict The paper's analytical machinery is solid and the integration-vs-summation comparison is a real contribution, but the 'always' claim in Sec. II C outruns the authors' own numerics and needs rewording plus error bars before the headline result can stand as stated. read the letter →

arxiv 2411.11710 v1 pith:4ZCEINB2 submitted 2024-11-18 cond-mat.stat-mech

classification cond-mat.stat-mech MSC 82C3194A1782C32 PACS 05.40.-a89.70.Cf
keywords mutualinformationnonlinearintegrationsummationtimescaleseparationstochasticmultilayernetworksinput-outputencodingoptimaldimensionalitybistability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper studies a minimal stochastic processing chain—an input unit, an intermediate processing unit, and an output unit, each containing many degrees of freedom on its own timescale—and asks how much information about the input survives to the output. It compares two wiring schemes: nonlinear summation, where each signal is passed through a $\tanh$ nonlinearity and then averaged, and nonlinear integration, where the weighted average is formed first and then passed through $\tanh$. Using a timescale-separated solution for the stationary distribution, the paper shows that integration almost always gives higher input-output mutual information than summation, for both fast and slow processing units, with fast processing making the advantage larger. The reason to care is architectural: if the claim holds, the order of pooling and nonlinear activation is a generic control knob for encoding accuracy, and the best processing-layer size is set by the input dimension rather than being a free parameter.

What carries the argument

The load-bearing object is the timescale-separated factorization of the joint stationary density into Gaussian conditionals: $p^{\mathrm{fp}}_{IOP} = p_I p_{P|I} p^{\mathrm{eff}}_{O|I}$ for fast processing and $p^{\mathrm{sp}}_{IOP} = p_I p_{P|I} p_{O|P}$ for slow processing, with the input as the slowest unit and each conditional covariance independent of the conditioning variable. The nonlinearity is pushed entirely into the conditional means; for the hyperbolic tangent, the required Gaussian average is evaluated exactly through a convergent series expansion, yielding closed-form means in both summation and integration schemes. With these means, fast-processing mutual information reduces to the output entropy minus a constant conditional-entropy term, while slow-processing mutual information is obtained by sampling the conditional distribution; output bistability is then quantified with Sarle's bimodality coefficient.

What would settle it

Run direct Langevin simulations of the same three-unit model—Gaussian input, $\tanh$ couplings, summation versus integration—over the full grid of coupling strengths, layer sizes, and processing-to-output weight variances without imposing the factorized solution, and estimate input-output mutual information from long trajectories at each point; the central claim stands only if $I_{IO}^{\mathrm{int}} > I_{IO}^{\mathrm{ns}}$ wherever the paper's formulas predict it, especially at strong coupling with small $M_P$ and large $\sigma_{OP}^2$, where only a few points are currently checked against simulation. A complementary check is to compute the first finite-timescale-ratio corrections to the factorized densities and see whether the sign of the difference survives.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central result is an ordering: with a slow input and a processing unit that is either faster or slower than the output, the input-output mutual information of nonlinear integration exceeds that of nonlinear summation, $I_{IO}^{\mathrm{int}} > I_{IO}^{\mathrm{ns}}$, across a wide range of coupling strengths, layer sizes, and weight distributions. The documented exception is a strong-coupling regime at small processing dimension and large variance of the processing-to-output weights, where summation can win because the few sampled weights push the activation into saturation. The same calculation uncovers two design principles: for a fixed input dimension $M_I$ there is an optimal processing dimension $M_P^*$, with low-dimensional inputs served best by high-dimensional embeddings and high-dimensional inputs by low-dimensional projections; and integration spontaneously produces a bimodal output distribution, read by the authors as input discrimination whose mode balance can be tuned by adding biases to the activation.

Load-bearing premise

The comparison stands or falls on the timescale-separated approximation imported from the authors' earlier work: at stationarity each unit's response to the unit before it is a Gaussian whose variance does not depend on the signal value, so all nonlinearity enters only through the conditional mean; if strong couplings produce corrections that change those variances, the information ordering between integration and summation could change.

Editorial extensions

If this is right

  • In multiscale processing chains with a slow input, combining signals before the nonlinear activation is a generic way to raise input-output mutual information relative to the activate-then-combine scheme.
  • A processing unit that is faster than the output increases information transfer on its own and amplifies the advantage of integration; slower processing reduces but does not erase it.
  • For a given input dimension there is an optimal processing-layer dimension, and the optimal strategy flips from high-dimensional embedding for small inputs to low-dimensional projection for large inputs.
  • Integration spontaneously yields a bistable output distribution, and introducing biases into the activation shifts the weight of the two modes, giving a tunable input-discrimination mechanism.
  • The same ordering and optimal-dimension behaviour carry over to chains with more than one processing unit, since the derivation is independent of the number of processing layers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the reason integration wins is plausibly that averaging before a saturating nonlinearity keeps fluctuations in the linear regime of $\tanh$; if so, the same ordering should hold for other sigmoidal activations, which the paper does not test.
  • Editorial extension: in practical reservoir or neural architectures, the result suggests placing the readout nonlinearity after the weighted sum of internal activities and choosing the hidden-layer width relative to the input dimension; this is a design heuristic the paper does not train or evaluate.
  • Editorial extension: the tunable output bistability could serve as a primitive classifier, with the active output mode labelling the input; turning the bimodality coefficient into a classification error rate would be a natural next step beyond the paper.
  • Editorial extension: the derivation assumes exact timescale separation, so whether integration remains superior at order-one timescale ratios is an open question that direct simulation at finite ratios could settle.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript studies a three-unit stochastic dynamical system (input, processing layer, output) with timescale separation and nonlinear activation functions. It derives a joint stationary distribution in the limits of fast and slow processing, expressed as products of Gaussian conditional distributions whose means depend nonlinearly on the conditioning variables. Using this factorization, the authors compute the input–output mutual information for two processing schemes—nonlinear summation (process-then-sum) and nonlinear integration (sum-then-process)—and report that integration yields higher mutual information over a wide parameter range, that fast processing enhances information, that an optimal processing-layer size exists, and that integration promotes tunable output bistability.

Significance. The analytical framework is a genuine strength: it avoids fitted parameters, provides an efficient sampling scheme, and is cross-checked against Langevin simulations in Fig. 1. If the ordering result is robust, the work offers a general principle for motif design in biological and artificial networks—integrate before nonlinearity to improve encoding. The practical relevance is high, as the model covers architectures such as random recurrent neural networks and signaling cascades. However, the headline claim is more general than the evidence presented, and the numerical support lacks uncertainty estimates; these issues must be resolved before the paper can serve as a foundation.

major comments (3)
  1. [Sec. II.C; Figs. 2i-j; SM Fig. S3] The assertion in Sec. II.C that, "independently of the internal timescale ordering, nonlinear integration always leads to higher mutual information with respect to summation" is false as written: the authors' own Figs. 2i-j and SM Fig. S3 show a regime with small processing dimension M_P, large interaction variance σ_OP, and strong couplings (g_PI = g_OP around 10) where I^ns_IO > I^int_IO. The main-text caption explicitly acknowledges this ("at large variances ... such advantage is achieved only for large enough M_P"). Please replace "always" with a qualified statement that explicitly carves out this regime, or restrict the claim to match the abstract's "over a wide range of parameters."
  2. [Sec. II.C; Figs. 2a-k; SM Figs. S3-S5] All mutual-information values are reported as averages over 10^3 random-matrix realizations, but no error bars, confidence intervals, or realization-to-realization statistics are provided. Because the central claim is an ordering between I^int_IO and I^ns_IO, the absence of uncertainty quantification makes it impossible to judge whether the reported reversal at small M_P is a statistically robust effect or sampling noise, and likewise whether the advantages in other panels are significant. Please provide bootstrap or standard-error estimates, and show the distribution of per-realization I_IO for at least one point in the reversal regime and one point in the advantage regime.
  3. [Sec. II.B; Eqs. (3)-(4); SM Eqs. (S34), (S42)] The factorization is imported from the authors' prior work and is controlled by timescale ratios rather than coupling strength, so I do not question its leading-order validity. However, the numerical comparisons are made at strong couplings (g = 5-10), while the only direct Langevin verification (Fig. 1) is at g = 5 for a single parameter set. A direct check of the mutual information at a strong-coupling point in the reversal region and at a point in the advantage region would confirm that the ordering is not an artifact of the conditional-Gaussian sampling at these parameters.
minor comments (5)
  1. [Eq. (1) and SM Eq. (S7)] There is a sign inconsistency in the definition of h_{O|I}: the main text defines h_{O|I} = ⟨p_{O|I} log2 p_{O|I}⟩_O without a minus sign, whereas SM Eq. (S7) correctly defines it as the conditional entropy with a minus sign. As written, Eq. (1) would give IIO = HO + H_{O|I}. Please fix the main-text definition so that h_{O|I} is the conditional entropy and Eq. (1) reads IIO = HO - ⟨h_{O|I}⟩_I.
  2. [Eq. (5)] In the conditional Gaussian p_{μ|ν} = N(m_{μ|ν}(x_ν), Σ_ν), the covariance should be that of the conditioned variable μ (e.g., Σ_P for p_{P|I} and Σ_O for p_{O|P}), not Σ_ν. This notation is inconsistent with Eq. (7) and with SM Eqs. (S20) and (S36).
  3. [Abstract and Sec. II.C] The abstract's "systematically enhance" and Sec. II.C's "always" should be harmonized; the latter is too strong given the exception stated later in the same section.
  4. [Fig. 2 caption] The caption already contains a clear caveat about the large-σ_OP regime; the main text should refer to this caveat in the same paragraph as the "always" sentence, rather than only in the later discussion.
  5. [SM Fig. S3 caption] The caption uses "Mp" with a lowercase p in "if Mp is small"; this should be M_P for consistency with the rest of the paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the mutual-information comparison is computed from derived distributions and verified against Langevin simulation, with no fitted parameter or definitional shortcut.

full rationale

The central claim is a comparison of input-output mutual information between nonlinear summation and integration. The mutual informations are evaluated from the factorized stationary measures (fast processing via Eq. (3) and Eq. (9); slow processing via Eq. (4) and sampling of Eq. (10)), and the integrals/entropies are computed from the model definitions in Eq. (2). No parameter is fitted to the target ordering, and no quantity is defined in terms of the mutual information it is used to predict. The timescale-separation factorization cited from the authors' prior work [36,55] is re-derived in the Supplemental Material order-by-order, yielding Eqs. (S34) and (S42), so the self-citations are not load-bearing in the present derivation. The slow-input condition is imported from [36] as a parameter-free prior theorem, not as an assumption containing the target result. The paper's own Fig. 2(i-j) and SM Fig. S3 document a strong-coupling, small-M_P regime where nonlinear summation outperforms integration, which contradicts the 'always' wording in Sec. II C; however, this is an internal consistency or scope issue, not a circular reduction. Thus, no circular step can be exhibited with a specific equation-to-equation reduction.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper contributes a new comparison inside an existing framework. No parameters are fitted to data; the model's coupling constants and noise variances are chosen inputs that are swept. The quantitative conclusions depend on the structural assumptions listed above and on finite-sample entropy estimation.

free parameters (2)
  • Coupling strengths g_PI, g_OP = 5 or 10 in main figures; scanned in Fig. 2
    Chosen by hand to define the numerical regime; the integration advantage reverses in a small-M_P, large-sigma_OP, strong-coupling corner, so these values are not neutral.
  • Random matrix variances sigma_PI, sigma_OP, sigma_II, sigma_PP = 1, 1, 0.9, 0.9 in default runs
    Chosen for stability and simplicity; the comparison is sensitive to sigma_OP (for example, Fig. S3), so this is a non-innocuous modeling choice.
assumptions (4)
  • domain assumption The timescale-separated stationary joint density factorizes as p_I p_{P|I} p^{eff}_{O|I} (fast processing) or p_I p_{P|I} p_{O|P} (slow processing), with conditional Gaussian factors, a result imported from refs. [36,55].
    Basis of all mutual information computations; appears in main-text Eqs. (3)-(4) and SM Eqs. (S34), (S42). Only representative cases are validated against Langevin simulation in Fig. 1.
  • domain assumption The input evolves independently, is the slowest unit, and no mutual information is generated when the input is not the slowest (ref. [36]).
    Defines the hierarchical, feed-forward architecture and the parameter regime studied; the condition is taken from refs. [36,55] and is not re-derived here.
  • domain assumption Intra-unit dynamics are linear, with the only nonlinearities being inter-unit tanh activation functions.
    Makes the conditional distributions Gaussian and is the reason the exact joint pdf is tractable; restricts the model's generality.
  • standard math The Gaussian average of tanh can be computed term-by-term from the z>0 geometric-series expansion, and the resulting infinite series converges to controlled accuracy.
    Used in SM Eqs. (S24), (S27), (S32) to obtain effective output means; convergence is demonstrated numerically in SM Fig. S2 but not proven as a uniform series.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multiscale nonlinear integration drives accurate encoding of input information." pith.science (2026). https://pith.science/paper/4ZCEINB2

@misc{pith2026241111710,
  author       = {Pith},
  title        = {Pith review of: Multiscale nonlinear integration drives accurate encoding of input information},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4ZCEINB2}},
  note         = {Machine review of arXiv:2411.11710}
}
read the original abstract

Biological and artificial systems encode information through several complex nonlinear operations, making their exact study a formidable challenge. These internal mechanisms often take place across multiple timescales and process external signals to enable functional output responses. In this work, we focus on two widely implemented paradigms: nonlinear summation, where signals are first processed independently and then combined; and nonlinear integration, where they are combined first and then processed. We study a general model where the input signal is propagated to an output unit through a processing layer via nonlinear activation functions. Further, we distinguish between the two cases of fast and slow processing timescales. We demonstrate that integration and fast-processing capabilities systematically enhance input-output mutual information over a wide range of parameters and system sizes, while simultaneously enabling tunable input discrimination. Moreover, we reveal that high-dimensional embeddings and low-dimensional projections emerge naturally as optimal competing strategies. Our results uncover the foundational features of nonlinear information processing with profound implications for both biological and artificial systems.

Figures

Figures reproduced from arXiv: 2411.11710 by the authors.

Figure 1
Figure 1. FIG. 1. (a-d) Output distribution and stochastic trajecto [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. (a-c) Mutual information between input and output [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. (a) Bimodality coefficient of output pdf, [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stochastic processes with multiple temporal scales: timescale separation and information

    cond-mat.stat-mech 2025-09 conditional novelty 5.0 of 10

    In the strong-timescale-separation limit, fast-to-slow couplings do not create statistical dependence between layers, while slow-to-fast feedback couplings do, and can propagate along minimal network paths.

Reference graph

Works this paper leans on

72 extracted references · 60 canonical work pages · cited by 1 Pith paper

  1. [1]

    Tkaˇ cik and W

    G. Tkaˇ cik and W. Bialek, Information processing in living systems, Annual Review of Condensed Matter Physics 7, 89 (2016)

  2. [2]

    Barkai and S

    N. Barkai and S. Leibler, Robustness in simple biochem- ical networks, Nature 387, 913 (1997)

  3. [3]

    S. K. Aoki, G. Lillacci, A. Gupta, A. Baumschlager, D. Schweingruber, and M. Khammash, A universal biomolecular integral feedback controller for robust per- fect adaptation, Nature 570, 533 (2019)

  4. [4]

    Flatt, D

    S. Flatt, D. M. Busiello, S. Zamuner, and P. De Los Rios, Abc transporters are billion-year-old maxwell demons, Communications Physics 6, 205 (2023)

  5. [5]

    Mochizuki, An analytical study of the number of steady states in gene regulatory networks, Journal of Theoretical Biology 236, 291 (2005)

    A. Mochizuki, An analytical study of the number of steady states in gene regulatory networks, Journal of Theoretical Biology 236, 291 (2005)

  6. [6]

    Trapnell, D

    C. Trapnell, D. Cacchiarelli, J. Grimsby, P. Pokharel, S. Li, M. Morse, N. J. Lennon, K. J. Livak, T. S. Mikkelsen, and J. L. Rinn, The dynamics and regula- tors of cell fate decisions are revealed by pseudotempo- ral ordering of single cells, Nature biotechnology 32, 381 (2014)

  7. [7]

    D. M. Mitrea and R. W. Kriwacki, Phase separation in biology; functional organization of a higher order, Cell Communication and Signaling 14, 1 (2016)

  8. [8]

    Klosin, F

    A. Klosin, F. Oltsch, T. Harmon, A. Honigmann, F. J¨ ulicher, A. A. Hyman, and C. Zechner, Phase sep- aration provides a mechanism to reduce noise in cells, Science 367, 464 (2020)

Show all 72 references
  1. [9]

    S. Vyas, M. D. Golub, D. Sussillo, and K. V. Shenoy, Computation through neural population dynamics, An- nual review of neuroscience 43, 249 (2020)

  2. [10]

    Barzon, G

    G. Barzon, G. Nicoletti, B. Mariani, M. Formentin, and S. Suweis, Criticality and network structure drive emer- gent oscillations in a stochastic whole-brain model, Jour- nal of Physics: Complexity 3, 025010 (2022)

  3. [11]

    Dubreuil, A

    A. Dubreuil, A. Valente, M. Beiran, F. Mastrogiuseppe, and S. Ostojic, The role of population structure in com- putations through neural dynamics, Nature neuroscience 25, 783 (2022)

  4. [12]

    J. D. Jordan, E. M. Landau, and R. Iyengar, Signaling networks: the origins of cellular multitasking, Cell 103, 193 (2000)

  5. [13]

    Cheong, A

    R. Cheong, A. Rhee, C. J. Wang, I. Nemenman, and A. Levchenko, Information transduction capacity of noisy biochemical signaling networks, science 334, 354 (2011)

  6. [14]

    M. T. Hagan, H. B. Demuth, and M. Beale, Neural net- work design (PWS Publishing Co., 1997)

  7. [15]

    Szanda la, Review and comparison of commonly used activation functions for deep neural networks, Bio- inspired neurocomputing , 203 (2021)

    T. Szanda la, Review and comparison of commonly used activation functions for deep neural networks, Bio- inspired neurocomputing , 203 (2021)

  8. [16]

    Karlik and A

    B. Karlik and A. V. Olgac, Performance analysis of vari- ous activation functions in generalized mlp architectures of neural networks, International Journal of Artificial In- telligence and Expert Systems 1, 111 (2011)

  9. [17]

    Ramachandran, B

    P. Ramachandran, B. Zoph, and Q. V. Le, Searching for activation functions, arXiv preprint arXiv:1710.05941 (2017)

  10. [18]

    Apicella, F

    A. Apicella, F. Donnarumma, F. Isgr` o, and R. Prevete, A survey on modern trainable activation functions, Neural Networks 138, 14 (2021)

  11. [19]

    Nwankpa, W

    C. Nwankpa, W. Ijomah, A. Gachagan, and S. Mar- shall, Activation functions: Comparison of trends in practice and research for deep learning, arXiv preprint arXiv:1811.03378 (2018). 7

  12. [20]

    A. Dack, B. Qureshi, T. E. Ouldridge, and T. Plesa, Recurrent neural chemical reaction networks that approximate arbitrary dynamics, arXiv preprint arXiv:2406.03456 (2024)

  13. [21]

    Hayou, A

    S. Hayou, A. Doucet, and J. Rousseau, On the impact of the activation function on deep neural networks training, in Proceedings of the 36th International Conference on Machine Learning, Proceedings of Machine Learning Re- search, Vol. 97, edited by K. Chaudhuri and R. Salakhut...

  14. [22]

    Floyd, A

    C. Floyd, A. R. Dinner, A. Murugan, and S. Vaikun- tanathan, Limits on the computational expressivity of non-equilibrium biophysical processes, arXiv preprint arXiv:2409.05827 (2024)

  15. [23]

    Barzon, D

    G. Barzon, D. M. Busiello, and G. Nicoletti, Excitation- inhibition balance controls information encoding in neu- ral populations, arXiv preprint arXiv:2406.03380 (2024)

  16. [24]

    S. E. Cavanagh, L. T. Hunt, and S. W. Kennerley, A diversity of intrinsic timescales underlie neural computa- tions, Frontiers in Neural Circuits 14, 615626 (2020)

  17. [25]

    Golesorkhi, J

    M. Golesorkhi, J. Gomez-Pilar, F. Zilio, N. Berberian, A. Wolff, M. C. Yagoub, and G. Northoff, The brain and its time: intrinsic neural timescales are key for input pro- cessing, Communications biology 4, 970 (2021)

  18. [26]

    Mariani, G

    B. Mariani, G. Nicoletti, M. Bisio, M. Maschietto, S. Vas- sanelli, and S. Suweis, Disentangling the critical sig- natures of neural activity, Scientific reports 12, 10770 (2022)

  19. [27]

    Zeraati, A

    R. Zeraati, A. Levina, J. H. Macke, and R. Gao, Neu- ral timescales from a computational perspective, arXiv preprint arXiv:2409.02684 (2024)

  20. [28]

    J. M. Parrondo, J. M. Horowitz, and T. Sagawa, Thermo- dynamics of information, Nature physics 11, 131 (2015)

  21. [29]

    Ito and T

    S. Ito and T. Sagawa, Information thermodynamics on causal networks, Physical review letters 111, 180603 (2013)

  22. [30]

    Nicoletti and D

    G. Nicoletti and D. M. Busiello, Mutual information disentangles interactions from changing environments, Physical Review Letters 127, 228301 (2021)

  23. [31]

    Nicoletti and D

    G. Nicoletti and D. M. Busiello, Mutual information in changing environments: non-linear interactions, out-of- equilibrium systems, and continuously-varying diffusivi- ties, Physical Review E 106, 014153 (2022)

  24. [32]

    Nicoletti, M

    G. Nicoletti, M. Bruzzone, S. Suweis, M. Dal Maschio, and D. M. Busiello, Information gain at the onset of ha- bituation to repeated stimuli, eLife 13 (2024)

  25. [33]

    I. R. Graf and B. B. Machta, A bifurcation integrates information from many noisy ion channels and allows for millikelvin thermal sensitivity in the snake pit organ, Proceedings of the National Academy of Sciences 121, e2308215121 (2024)

  26. [34]

    H. H. Mattingly, K. Kamino, B. B. Machta, and T. Emonet, Escherichia coli chemotaxis is information limited, Nature physics 17, 1426 (2021)

  27. [35]

    Bauer and W

    M. Bauer and W. Bialek, Information bottleneck in molecular sensing, PRX Life 1, 023005 (2023)

  28. [36]

    Nicoletti and D

    G. Nicoletti and D. M. Busiello, Information propaga- tion in multilayer systems with higher-order interactions across timescales, Physical Review X 14, 021007 (2024)

  29. [37]

    Tostevin and P

    F. Tostevin and P. R. Ten Wolde, Mutual information be- tween input and output trajectories of biochemical net- works, Physical review letters 102, 218101 (2009)

  30. [38]

    Moor and C

    A.-L. Moor and C. Zechner, Dynamic information trans- fer in stochastic biochemical networks, Physical Review Research 5, 013032 (2023)

  31. [39]

    Nicoletti and D

    G. Nicoletti and D. M. Busiello, Tuning transduction from hidden observables to optimize information harvest- ing, Physical Review Letters 133, 158401 (2024)

  32. [40]

    De Domenico, A

    M. De Domenico, A. Sol´ e-Ribalta, E. Cozzo, M. Kivel¨ a, Y. Moreno, M. A. Porter, S. G´ omez, and A. Arenas, Mathematical formulation of multilayer networks, Phys- ical Review X 3, 041022 (2013)

  33. [41]

    Ghavasieh, C

    A. Ghavasieh, C. Nicolini, and M. De Domenico, Statis- tical physics of complex information dynamics, Physical Review E 102, 052304 (2020)

  34. [42]

    W. Ma, A. Trusina, H. El-Samad, W. A. Lim, and C. Tang, Defining network topologies that can achieve biochemical adaptation, Cell 138, 760 (2009)

  35. [43]

    S. J. Rahi, J. Larsch, K. Pecani, A. Y. Katsov, N. Man- souri, K. Tsaneva-Atanasova, E. D. Sontag, and F. R. Cross, Oscillatory stimuli differentiate adapting circuit topologies, Nature methods 14, 1010 (2017)

  36. [44]

    T.-M. Yi, Y. Huang, M. I. Simon, and J. Doyle, Ro- bust perfect adaptation in bacterial chemotaxis through integral feedback control, Proceedings of the National Academy of Sciences 97, 4649 (2000)

  37. [45]

    T. M. Cover, Elements of information theory (John Wi- ley & Sons, 1999)

  38. [46]

    Sompolinsky, A

    H. Sompolinsky, A. Crisanti, and H.-J. Sommers, Chaos in random neural networks, Physical review letters 61, 259 (1988)

  39. [47]

    Kadmon and H

    J. Kadmon and H. Sompolinsky, Transition to chaos in random neuronal networks, Physical Review X 5, 041030 (2015)

  40. [48]

    Engelken, F

    R. Engelken, F. Wolf, and L. F. Abbott, Lyapunov spec- tra of chaotic recurrent neural networks, Physical Review Research 5, 043044 (2023)

  41. [49]

    Lukoˇ seviˇ cius and H

    M. Lukoˇ seviˇ cius and H. Jaeger, Reservoir computing ap- proaches to recurrent neural network training, Computer Science Review 3, 127 (2009)

  42. [50]

    Maheswaranathan, A

    N. Maheswaranathan, A. Williams, M. Golub, S. Gan- guli, and D. Sussillo, Universality and individuality in neural dynamics across large populations of recurrent networks, Advances in neural information processing sys- tems 32 (2019)

  43. [51]

    L. N. Driscoll, K. Shenoy, and D. Sussillo, Flexible mul- titask computation in recurrent networks utilizes shared dynamical motifs, Nature Neuroscience 27, 1349 (2024)

  44. [52]

    Tanaka, R

    G. Tanaka, R. Nakane, T. Yamane, D. Nakano, S. Takeda, S. Nakagawa, and A. Hirose, Exploiting het- erogeneous units for reservoir computing with simple ar- chitecture, in Neural Information Processing: 23rd Inter- national Conference, ICONIP 2016, Kyoto, Japan, Octo- ber 16–21...

  45. [53]

    Z. K. Malik, A. Hussain, and Q. J. Wu, Multilayered echo state machine: A novel architecture and algorithm, IEEE Transactions on cybernetics 47, 946 (2016)

  46. [54]

    We also study the simplest case of an input-output system without a processing unit

    See supplemental material for analytical derivations, de- tails on sampling methods, and additional numerical re- sults. We also study the simplest case of an input-output system without a processing unit

  47. [55]

    Nicoletti and D

    G. Nicoletti and D. M. Busiello, Information propagation in gaussian processes on multilayer networks, Journal of Physics: Complexity 5, 045004 (2024)

  48. [56]

    Vasicek, A test for normality based on sample en- tropy, Journal of the Royal Statistical Society Series B: Statistical Methodology 38, 54 (1976)

    O. Vasicek, A test for normality based on sample en- tropy, Journal of the Royal Statistical Society Series B: Statistical Methodology 38, 54 (1976). 8

  49. [57]

    L. F. Kozachenko and N. N. Leonenko, Sample estimate of the entropy of a random vector, Problemy Peredachi Informatsii 23, 9 (1987)

  50. [58]

    Lu and J

    C. Lu and J. Peltonen, Enhancing nearest neighbor based entropy estimator for high dimensional distributions via bootstrapping local ellipsoid, in Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34 (2020) pp. 5013–5020

  51. [59]

    Multiscale nonlinear integration drives accurate encoding of input information

    R. Pfister, K. A. Schwarz, M. Janczyk, R. Dale, and J. B. Freeman, Good things peak in pairs: a note on the bimodality coefficient, Frontiers in psychology 4, 700 (2013). 1 Supplemental Material: “Multiscale nonlinear integration drives accurate encoding of input information” ...

  52. [60]

    (S2) and (S3), as the integral over ⃗ xP explicitly depends on the form employed for activation function implementing interactions between the units

    Nonlinear summation We now need to distinguish the two classes of nonlinear couplings in Eqs. (S2) and (S3), as the integral over ⃗ xP explicitly depends on the form employed for activation function implementing interactions between the units. We first consider the case of a n...

  53. [61]

    Contrary to the previous case, we cannot reduce the effective operator to a sum of independent one-dimensional integrals

    Nonlinear integration We now switch to the case of nonlinear integration, Eq (S3). Contrary to the previous case, we cannot reduce the effective operator to a sum of independent one-dimensional integrals. Indeed, we have Lint,eff O (⃗ xI , ⃗ xO) = MOX i=1 ∂ ∂xi O   MOX j=1 A...

  54. [62]

    fast processing

    Computing the mutual information In both scenarios, we have an exact expression for peff,st O|I which is a Gaussian distribution in the output state ⃗ xO with a nonlinear dependence on the input state ⃗ xI . Therefore, we only need to solve the last order of the Fokker-Planck ...

  55. [63]

    sample {⃗ xI }Nsam i=1 from the independent Gaussian distribution of the input

  56. [64]

    (S27) or Eq

    compute the means ⃗ mO|I ({⃗ xI }i) through either Eq. (S27) or Eq. (S32), depending on the nonlinearity, for each sample i

  57. [65]

    slow processing

    for all i, sample ⃗ xO from the multivariate Gaussian with covariance ˆΣO and means ⃗ mO|I ({⃗ xI }i). Then, the entropy HO of the output distribution can be estimated from the samples {⃗ xO}i [56, 57]. Since we focus on one-dimensional outputs, such estimates are especially r...

  58. [66]

    , Nsam,I

    sample a fixed input ⃗ x(i) I ∼ N(0, ˆΣI ) for i = 1, . . . , Nsam,I

  59. [67]

    for each input sample ⃗ x(i) I , compute ⃗ mP |I ⃗ x(i) I , and extract the samples ⃗ x(i,j) P from N ⃗ mP |I ⃗ x(i) I , ˆΣP for j = 1, . . . , Nsam

  60. [68]

    for each processing sample ⃗ x(i,j) P , compute the mean ⃗ mO|P ⃗ x(i,j) P and extract the corresponding output ⃗ x(i,j) O from N ⃗ mO|P ⃗ x(i,j) P , ˆΣO

  61. [69]

    for each input sample ⃗ x(i) I , estimate the entropy hO|I ⃗ x(i) I of the conditional distribution pst O|I from the output samples {⃗ xO}i,j, using any numerical estimator (e.g., Vasicek [56] or Kozachenko-Leonenko [57])

  62. [70]

    estimate the conditional entropy HO|I via importance sampling, HO|I ≈ Nsam,IX i=1 hO|I ⃗ x(i) I (S45)

  63. [71]

    from all the output samples {⃗ xO}i,j, estimate the entropy HO via any numerical estimator

  64. [72]

    compute the mutual information as IIO = HO − HO|I . Once more, since we focus on one-dimensional outputs, this sampling scheme avoids any issue with the curse of dimensionality, allowing us to explore large processing (and input) dimensions. 12 FIG. S3. Mutual information betw...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.