Pith. sign in

REVIEW 2 major objections 4 minor 3 cited by

Optimization and variability can coexist

T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Near-optimal biological function does not require fine tuning; variability is a predicted feature of sloppy performance landscapes.

desk verdict A genuinely useful survey of sloppy landscapes across five biological systems, but the headline scaling law rests on a quadratic approximation that breaks down on the very spectra that make it work. read the letter →

arxiv 2505.23398 v1 pith:E3EUNMC6 submitted 2025-05-29 q-bio.QM cond-mat.dis-nn

classification q-bio.QMcond-mat.dis-nn
keywords optimizationprinciplesloppymodelsparametervariabilitymaximumentropyHessianspectrumsoftmodesbiologicalphysicsinformationtheory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that optimization and variability in biology are not in conflict. Across five very different systems—photoreceptor arrangements, gene expression in fly embryos, ion-channel copy numbers in a rhythm-generating circuit, recurrent networks, and deep networks—the performance landscape near the optimum is 'sloppy': most parameter combinations are only weakly constrained, so many settings give nearly the same performance. The authors then make this quantitative: if a population's parameters vary as broadly as possible subject only to a constraint on average performance, then, for such sloppy landscapes, the entropy in parameter space can be extensive (finite entropy per parameter) even as average performance approaches the optimum arbitrarily closely. This removes the fine-tuning objection to optimization as a general principle and predicts that substantial parameter variability should be the norm rather than a sign of failure.

What carries the argument

The central object is the Hessian matrix $H_{ij} = -\partial^2 F/\partial \theta_i \partial \theta_j$ evaluated at the performance optimum; its eigenvalues $\lambda_\mu$ measure the stiffness of different parameter combinations, with small eigenvalues corresponding to soft modes. A 'sloppy' spectrum has eigenvalues falling off geometrically ($\lambda_\mu = \lambda_{\max} e^{-\alpha \mu}$), i.e., uniformly on a logarithmic scale. The argument is carried by the maximum-entropy distribution $P(\theta) \propto \exp(F(\theta)/T)$, a Boltzmann distribution with performance as negative energy, which yields the entropy $S = \frac{1}{2} \sum_\mu \ln(2\pi e T/\lambda_\mu)$ and the central identity Eq. (31) connecting $\Delta F$, $S$, $K$, $\lambda_{\max}$, and $\alpha$. This identity converts observed soft modes into a quantitative prediction about parameter variability at near-optimal average performance.

What would settle it

Measure, in a real system such as the stomatogastric ganglion or a laboratory selection experiment, both the Hessian eigenvalue spectrum and the joint distribution of parameters across individuals; if the observed parameter entropy falls far below the maximum-entropy value predicted from the measured spectrum and the observed average performance via Eq. (31), or if a high-dimensional system with near-optimal performance shows a narrow spectrum of eigenvalues with no soft tail, then the claimed coexistence of optimization and variability fails for that system.

Watch

Extended reading notes

Core claim

The paper's central claim is that near-optimal performance does not require fine tuning because the Hessian of functional performance at the optimum generically has a spectrum of eigenvalues spread over many decades, with a high density of small eigenvalues ('soft modes'). Using a maximum-entropy population model, the authors derive an exact relationship between the average performance gap $\Delta F$, the entropy $S$ of the parameter distribution, the number of parameters $K$, and the sloppiness exponent $\alpha$: $\Delta F = \frac{K\lambda_{\max}}{4\pi e}\,\exp\left(\frac{2S}{K} - \frac{\alpha(K+1)}{2}\right)$. For a sloppy spectrum of the form $\lambda_\mu = \lambda_{\max} e^{-\alpha \mu}$, this implies that as $K$ grows, one can hold a finite entropy per parameter while having $\Delta F \to 0$: optimization and variability coexist. The argument is illustrated with concrete systems in which the Hessian spectra are measured or computed, from retinal receptor arrays to trained deep networks.

Load-bearing premise

The prediction rests on the assumption that a population is as variable as possible subject only to the constraint on average performance; if mutation biases, history, or out-of-equilibrium dynamics actually shape parameter distributions, the predicted link between average performance and parameter entropy need not hold.

Editorial extensions

If this is right

  • Observed variation in biological parameters—protein copy numbers, synaptic strengths, receptor positions—is not evidence against optimization; it is what optimization on a sloppy landscape predicts.
  • In high-dimensional parameter spaces with sloppy spectra, populations should be spread across a large volume of parameter space while maintaining near-optimal average performance, making tightly controlled parameters the exception that demands explanation.
  • The maximum-entropy relation gives a quantitative tradeoff: at fixed average performance, the entropy per parameter that a population can sustain grows with the dynamic range of Hessian eigenvalues, so soft modes act as a resource for variability.
  • The results unify optimization arguments across gene regulation, neural circuits, and deep networks, suggesting a common principle rather than a collection of special cases.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the sloppy-spectrum identity holds, then measuring the Hessian spectrum of any high-dimensional fitness landscape immediately bounds how much variability a population can exhibit at a given fitness cost; this could be tested in laboratory evolution experiments where fitness landscapes are measurable.
  • The maximum-entropy premise is the soft spot: real populations are shaped by mutation bias and historical constraint, so the prediction is a baseline, and deviations from Eq. (31) could be used to quantify how strongly history constrains variation.
  • The same logic may explain protein family sequence entropy: structures that are stable for many sequences correspond to sloppy mappings, making extensive sequence entropy at almost no functional cost a generic feature of evolvable proteins.
  • A direct extension to trained deep networks: the scale-invariant tail of the Hessian spectrum implies that pruning or compressing networks by removing soft modes should be nearly lossless until the spectrum is truncated, which could be tested as a compression principle.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper argues that biological systems can be close to optimal while their parameters vary widely, because the performance landscape is sloppy: many parameter combinations are weakly constrained. It presents Hessian eigenvalue spectra for five systems (photoreceptor arrays, gap-gene readout in the fly embryo, a stomatogastric ganglion model, a linear recurrent network, and a deep ResNet on CIFAR-10) and shows that eigenvalues are roughly log-uniform over many decades. The theoretical part posits a maximum-entropy distribution of parameters at fixed mean performance and derives, for an exponential eigenvalue spectrum, Eq (31): the average performance deficit vanishes as K grows while the entropy per parameter remains finite. The authors conclude that optimization and variability coexist, and that variability is a prediction rather than a retreat from optimality.

Significance. If the central claim is correct, it resolves a long-standing tension between optimization principles and observed biological variability, and it gives a concrete, falsifiable scaling prediction. The paper's strengths are its breadth of examples, the explicit derivation through Eqs (1)-(31), and the use of data-driven Hessian spectra rather than purely synthetic models. The authors also provide a clearly stated null model (maximum entropy at fixed mean performance) and a specific prediction for how the performance-entropy tradeoff scales with dimensionality. The deep-network Fisher-information approximation in Box 5 is a useful methodological contribution. However, the main quantitative claim rests on a harmonic calculation whose self-consistency is not established; this is the principal obstacle to acceptance.

major comments (2)
  1. [Synthesis, Eqs (27)-(31)] Equation (31) is derived by combining the quadratic expansion (1) with the maximum-entropy distribution (26). For the exponential spectrum (29) and fixed entropy density S/K = s, Eq (28) gives T = λmax/(2πe) exp(2s − α(K+1)/2). The variance along the softest mode is then T/λmin = (1/(2πe)) exp(2s + α(K−1)/2), which diverges as K→∞. Thus typical fluctuations along the softest eigenvector grow without bound, so the Taylor expansion in Eq (1) is not controlled on the support of P(θ); Eqs (27) and (28) effectively treat F as globally quadratic. The conclusion that ΔF→0 with finite entropy per parameter is therefore not established by this harmonic calculation. The paper should either show that the relevant examples have globally quadratic performance, or replace Eq (31) with a derivation that includes anharmonic terms and states precisely when the local-Hessian prediction is valid.
  2. [Synthesis, Eq (26)] The maximum-entropy distribution is assumed rather than derived from evolutionary, developmental, or learning dynamics. The text's claim that 'the only consistent way' to make 'as variable as possible' precise is maximum entropy is a modeling assumption; mutation biases, selection dynamics, and historical constraints can produce distributions with lower entropy at the same mean performance. The quantitative prediction Eq (31) depends on this assumption. The paper should explicitly frame Eq (31) as a prediction of the max-entropy null model and, ideally, test whether the observed systems' parameter distributions satisfy the predicted entropy-performance relation, rather than reporting only the Hessian spectra.
minor comments (4)
  1. [Box 3] The last sentence of Box 3 is incomplete: 'taking the form of a sigmoid In small circuits...' should be completed or rephrased.
  2. [Box 2] The notation 'θ has ||Z|| = 2N dimensions' conflates the number of Voronoi centers with the number of discrete regions ||Z||. If there are N centers each with two components, the parameter dimension is 2N, not ||Z||; please clarify.
  3. [Figures 3 and 5] The Hessian spectra in Figs 3C and 5C are plotted without uncertainty estimates. Since these spectra are estimated from numerical simulations and finite sample averages, including bootstrap or other error bars would strengthen the claims of convergence and reproducibility.
  4. [Eq (32)] For Eq (32), the text says finite entropy per parameter requires the dynamic range of ln λ to grow with K, but the integral includes λmin without a corresponding criterion relating λmin to K. Please state the precise condition (e.g., λmin ~ e^{-cK}) and its implications for the ρ(λ)~1/λ spectra shown in Fig 5.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Eq (31) follows from the stated maximum-entropy and quadratic-expansion assumptions, with spectra measured independently.

full rationale

The central derivation, Eqs (26)-(31), is a closed-form mathematical consequence of two explicitly stated assumptions: the local quadratic expansion of the performance function, Eq (1), and the maximum-entropy ensemble consistent with a fixed mean performance, Eq (26). Equation (27) is the standard equipartition result for a Gaussian Boltzmann distribution, and Eq (28) is the corresponding Gaussian entropy; substituting the exponential spectrum (29) into these identities produces Eq (31) directly. No parameter is fitted to the entropy or to the predicted variability, and no equation reduces to an earlier fitted value. The Hessian spectra in the biological examples are computed from data or simulations independently of the entropy formula, so the examples serve as evidence for sloppy spectra rather than as inputs that force the conclusion. The self-citations that appear, e.g. [26] for the natural-image power spectrum and [35] for the information-bottleneck/Voronoi solution used in the fly example, are not load-bearing for the main entropy-performance identity: those results are externally measurable or mathematically derived, and the central argument would survive if the examples were replaced. The possible failure of the quadratic expansion on exponentially soft modes, raised by the skeptic, is an internal-consistency or correctness concern about the derivation's domain of validity, not a circularity. No circular step is present.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the maximum entropy model of population variability, the quadratic approximation of the performance landscape, and the assumption that a sloppy exponential eigenvalue spectrum is representative. The spectra themselves are computed from models or data rather than postulated, but the max entropy distribution is an assumption about evolutionary exploration.

free parameters (3)
  • effective temperature T = not estimated
    Introduced in Eq (26) as the Lagrange multiplier enforcing the mean-performance constraint Eq (25); the paper gives no value and it is eliminated in Eq (31). It is a chosen modeling parameter rather than a measured biological quantity.
  • dimensionless signal-to-noise ratio SNR = varied over range
    In the retina example (Box 1), SNR is the only free parameter and is scanned to match human cone estimates; the qualitative result holds across the range.
  • rank n of Fisher approximation = 1500
    In Box 5, the deep-network Hessian is approximated with n=1500 samples; the choice affects the spectrum, although convergence is claimed.
assumptions (5)
  • domain assumption Maximum entropy distribution for parameter variability
    Equation (26) asserts P(θ) ∝ exp(F/T); this assumes populations explore parameter space as widely as possible under a single mean-performance constraint, with no other costs or history.
  • domain assumption Quadratic truncation of performance landscape
    Equation (1) keeps only the Hessian term; if the landscape is nonsmooth or the optimum is degenerate with zero eigenvalues, the entropy formula in Eq (28) breaks down.
  • ad hoc to paper Sloppy spectrum with exponential eigenvalue spacing
    Equation (29), λμ = λmax e^{-αμ}, is an idealized log-uniform spectrum used to derive extensivity; the examples approximate it but none realizes it exactly.
  • domain assumption Natural image power law spectrum
    Box 1 uses S(k) = A/|k|^{2-η} with η ≈ 0.2 from Ref [26]; this is an empirical input for the retina calculation.
  • standard math Fisher information equals Hessian at zero training loss
    Equation (19) identifies the Hessian with the Fisher information when the model achieves perfect classification; this is a standard result in statistics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Optimization and variability can coexist." pith.science (2026). https://pith.science/paper/E3EUNMC6

@misc{pith2026250523398,
  author       = {Pith},
  title        = {Pith review of: Optimization and variability can coexist},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E3EUNMC6}},
  note         = {Machine review of arXiv:2505.23398}
}
read the original abstract

Many biological systems perform close to their physical limits, but promoting this optimality to a general principle seems to require implausibly fine tuning of parameters. Using examples from a wide range of systems, we show that this intuition is wrong. Near an optimum, functional performance depends on parameters in a "sloppy'' way, with some combinations of parameters being only weakly constrained. Absent any other constraints, this predicts that we should observe widely varying parameters, and we make this precise: the entropy in parameter space can be extensive even if performance on average is very close to optimal. This removes a major objection to optimization as a general principle, and rationalizes the observed variability.

Figures

Figures reproduced from arXiv: 2505.23398 by the authors.

Figure 1
Figure 1. FIG. 1: Sampling the visual world. (a) A segment of the fly’s [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2: Optimizing the response to morphogens. (a) (left) [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3: Ion channel copy numbers in a small neural cir [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: FIG. 4: A linear recurrent network of [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5: Eigenvalue spectrum of the Fisher information matrix [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Invariant non-equilibrium dynamics of transcriptional regulation optimize information flow

    q-bio.MN 2025-07 conditional novelty 7.0 of 10

    A four-state non-equilibrium promoter model reproduces the invariant switching correlation time seen in Drosophila and links it to maximized information transmission under a switching speed limit.

  2. Social-spatial dependencies for learning visual navigation

    cs.NE 2026-07 conditional novelty 6.0 of 10

    Neural-network agents trained in social environments learn hybrid navigation strategies that combine individual landmark use with social following, with strategy shifts driven by the ratio of skilled to unskilled soci...

  3. Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    RNNs can sustain power-law forgetting and multi-time-scale learning when heavy-tailed fluctuations in SGD balance the collapse tendency toward short time scales, governed by a spectral exponent β.

Reference graph

Works this paper leans on

74 extracted references · 71 canonical work pages · cited by 3 Pith papers

  1. [1]

    exp −|⃗ x|2/2x2 0 , and assume that the width of this blur is matched to the over- all density of receptorsρ= 1/x 2

  2. [2]

    morphogen

    In this setting we find that the mutual information between the set of responses{R n}and the imageϕ(⃗ x) is I= 1 2 log2 det 1+ ˆC ,(5) and theN×Nmatrix ˆChas elements Cnm = A xη 0σ2 Z ∞ 0 dz 2π e−z2 z1−η J0(zrnm/x0),(6) whereJ 0 is a Bessel function andr nm =|⃗ xn −⃗ xm|. We can choosex 0 as the unit of distance, so the only free parameter is a dimensionl...

  3. [3]

    & Baylor, D

    Rieke, F. & Baylor, D. A. Single photon detection by rod cells of the retina.Reviews of Modern Physics70, 1027–1036 (1998)

  4. [4]

    Barlow, H. B. The size of ommatidia in apposition eyes. Journal of Experimental Biology29, 667–674 (1952)

  5. [5]

    Snyder, A. W. Acuity of compound eyes: Physical lim- itations and design.Journal of Comparative Physiology 116, 161–182 (1977)

  6. [6]

    Berg, H. C. & Purcell, E. M. Physics of chemoreception. Biophysical Journal20, 193–219 (1977)

  7. [7]

    Bialek, W.Biophysics: Searching for Principles(Prince- ton University Press, Princeton NJ, 2012)

  8. [8]

    Ambitions for theory in the physics of life

    Bialek, W. Ambitions for theory in the physics of life. In Bitbol, A.-F., Mora, T., Nemenman, I. & Walczak, A. M. (eds.)Les Houches Summer School Lecture Notes (SciPost Physics Lecture Notes 84, 2024)

Show all 74 references
  1. [9]

    Gould, S. J. & Lewontin, R. C. The spandrels of san marco and the panglossian paradigm: A critique of the adaptationist programme.Proceedings of the Royal So- ciety B: Biological Sciences205, 581–598 (1979)

  2. [10]

    Spudich, J. L. & Koshland, D. E., Jr. Non-genetic indi- viduality: Chance in the single cell.Nature262, 467–471 (1976)

  3. [11]

    A., Bucher, D

    Prinz, A. A., Bucher, D. & Marder, E. Similar network activity from disparate circuit parameters.Nature Neu- roscience7, 1345–1352 (2004). 11

  4. [12]

    & Goaillard, J

    Marder, E. & Goaillard, J. M. Variability, compensation and homeostasis in neuron and network function.Nature Reviews Neuroscience7, 563–574 (2006)

  5. [13]

    The roles of mutation, inbreeding, crossbreed- ing and selection in evolution

    Wright, S. The roles of mutation, inbreeding, crossbreed- ing and selection in evolution. In Jones, D. F. (ed.)Pro- ceedings of the Sixth International Congress of Genetics, vol I, 356–366 (Genetics Society of America, 1932)

  6. [14]

    G., Schenk, M

    Szendro, I. G., Schenk, M. F., Franke, J., Krug, J. & de Visser, J. A. G. M. Quantitative analyses of empir- ical fitness landscapes.Journal of Statistical Mechanics P01005 (2013)

  7. [15]

    K.et al.Perspective: Sloppiness and emergent theories in physics, biology, and beyond.Jour- nal of Chemical Physics143, 010901 (2015)

    Transtrum, M. K.et al.Perspective: Sloppiness and emergent theories in physics, biology, and beyond.Jour- nal of Chemical Physics143, 010901 (2015)

  8. [16]

    Brown, K. S. & Sethna, J. P. Statistical mechanical ap- proaches to models with many poorly known parameters. Physical Review E68, 021904 (2003)

  9. [17]

    N.et al.Universally sloppy parameter sensitivities in systems biology models.PLoS Computa- tional Biology3, e189 (2007)

    Gutenkunst, R. N.et al.Universally sloppy parameter sensitivities in systems biology models.PLoS Computa- tional Biology3, e189 (2007)

  10. [18]

    J.et al.Sloppy–model universality class and the Vandermonde matrix.Physical Review Letters 97, 150601 (2006)

    Waterfall, J. J.et al.Sloppy–model universality class and the Vandermonde matrix.Physical Review Letters 97, 150601 (2006)

  11. [19]

    N., Wilber, H., Townsend, A

    Quinn, K. N., Wilber, H., Townsend, A. & Sethna, J. P. Chebyshev approximation and the global geometry of model predictions.Physical Review Letters122, 158302 (2019)

  12. [20]

    N., Abbott, M

    Quinn, K. N., Abbott, M. C., Transtrum, M. K., Machta, B. B. & Sethna, J. P. Information geometry for mul- tiparameter models: New perspectives on the origin of simplicity.Reports on Progress in Physics86, 035901 (2022)

  13. [21]

    R., Gregor, T., Bialek, W

    Sokolowski, T. R., Gregor, T., Bialek, W. & Tkaˇ cik, G. Deriving a genetic regulatory network from an optimiza- tion principle.Proceedings of the National Academy of Sciences (USA)122, e2402925121 (2025)

  14. [22]

    Verlag (Springer, Berlin, 1976)

    Strausfeld, N.Atlas of an Insect Brain. Verlag (Springer, Berlin, 1976)

  15. [23]

    B., Lennie, P

    Roorda, A., Mehta, A. B., Lennie, P. & Williams, D. R. Packing arrangement of the three cone classes in primate retina.Vision Research41, 1291–1306 (2001)

  16. [24]

    Sz´ el,´A., R¨ ohlich, P., Caff´ e, A. R. & van Veen, T. Distri- bution of cone photoreceptors in the mammalian retina. Microscopy Research and Technique35, 445–462 (1996)

  17. [25]

    M., Bui, D

    Sherry, D. M., Bui, D. D. & Degrip, W. J. Identifica- tion and distribution of photoreceptor subtypes in the neotenic tiger salamander retina.Visual Neuroscience 15, 1175–1187 (1998)

  18. [26]

    Jiao, Y.et al.Avian photoreceptor patterns represent a disordered hyperuniform solution to a multiscale packing problem.Physical Review E89, 022721 (2014)

  19. [27]

    Leertouwer, H. L. Cover photograph, June 28.Science 252(1991)

  20. [28]

    Ruderman, D. L. & Bialek, W. Statistics of natural im- ages: Scaling in the woods.Physical Review Letters73, 814–817 (1994)

  21. [29]

    Positional information and the spatial pat- tern of cellular differentiation.Journal of Theoretical Bi- ology25, 1–47 (1969)

    Wolpert, L. Positional information and the spatial pat- tern of cellular differentiation.Journal of Theoretical Bi- ology25, 1–47 (1969)

  22. [30]

    O., Tkaˇ cik, G., Wieschaus, E

    Dubuis, J. O., Tkaˇ cik, G., Wieschaus, E. F., Gregor, T. & Bialek, W. Positional information, in bits.Proceedings of the National Academy of Sciences (USA)110, 16301– 16308 (2013)

  23. [31]

    O., Petkova, M

    Tkaˇ cik, G., Dubuis, J. O., Petkova, M. D. & Gregor, T. Positional information, positional error, and read–out precision in morphogenesis: a mathematical framework. Genetics199, 39–59 (2015)

  24. [32]

    & Gregor, T

    Tkaˇ cik, G. & Gregor, T. The many bits of positional information.Development148, dev176065 (2021)

  25. [33]

    A.The Making of a Fly: The Genetics of Animal Design(Blackwell Scientific, Oxford, 1992)

    Lawrence, P. A.The Making of a Fly: The Genetics of Animal Design(Blackwell Scientific, Oxford, 1992)

  26. [34]

    & N¨ usslein-Volhard, C

    Wieschaus, E. & N¨ usslein-Volhard, C. The Heidelberg screen for pattern mutants ofdrosophila: A personal ac- count.Annual Review of Cell and Developmental Biology 32, 1–46 (2016)

  27. [35]

    D., Tkaˇ cik, G., Bialek, W., Wieschaus, E

    Petkova, M. D., Tkaˇ cik, G., Bialek, W., Wieschaus, E. F. & Gregor, T. Optimal decoding of cellular identities in a genetic network.Cell176, 844–855 (2019)

  28. [36]

    The gap gene network.Cellular and Molecular Life Sciences68, 243–274 (2011)

    Jaeger, J. The gap gene network.Cellular and Molecular Life Sciences68, 243–274 (2011)

  29. [37]

    D., Gregor, T., Wieschaus, E

    Bauer, M., Petkova, M. D., Gregor, T., Wieschaus, E. F. & Bialek, W. Trading bits in the readout from a ge- netic network.Proceedings of the National Academy of Sciences (USA)118, e2109011118 (2021)

  30. [38]

    & Bialek, W

    Bauer, M. & Bialek, W. Information bottleneck in molec- ular sensing.PRX Life1, 023005 (2023)

  31. [39]

    McGough, L.et al.Finding the last bits of positional information.PRX Life2, 013016 (2024)

  32. [40]

    Furlong, E. E. & Levine, M. Developmental enhancers and chromosome topology.Science361, 1341 (2018)

  33. [41]

    D., Gregor, T., Wieschaus, E

    Bauer, M., Petkova, M. D., Gregor, T., Wieschaus, E. F. & Bialek, W. Trading bits in the readout from a genetic network (2020). arXiv:2012.15817 [q-bio.MN]

  34. [42]

    Tishby, N., Pereira, F. C. & Bialek, W. The informa- tion bottleneck method. In Hajek, B. & Sreenivas, R. S. (eds.)Proceedings of the 37th Annual Allerton Confer- ence on Communication, Control and Computing, 368– 377 (University of Illinois, 1999). arXiv:physics/0004057 [phys...

  35. [43]

    Johnston, D. & Wu, S. M.-S.Foundations of Cellular Neurophysiology(MIT Press, Cambridge MA, 1994)

  36. [44]

    Bevan, M. D. & Wilson, C. J. Mechanisms underlying spontaneous oscillation and rhythmic firing in rat subtha- lamic neurons.Journal of Neuroscience19, 7617–7628 (1999)

  37. [45]

    & Abbott, L

    LeMasson, G., Marder, E. & Abbott, L. F. Activity– dependent regulation of conductances in model neurons. Science259, 1915–1917 (1993)

  38. [46]

    J.The Physiology of Excitable Cells(Cam- bridge University Press, Sunderland MA and Cambridge UK, 1998)

    Aidley, D. J.The Physiology of Excitable Cells(Cam- bridge University Press, Sunderland MA and Cambridge UK, 1998). Third Edition (2001) Editor B. Hille,Ion Channels of Excitable Membranes

  39. [47]

    Abbott, L. F. & Marder, E. Modeling small networks (1998)

  40. [48]

    & Marder, E

    Gorur-Shandilya, S., Hoyland, A. & Marder, E. Xolotl: an intuitive and approachable neuron and network simu- lator for research and teaching.Frontiers in neuroinfor- matics12, 87 (2018)

  41. [49]

    Hennequin, G., Vogels, T. P. & Gerstner, W. Optimal control of transient dynamics in balanced networks sup- ports generation of complex movements.Neuron82, 1394–1406 (2014)

  42. [50]

    M., Kaufman, M

    Sussillo, D., Churchland, M. M., Kaufman, M. T. & Shenoy, K. V. A neural network that finds a naturalis- tic solution for the production of muscle activity.Nature Neuroscience18, 1025–1033 (2015)

  43. [51]

    & Barak, 12 O

    Susman, L., Mastrogiuseppe, F., Brenner, N. & Barak, 12 O. Quality of internal representation shapes learning per- formance in feedback neural networks.Physical Review Research3, 013176 (2021)

  44. [52]

    & Hinton, G

    LeCun, Y., Bengio, Y. & Hinton, G. Deep learning.Na- ture521, 436–444 (2015)

  45. [53]

    In Chiappa, S

    Thomas, V.et al.On the interplay between noise and curvature and its effect on optimization and generaliza- tion. In Chiappa, S. & Calandra, R. (eds.)Proceedings of the Twenty Third International Conference on Artifi- cial Intelligence and Statistics, vol. 108 ofProceedings of...

  46. [54]

    U., Dauphin, Y

    Sagun, L., Evci, U., G¨ uney, V. U., Dauphin, Y. N. & Bottou, L. Empirical analysis of the hessian of over- parametrized neural networks. In6th International Con- ference on Learning Representations, ICLR 2018, Van- couver, BC, Canada, April 30 - May 3, 2018, Work- shop Track ...

  47. [55]

    & Sun, J

    He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition.arXiv:1512.03385(2015). [cs.CV]

  48. [56]

    Proper ResNet implementation for CI- F AR10/CIF AR100 in PyTorch.https://github.com/ akamaster/pytorch_resnet_cifar10(2018)

    Idelbayev, Y. Proper ResNet implementation for CI- F AR10/CIF AR100 in PyTorch.https://github.com/ akamaster/pytorch_resnet_cifar10(2018)

  49. [57]

    Learning multiple layers of features from tiny images.https://www.cs.toronto.edu/ ~kriz/ cifar.html(2009)

    Krishevsky, A. Learning multiple layers of features from tiny images.https://www.cs.toronto.edu/ ~kriz/ cifar.html(2009)

  50. [58]

    Gur-Ari, G., Roberts, D. A. & Dyer, E. Gradient descent happens in a tiny subspace (2018). arXiv:1812.04754 [cs.LG]

  51. [59]

    Shannon, C. E. A mathematical theory of communica- tion.Bell System Technical Journal27, 379–423 (1948). & 623–656

  52. [60]

    Jaynes, E. T. Information theory and statistical mechan- ics.Physical Review106, 620–630 (1957)

  53. [61]

    Stroud, R. M. A family of protein cutting proteins.Sci- entific American231, 74–88 (1974)

  54. [62]

    E. L. L. Sonnhammer, S. R. E. & Durbin, R. Pfam: A comprehensive database of protein families based on seed alignments.Proteins28, 405 (1997)

  55. [63]

    Blum, M.et al.The interpro protein families and do- mains database: 20 years on.Nucleic Acids Research 49, D344–D354 (2020)

  56. [64]

    G., Liu, L

    Lapedes, A., Giraud, B. G., Liu, L. C. & Stormo, G. A maximum entropy formalism for disentangling chains of correlated sequence positions.LANL Report98–1094 (1998). LA-UR

  57. [65]

    & Ranganathan, R

    Bialek, W. & Ranganathan, R. Rediscovering the power of pairwise interactions (2007). arXiv:0712.4397 [q- bio.QM]

  58. [66]

    A., Szurmant, H., Hoch, J

    Weigt, M., White, R. A., Szurmant, H., Hoch, J. A. & Hwa, T. Identification of direct residue contacts in protein–protein interaction by message passing.Proceed- ings of the National Academy of Sciences (USA)106, 67–72 (2009)

  59. [67]

    S.et al.Protein 3d structure computed from evolutionary sequence variation.PLoS One6, e28766 (2011)

    Marks, D. S.et al.Protein 3d structure computed from evolutionary sequence variation.PLoS One6, e28766 (2011)

  60. [68]

    P.et al.An evolution-based model for design- ing chorismate mutase enzymes.Science369, 440–445 (2020)

    Russ, W. P.et al.An evolution-based model for design- ing chorismate mutase enzymes.Science369, 440–445 (2020)

  61. [69]

    P., Chakraborty, A

    Barton, J. P., Chakraborty, A. K., Cocco, S., Jacquin, H. & Monasson, R. On the entropy of protein families. Journal of Statistical Physics162, 1267–1293 (2016)

  62. [70]

    & Wingreen, N

    Li, H., Helling, R., Tang, C. & Wingreen, N. Emergence of preferred structures in a simple model of protein fold- ing.Science273, 666–669 (1996)

  63. [71]

    & Wingreen, N

    Li, H., Tang, C. & Wingreen, N. S. Are protein folds atypical?Proceedings of the National Academy of Sci- ences (USA)95, 4987–4990 (1998)

  64. [72]

    & Hartl, D

    Carneiro, M. & Hartl, D. L. Adaptive landscapes and protein evolution.Proceedings of the National Academy of Sciences (USA)107, 1747–1751 (2010)

  65. [73]

    & Wagner, A

    Papkou, A., Garcia-Pastor, L., Escudero, J. & Wagner, A. A rugged yet easily navigable fitness landscape.Sci- ence382, eadh3860 (2023)

  66. [74]

    & Hermundstad, A

    Ma, T. & Hermundstad, A. M. A vast space of compact strategies for effective decisions.Science Advances10, eadj4064 (2024)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.