{"id":"0faea7e4-cfba-4b6c-9ca9-96aa3a293002","arxiv_id":"2505.23398","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Near an optimum, performance landscapes have many soft parameter directions, so the entropy of parameter variation can be extensive even when average performance is essentially optimal.","lead":"This paper argues that biological systems can sit very close to optimal performance while their underlying parameters vary widely, because performance depends on many parameter combinations only weakly. If right, it removes a common objection to the idea that evolution optimizes function, while also explaining observed variability.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq (31) rests on the quadratic expansion holding over the full max-entropy distribution; for exponential spectra the softest-mode variance T/λmin diverges as K→∞, so the derivation is not self-consistent.","rationale":"The paper's central claim is formalized by Eq (31): for exponential sloppy spectra, the average performance gap can vanish while the entropy density stays finite. The derivation is a chain of equalities—quadratic expansion (1), max-entropy form (26), equipartition (27), entropy (28), and the exponential spectrum (29). Each step is algebraically correct if F is globally quadratic. The problem is that the asymptotic statement requires K→∞ with fixed s = S/K, which forces T ~ e^{-αK/2}. The variance along the softest mode then scales as T/λmin ~ e^{αK/2}, so the typical configurations of the max-entropy distribution move far outside the regime where the Hessian controls the landscape. Equations (27) and (28) are therefore not self-consistent for the very spectra used to reach the conclusion. This is a load-bearing gap: without a bound on anharmonic corrections or a nonlinear treatment of the soft modes, the paper's precise claim is not established by the mathematics. The empirical spectra in Figs 3–5 and the retina example support the qualitative idea that soft modes allow variability, and I do not dispute that near-optimal performance can coexist with variation. But the quantitative statement (31) requires the quadratic approximation to hold over the support of the ensemble, and that condition demonstrably fails in the K→∞ limit for the exponential spectrum. This concern is distinct from the reader's emphasis on the maximum-entropy assumption: even granting Eq (26), the harmonic approximation is internally inconsistent in the limit central to the paper. The reader's CONDITIONAL verdict is therefore appropriate, but the requested revision should include a self-consistency check of the harmonic expansion (or a non-perturbative treatment of soft modes), rather than only addressing the maximum-entropy assumption and code availability.","tokens_in":15136,"tokens_out":15417,"duration_ms":149297,"concrete_test":"Simulate a K-dimensional toy landscape F(θ)=−(1/2)Σ_{μ=1}^K e^{-αμ} θ_μ^2 − c Σ θ_μ^4 with fixed α>0 and c>0. For K=20, 40, 80, find T such that S/K = s (estimate S by Monte Carlo from P∝exp(F/T)), and compute ΔF exactly. Compare to Eq (31); if the exact ΔF does not vanish as K grows, or deviates increasingly from Eq (31), the quadratic approximation is the point of failure. A c=0 control should reproduce Eq (31), isolating the role of anharmonicity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equations (27)–(31) follow from the quadratic expansion (1) and the maximum-entropy distribution (26). For the exponential spectrum (29), holding the entropy density s = S/K fixed requires T = λmax/(2πe) exp(2s − α(K+1)/2). The variance along the softest mode is then T/λmin = (2πe)^{-1} exp(2s + α(K−1)/2), which diverges as K→∞. Thus typical fluctuations along the softest eigenvector grow without bound, so the Taylor expansion (1) is not controlled on the support of P(θ). Equations (28) and (27) presuppose a globally quadratic F along every mode; for the very spectra that make ΔF in (31) vanish, this presupposition fails. The central claim—finite entropy per parameter while average performance approaches optimal—is therefore not supported by the harmonic derivation; it rests on an uncontrolled extrapolation of the local Hessian to exponentially soft modes. This is a mathematical gap internal to the argument, distinct from the biological plausibility of the maximum-entropy ensemble.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that biological systems can be close to optimal while their parameters vary widely, because the performance landscape is sloppy: many parameter combinations are weakly constrained. It presents Hessian eigenvalue spectra for five systems (photoreceptor arrays, gap-gene readout in the fly embryo, a stomatogastric ganglion model, a linear recurrent network, and a deep ResNet on CIFAR-10) and shows that eigenvalues are roughly log-uniform over many decades. The theoretical part posits a maximum-entropy distribution of parameters at fixed mean performance and derives, for an exponential eigenvalue spectrum, Eq (31): the average performance deficit vanishes as K grows while the entropy per parameter remains finite. The authors conclude that optimization and variability coexist, and that variability is a prediction rather than a retreat from optimality.","tokens_in":15377,"tokens_out":6141,"duration_ms":69488,"significance":"If the central claim is correct, it resolves a long-standing tension between optimization principles and observed biological variability, and it gives a concrete, falsifiable scaling prediction. The paper's strengths are its breadth of examples, the explicit derivation through Eqs (1)-(31), and the use of data-driven Hessian spectra rather than purely synthetic models. The authors also provide a clearly stated null model (maximum entropy at fixed mean performance) and a specific prediction for how the performance-entropy tradeoff scales with dimensionality. The deep-network Fisher-information approximation in Box 5 is a useful methodological contribution. However, the main quantitative claim rests on a harmonic calculation whose self-consistency is not established; this is the principal obstacle to acceptance.","major_comments":[{"comment":"Equation (31) is derived by combining the quadratic expansion (1) with the maximum-entropy distribution (26). For the exponential spectrum (29) and fixed entropy density S/K = s, Eq (28) gives T = λmax/(2πe) exp(2s − α(K+1)/2). The variance along the softest mode is then T/λmin = (1/(2πe)) exp(2s + α(K−1)/2), which diverges as K→∞. Thus typical fluctuations along the softest eigenvector grow without bound, so the Taylor expansion in Eq (1) is not controlled on the support of P(θ); Eqs (27) and (28) effectively treat F as globally quadratic. The conclusion that ΔF→0 with finite entropy per parameter is therefore not established by this harmonic calculation. The paper should either show that the relevant examples have globally quadratic performance, or replace Eq (31) with a derivation that includes anharmonic terms and states precisely when the local-Hessian prediction is valid.","section":"Synthesis, Eqs (27)-(31)"},{"comment":"The maximum-entropy distribution is assumed rather than derived from evolutionary, developmental, or learning dynamics. The text's claim that 'the only consistent way' to make 'as variable as possible' precise is maximum entropy is a modeling assumption; mutation biases, selection dynamics, and historical constraints can produce distributions with lower entropy at the same mean performance. The quantitative prediction Eq (31) depends on this assumption. The paper should explicitly frame Eq (31) as a prediction of the max-entropy null model and, ideally, test whether the observed systems' parameter distributions satisfy the predicted entropy-performance relation, rather than reporting only the Hessian spectra.","section":"Synthesis, Eq (26)"}],"minor_comments":[{"comment":"The last sentence of Box 3 is incomplete: 'taking the form of a sigmoid In small circuits...' should be completed or rephrased.","section":"Box 3"},{"comment":"The notation 'θ has ||Z|| = 2N dimensions' conflates the number of Voronoi centers with the number of discrete regions ||Z||. If there are N centers each with two components, the parameter dimension is 2N, not ||Z||; please clarify.","section":"Box 2"},{"comment":"The Hessian spectra in Figs 3C and 5C are plotted without uncertainty estimates. Since these spectra are estimated from numerical simulations and finite sample averages, including bootstrap or other error bars would strengthen the claims of convergence and reproducibility.","section":"Figures 3 and 5"},{"comment":"For Eq (32), the text says finite entropy per parameter requires the dynamic range of ln λ to grow with K, but the integral includes λmin without a corresponding criterion relating λmin to K. Please state the precise condition (e.g., λmin ~ e^{-cK}) and its implications for the ρ(λ)~1/λ spectra shown in Fig 5.","section":"Eq (32)"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the paper is a broad, thought-provoking synthesis and the examples are genuinely illustrative. The main obstacle is mathematical, not empirical: the derivation of Eq (31) is not self-consistent for exponential spectra because the soft-mode variance diverges. This is fixable by either adding an explicit anharmonic treatment or reframing the claim as a conjecture supported by the examples, but it needs to be addressed before publication. I saw no citation-ethics concerns; the self-citations are clearly tied to prior data and methods and are appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about 2505.23398. First, the paper does something genuinely new: it recasts sloppiness as a property of functional performance, not model fit, and it shows across five very different systems—retina, fly embryo, stomatogastric ganglion, recurrent networks, deep nets—that Hessian spectra are roughly log-uniform. That synthesis is the most valuable part, and the examples are well chosen. Second, the central scaling law, Eq (31), has a self-consistency problem that the paper does not address.\n\nFor an exponential spectrum λμ = λmax e^{-αμ}, holding the entropy density S/K fixed forces the effective temperature T to grow exponentially with K. The variance along the softest mode is then T/λmin, which also grows exponentially. So typical fluctuations along that mode are huge, and the quadratic expansion of F around the optimum is not controlled on the support of the max-entropy distribution. Equations (27)–(31) presuppose a globally quadratic F; for the very spectra that make ΔF vanish, this is not a valid approximation unless the landscape is exactly quadratic. That is a real gap in the formal argument, distinct from the biological plausibility of the max-entropy assumption.\n\nOn the empirical side, the spectra in Figures 1–5 have no error bars, and no code or data are provided, so independent reproduction is hard. The deep-network Hessian is estimated with a finite-rank approximation; they show convergence, but explicit checks would help. The max-entropy distribution in Eq (26) is assumed, not derived from evolutionary dynamics; the authors acknowledge this is a hypothesis, but the title and abstract present the coexistence result more strongly than the assumption warrants.\n\nWho should read this? People working on sloppy models, optimization principles in biology, and the geometry of learning. The examples and the entropy/performance tradeoff idea deserve a serious referee. But the formal result needs to be fixed or properly qualified—for instance by including anharmonicity or stating conditions under which the local Hessian controls the full distribution. I would send it to peer review, with the expectation of major revision.","headline":"A genuinely useful survey of sloppy landscapes across five biological systems, but the headline scaling law rests on a quadratic approximation that breaks down on the very spectra that make it work.","tokens_in":15868,"tokens_out":4979,"would_cite":false,"duration_ms":53502,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Near-optimal biological function does not require fine tuning; variability is a predicted feature of sloppy performance landscapes.","keywords":["optimization principle","sloppy models","parameter variability","maximum entropy","Hessian spectrum","soft modes","biological physics","information theory"],"falsifier":"Measure, in a real system such as the stomatogastric ganglion or a laboratory selection experiment, both the Hessian eigenvalue spectrum and the joint distribution of parameters across individuals; if the observed parameter entropy falls far below the maximum-entropy value predicted from the measured spectrum and the observed average performance via Eq. (31), or if a high-dimensional system with near-optimal performance shows a narrow spectrum of eigenvalues with no soft tail, then the claimed coexistence of optimization and variability fails for that system.","tokens_in":14979,"feed_emoji":"🧬","tokens_out":8681,"duration_ms":87405,"temperature":0.7,"pith_summary":"This paper argues that optimization and variability in biology are not in conflict. Across five very different systems—photoreceptor arrangements, gene expression in fly embryos, ion-channel copy numbers in a rhythm-generating circuit, recurrent networks, and deep networks—the performance landscape near the optimum is 'sloppy': most parameter combinations are only weakly constrained, so many settings give nearly the same performance. The authors then make this quantitative: if a population's parameters vary as broadly as possible subject only to a constraint on average performance, then, for such sloppy landscapes, the entropy in parameter space can be extensive (finite entropy per parameter) even as average performance approaches the optimum arbitrarily closely. This removes the fine-tuning objection to optimization as a general principle and predicts that substantial parameter variability should be the norm rather than a sign of failure.","feed_headline":"Sloppy landscapes let organisms stay near-optimal and variable","feed_subtitle":"A maximum-entropy argument shows parameter entropy can stay high while average performance approaches its optimum.","key_machinery":"The central object is the Hessian matrix $H_{ij} = -\\partial^2 F/\\partial \\theta_i \\partial \\theta_j$ evaluated at the performance optimum; its eigenvalues $\\lambda_\\mu$ measure the stiffness of different parameter combinations, with small eigenvalues corresponding to soft modes. A 'sloppy' spectrum has eigenvalues falling off geometrically ($\\lambda_\\mu = \\lambda_{\\max} e^{-\\alpha \\mu}$), i.e., uniformly on a logarithmic scale. The argument is carried by the maximum-entropy distribution $P(\\theta) \\propto \\exp(F(\\theta)/T)$, a Boltzmann distribution with performance as negative energy, which yields the entropy $S = \\frac{1}{2} \\sum_\\mu \\ln(2\\pi e T/\\lambda_\\mu)$ and the central identity Eq. (31) connecting $\\Delta F$, $S$, $K$, $\\lambda_{\\max}$, and $\\alpha$. This identity converts observed soft modes into a quantitative prediction about parameter variability at near-optimal average performance.","core_discovery":"The paper's central claim is that near-optimal performance does not require fine tuning because the Hessian of functional performance at the optimum generically has a spectrum of eigenvalues spread over many decades, with a high density of small eigenvalues ('soft modes'). Using a maximum-entropy population model, the authors derive an exact relationship between the average performance gap $\\Delta F$, the entropy $S$ of the parameter distribution, the number of parameters $K$, and the sloppiness exponent $\\alpha$: $\\Delta F = \\frac{K\\lambda_{\\max}}{4\\pi e}\\,\\exp\\left(\\frac{2S}{K} - \\frac{\\alpha(K+1)}{2}\\right)$. For a sloppy spectrum of the form $\\lambda_\\mu = \\lambda_{\\max} e^{-\\alpha \\mu}$, this implies that as $K$ grows, one can hold a finite entropy per parameter while having $\\Delta F \\to 0$: optimization and variability coexist. The argument is illustrated with concrete systems in which the Hessian spectra are measured or computed, from retinal receptor arrays to trained deep networks.","pith_inferences":["If the sloppy-spectrum identity holds, then measuring the Hessian spectrum of any high-dimensional fitness landscape immediately bounds how much variability a population can exhibit at a given fitness cost; this could be tested in laboratory evolution experiments where fitness landscapes are measurable.","The maximum-entropy premise is the soft spot: real populations are shaped by mutation bias and historical constraint, so the prediction is a baseline, and deviations from Eq. (31) could be used to quantify how strongly history constrains variation.","The same logic may explain protein family sequence entropy: structures that are stable for many sequences correspond to sloppy mappings, making extensive sequence entropy at almost no functional cost a generic feature of evolvable proteins.","A direct extension to trained deep networks: the scale-invariant tail of the Hessian spectrum implies that pruning or compressing networks by removing soft modes should be nearly lossless until the spectrum is truncated, which could be tested as a compression principle."],"forward_implications":["Observed variation in biological parameters—protein copy numbers, synaptic strengths, receptor positions—is not evidence against optimization; it is what optimization on a sloppy landscape predicts.","In high-dimensional parameter spaces with sloppy spectra, populations should be spread across a large volume of parameter space while maintaining near-optimal average performance, making tightly controlled parameters the exception that demands explanation.","The maximum-entropy relation gives a quantitative tradeoff: at fixed average performance, the entropy per parameter that a population can sustain grows with the dynamic range of Hessian eigenvalues, so soft modes act as a resource for variability.","The results unify optimization arguments across gene regulation, neural circuits, and deep networks, suggesting a common principle rather than a collection of special cases."],"supporting_citations":[{"why":"Introduces the concept of sloppy models, the starting point for the paper's treatment of soft modes in performance landscapes.","marker":"[13]"},{"why":"Argues that smooth models generically produce sloppy spectra with eigenvalues spread over many decades, supporting the universality of the observed spectra.","marker":"[16]"},{"why":"Shows that the crab stomatogastric rhythm can arise from widely disparate ion-channel copy numbers, a key empirical example of variability at near-optimal function.","marker":"[9]"},{"why":"Supplies the scale-invariant natural image statistics used to compute information in the receptor array example.","marker":"[26]"},{"why":"Establishes that four gap genes encode position with about one percent accuracy, setting up the expression-level optimization problem.","marker":"[28]"},{"why":"Solves the readout optimization for gap-gene expression, giving the performance function and threshold parameterization used in the embryo example.","marker":"[35]"},{"why":"Provides the maximum-entropy formalism used to derive the predicted parameter distribution.","marker":"[58]"},{"why":"Documents soft modes and scale-invariant Hessian spectra in over-parameterized neural networks, supporting the deep network example.","marker":"[52]"},{"why":"Shows a broad space of near-optimal strategies in discrete decision tasks, an analogous example of soft modes in strategy space.","marker":"[72]"}],"fun_headline_variants":["Optimal without fine-tuning: sloppy modes make variability free","Extensive entropy near optimum: why variability comes free","Sloppy spectra: the key to being both optimal and variable","Near-optimal but variable: a maximum-entropy resolution"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The prediction rests on the assumption that a population is as variable as possible subject only to the constraint on average performance; if mutation biases, history, or out-of-equilibrium dynamics actually shape parameter distributions, the predicted link between average performance and parameter entropy need not hold.","fun_headline_variants_meta":{"raw":{"variants":["Optimal without fine-tuning: sloppy modes make variability free","Extensive entropy near optimum: why variability comes free","Sloppy spectra: the key to being both optimal and variable","Near-optimal but variable: a maximum-entropy resolution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0002,"raw_usage":{"total_tokens":1331,"prompt_tokens":856,"completion_tokens":475,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":406}},"tokens_in":472,"tokens_out":475,"duration_ms":5901,"temperature":1.0,"reasoning_tokens":406,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:46:32.732664+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure, in a real system such as the stomatogastric ganglion or a laboratory selection experiment, both the Hessian eigenvalue spectrum and the joint distribution of parameters across individuals; if the observed parameter entropy falls far below the maximum-entropy value predicted from the measured spectrum and the observed average performance via Eq. (31), or if a high-dimensional system with near-optimal performance shows a narrow spectrum of eigenvalues with no soft tail, then the claimed coexistence of optimization and variability fails for that system.","supporting_citations":[{"cited_title":"The roles of mutation, inbreeding, crossbreed- ing and selection in evolution","cited_arxiv_id":null,"evidence_quote":"Introduces the concept of sloppy models, the starting point for the paper's treatment of soft modes in performance landscapes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues that smooth models generically produce sloppy spectra with eigenvalues spread over many decades, supporting the universality of the observed spectra."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that the crab stomatogastric rhythm can arise from widely disparate ion-channel copy numbers, a key empirical example of variability at near-optimal function."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the scale-invariant natural image statistics used to compute information in the receptor array example."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that four gap genes encode position with about one percent accuracy, setting up the expression-level optimization problem."},{"cited_title":"D., Tkaˇ cik, G., Bialek, W., Wieschaus, E","cited_arxiv_id":null,"evidence_quote":"Solves the readout optimization for gap-gene expression, giving the performance function and threshold parameterization used in the embryo example."},{"cited_title":"& Hinton, G","cited_arxiv_id":null,"evidence_quote":"Documents soft modes and scale-invariant Hessian spectra in over-parameterized neural networks, supporting the deep network example."},{"cited_title":"& Hartl, D","cited_arxiv_id":null,"evidence_quote":"Shows a broad space of near-optimal strategies in discrete decision tasks, an analogous example of soft modes in strategy space."}],"review_version":1}