REVIEW 4 major objections 5 minor 40 references
Mutual Information Rate -- Linear Noise Approximation and Exact Computation
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read For discrete biochemical networks the Gaussian approximation never reaches the exact information rate, and in nonlinear continuous systems its bias direction depends on the linearization route.
desk verdict Useful, clearly written, and honest, but the missing convergence diagnostics for the PWS ground truth make the novel quantitative claims—especially the low-copy-number optimum—provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the mutual information rate R(S,X), the long-time limit of trajectory mutual information. Path Weight Sampling (PWS) estimates it exactly by Monte Carlo: sample input trajectories s_T from P(s_T), output trajectories x_T from P(x_T|s_T), and average log P(x_T|s_T) − log P(x_T); for Langevin dynamics the path weight is the Onsager–Machlup action. The Gaussian approximation replaces trajectory statistics by a jointly Gaussian model, reducing the rate to a frequency integral over the coherence φ_sx(ω)=|S_sx(ω)|²/(S_ss(ω)S_xx(ω)). The two variants differ in how the spectra are obtained: analytically from the linear noise approximation (LNA), or numerically from simulat
What would settle it
Recompute the four-reaction network's information rate with an independent exact method—e.g., numerical integration of the stochastic filtering equation—at large copy numbers (s̄,x̄ ≥ 100) and compare with the LNA closed form; agreement would refute the claim that the Gaussian bound never tightens. For the nonlinear Langevin system, scan parameters with fast response and high gain and check whether the empirical Gaussian estimate ever exceeds the PWS rate; the paper predicts it never does.
Extended reading notes
Core claim
The paper's central claim is that the Gaussian approximation of the mutual information rate—exact only for linear systems with additive Gaussian noise—systematically misestimates the true rate in non-Gaussian systems. In a four-reaction linear network, the LNA-based Gaussian rate R=(λ/2)(√(1+ρ/λ)−1) is a strict lower bound that does not tighten as copy numbers grow; a reaction-based discrete approximation R=(λ/2)(√(1+2ρ/λ)−1) matches exact PWS results over nearly all parameters, except low input copy numbers where the true rate peaks above both. In a continuous Langevin system with a Hill-type nonlinearity, linearizing dynamics via the LNA overestimates the exact rate, while estimating spect
Load-bearing premise
The conclusions treat the Path Weight Sampling results as ground truth, but PWS is exact only with infinitely many Monte Carlo samples and, for the Langevin extension, an infinitesimal time step; the paper does not report convergence checks, so any systematic bias in the PWS estimates would change the measured deviations of the Gaussian approximation.
Editorial extensions
If this is right
- For linear discrete reaction networks, the Gaussian approximation is a strict lower bound on the true information rate and never becomes tight at high copy numbers; estimates from it should be interpreted as a floor, not a prediction.
- The discrete approximation of [4] is accurate over a wide parameter range (output copy numbers down to ≪1 and input copy numbers ≳10), making it a cheap replacement for exact computation in those regimes.
- For continuous nonlinear systems, the empirical Gaussian estimate is a lower bound for Gaussian inputs, while the LNA-based estimate is not bounded; across the tested parameters the empirical route is closer to the exact rate.
- Practical usage rule: Gaussian approximations are reliable for small gain or slow output response; exact methods like PWS are needed for low input copy numbers (≲10) in discrete networks and for fast, strongly nonlinear continuous systems.
- The true information rate of the discrete network has an optimum at low input copy number (s̄ ≈ 1), above both analytic approximations, so reducing input copy number can increase information transmission in this regime.
Reading between the lines
- The distinction between reaction-based and state-based readouts implies the Gaussian bound is the information accessible to a decoder that cannot resolve individual reaction events; a testable prediction is that a downstream decoder with finer time resolution approaches the higher PWS rate.
- The opposite bias directions likely extend beyond Hill kinetics to any saturating input–output relation, so repeating the comparison with Michaelis–Menten or exponential saturation would test whether the asymmetry is generic.
- The low-input-copy-number peak suggests intrinsic noise can carry information rather than only degrade it; if real signaling networks tune input copy numbers near this optimum, that would be evidence of noise exploitation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript evaluates the Gaussian approximation (GA) to the mutual information rate in two model problems, using Path Weight Sampling (PWS) as an exact benchmark. For a linear discrete reaction network, the LNA-based GA is shown to be a lower bound that is not tight even at large copy numbers, whereas a recently proposed 'discrete approximation' agrees with PWS except at low input copy numbers, where PWS indicates a nonmonotonic maximum. For a nonlinear Langevin system with Hill-type activation, the LNA-based GA overestimates and an empirical Gaussian estimate underestimates the exact rate, with errors increasing with gain and decreasing with response time. The paper derives closed-form expressions, discusses the mechanism behind the discreteness effect, and gives practical guidance on when approximate methods suffice.
Significance. If the numerical results are reliable, the paper provides a useful quantitative caution against routine use of the Gaussian approximation and gives concrete guidance. Its analytic derivations are self-consistent, and benchmarking against an independent exact Monte Carlo method is appropriate. The extension of PWS to Langevin dynamics is a useful methodological contribution. However, the central quantitative claims—especially the low-copy-number optimum and the signed deviations in the nonlinear system—rest entirely on PWS estimates for which no convergence evidence or uncertainty is reported. As it stands, the paper's empirical conclusions are plausible but not fully substantiated.
major comments (4)
- [§II.C, Eqs. (9), (11); §III.A–B] The PWS benchmark is defined by three limits: N→∞ in Eq. (9), T→∞ in Eq. (4), and Δt→0 in Eq. (11). None of these limits is documented for any figure: no sample sizes, trajectory lengths, time steps, convergence checks, or error bars are reported. This is load-bearing because the central quantitative claims—the non-tight lower bound in Fig. 1a, the nonmonotonic optimum in Fig. 1b, and the signed deviations in Figs. 3–4—are deviations from PWS. Without evidence that the PWS estimates have converged, the reported deviations could be numerical artifacts. Please add convergence diagnostics and error bars/confidence intervals for all PWS points.
- [§III.A, Fig. 1b] The low-copy-number maximum of the information rate is a central new observation, yet the paper itself states that 'a precise characterization of this finding' is left for future work. Because both analytic approximations are independent of κ, the entire phenomenon rests on the PWS points for small κ. The figure shows no error bars and no dependence on trajectory length or sample size. The authors should show that the nonmonotonicity is stable under increasing T and N, and provide uncertainty estimates, before concluding that low input copy numbers maximize the rate.
- [§III.B, Eq. (24), Fig. 4] The empirical Gaussian rate is computed by Welch's method from simulated trajectories. Finite-sample spectral and coherence estimates are biased; the manuscript does not report window length, overlap, number of independent realizations, or statistical uncertainty. The claimed lower-bound behavior relies on the asymptotic Mitra–Stark argument, so the numerical verification in Fig. 4 must be shown not to be an artifact of spectral estimation bias. Please provide the estimation parameters and error bars, and ideally a check that the empirical Gaussian estimate converges to the true spectral coherence from below as data length increases.
- [§II.C, Eq. (11)] For diffusive systems, the path weight Eq. (11) is valid only in the limit Δt→0 and, for multiplicative noise, up to a normalization constant that depends on x_i and therefore does not cancel in Eq. (9). The paper does not report Δt or a convergence study for the Euler–Maruyama discretization. While the case studies appear to use additive noise (constant σ), the general claim that PWS has been extended to diffusive dynamics needs a statement of the discretization error and a convergence check for the specific systems studied.
minor comments (5)
- [§II.C] Typo: 'Euler-Mayurama' should be 'Euler–Maruyama'.
- [§II.C, Eq. (9)] The displayed estimator appears to be missing an explicit 1/N prefactor; as written it reads like a sum rather than an average.
- [Figs. 3–4] The colorbar labels are hard to parse ('10 1 100 101'); they should be formatted as 10^{-1}, 10^0, 10^1, and the units of τx should be stated.
- [General] No code or data availability statement is provided. For reproducibility, the PWS implementation, spectral estimation parameters, and simulation details should be included in an appendix or public repository.
- [Abstract/Introduction] The phrase 'even when the system's statistics are nearly Gaussian' is not quantified. A brief measure of non-Gaussianity (e.g., skewness of the stationary distribution or a path-space diagnostic) would strengthen the claim.
Circularity Check
No significant circularity: approximations are independent of the PWS benchmark and not fitted to it.
full rationale
The paper's central comparisons use Path Weight Sampling (PWS) from Ref. [5] as an exact Monte Carlo benchmark. PWS is published independently, is not fitted to any of the approximate formulas, and does not incorporate the target results of this paper. The Gaussian approximation formulas (Eqs. 18, 23) and the discrete approximation (Eq. 19) are derived analytically from the underlying stochastic models (LNA and filtering approximations, respectively) before comparison with PWS. No parameter is calibrated to PWS data, and the reported deviations are observed rather than enforced by construction. The claimed lower-bound property of the empirical Gaussian estimate is supported by an independent theorem (Mitra and Stark), not by a self-citation. Self-citations to [4] and [24] provide context and mechanistic explanation, but the load-bearing quantitative evidence comes from PWS, an external exact method. The absence of explicit PWS convergence statistics (sample size, trajectory length, time step) is a reproducibility/correctness concern, not circularity. Therefore the derivation chain is self-contained with respect to the evaluated approximations.
Assumptions & free parameters
assumptions (4)
- domain assumption Path Weight Sampling (PWS) computes the exact mutual information rate for the considered systems
- standard math The Onsager-Machlup discretized path weight (Eq. 11) accurately represents the conditional trajectory likelihood for the Langevin systems
- domain assumption The processes are stationary and ergodic, so the mutual information rate in Eq. (4) exists and can be estimated from finite trajectory Monte Carlo
- domain assumption The chemical master equation is the exact model for the discrete reaction system
Cite this review
Pith. "Pith review of Mutual Information Rate -- Linear Noise Approximation and Exact Computation." pith.science (2026). https://pith.science/paper/Z7SVVCOF
@misc{pith2026250821220,
author = {Pith},
title = {Pith review of: Mutual Information Rate -- Linear Noise Approximation and Exact Computation},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z7SVVCOF}},
note = {Machine review of arXiv:2508.21220}
}
read the original abstract
Efficient information processing is crucial for both living organisms and engineered systems. The mutual information rate, a core concept of information theory, quantifies the amount of information shared between the trajectories of input and output signals, and enables the quantification of information flow in dynamic systems. A common approach for estimating the mutual information rate is the Gaussian approximation which assumes that the input and output trajectories follow Gaussian statistics. However, this method is limited to linear systems, and its accuracy in nonlinear or discrete systems remains unclear. In this work, we assess the accuracy of the Gaussian approximation for non-Gaussian systems by leveraging Path Weight Sampling (PWS), a recent technique for exactly computing the mutual information rate. In two case studies, we examine the limitations of the Gaussian approximation. First, we focus on discrete linear systems and demonstrate that, even when the system's statistics are nearly Gaussian, the Gaussian approximation fails to accurately estimate the mutual information rate. Second, we explore a continuous diffusive system with a nonlinear transfer function, revealing significant deviations between the Gaussian approximation and the exact mutual information rate as nonlinearity increases. Our results provide a quantitative evaluation of the Gaussian approximation's performance across different stochastic models and highlight when more computationally intensive methods, such as PWS, are necessary.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Signal The input signal is generated by a birth-death process, ∅ κ − −⇀↽− −λ S, (A1) Its dynamics in Langevin form are, ˙s = κ − λs(t) + ηs(t), (A2) yielding the steady state signal concentration ¯ s = κ/λ. The independent Gaussian white noise process ηs(t) sum- marizes all reactions that contribute to fluctuations in S. The strength of the noise term in ...
-
[2]
(A5) We define the activation level a(s) to be a Hill function, a(s) = s(t)n K n + s(t)n
Linear approximation We now consider the readout X, which is produced via a nonlinear activation function a(s): S ρ a (s) − − − →S + X, X µ − − → ∅. (A5) We define the activation level a(s) to be a Hill function, a(s) = s(t)n K n + s(t)n . (A6) Such a dependency, in which K sets the concentration of S at which the activation is half-maximal and n sets the...
-
[3]
Information rate Following Tostevin & Ten Wolde [1, 2], we can express the Gaussian information rate as follows, R(S; X ) = 1 4π Z ∞ −∞ dω log 1 + |K(ω)|2 |N(ω)|2 Sss(ω) , (A13) where |K(ω)|2 is the frequency dependent gain and |N(ω)|2 is the frequency dependent noise of the output process X . If the intrinsic noise of the network is not cor- related to t...
-
[4]
Linear network To disambiguate differences in the information rate caused by the linear approximation of our nonlinear reac- tion network on the one hand and the Gaussian approx- imation of the underlying jump process on the other, we consider the information rate of a linear network. Any difference between the exact information rate and the Gaussian info...
-
[5]
Mutual information between in- and output trajectories of biochemical networks
F. Tostevin and P. R. ten Wolde, Mutual Information between Input and Output Trajectories of Biochemical Networks, Physical Review Letters 102, 218101 (2009), 0901.0280
work page Pith review arXiv 2009
-
[6]
F. Tostevin and P. R. ten Wolde, Mutual information in time-varying biochemical systems, Phys. Rev. E 81, 061917 (2010)
work page 2010
-
[7]
H. H. Mattingly, K. Kamino, B. B. Machta, and T. Emonet, Escherichia coli chemotaxis is information limited, Nature Physics 17, 1426 (2021)
work page 2021
-
[8]
A.-L. Moor and C. Zechner, Dynamic information trans- fer in stochastic biochemical networks, Physical Review Research 5, 013032 (2023)
work page 2023
Show all 40 references
-
[9]
Reinhardt, G
M. Reinhardt, G. Tkaˇ cik, and P. R. ten Wolde, Path Weight Sampling: Exact Monte Carlo Computation of the Mutual Information between Stochastic Trajectories, Physical Review X 13, 041017 (2023)
2023
-
[10]
L. Hahn, A. M. Walczak, and T. Mora, Dynamical In- formation Synergy in Biochemical Signaling Networks, Physical Review Letters 131, 128401 (2023)
2023
-
[11]
J. E. Segall, S. M. Block, and H. C. Berg, Temporal com- parisons in bacterial chemotaxis., Proceedings of the Na- tional Academy of Sciences 83, 8987 (1986)
1986
-
[12]
M. W. Covert, T. H. Leung, J. E. Gaston, and D. Bal- timore, Achieving stability of lipopolysaccharide-induced nf-κb activation, Science 309, 1854 (2005)
2005
-
[13]
Strong, R
S. Strong, R. D. R. Van Steveninck, W. Bialek, R. Koberle, et al., On the application of information the- ory to neural spike trains, in Pac Symp Biocomput, Vol. 1998 (1998) pp. 621–632
1998
-
[14]
C. E. Shannon, A Mathematical Theory of Communica- tion, Bell System Technical Journal 27, 379 (1948)
1948
-
[15]
E. Ziv, I. Nemenman, and C. H. Wiggins, Optimal signal processing in small stochastic biochemical networks, PloS one 2, e1077 (2007)
2007
-
[16]
Cheong, A
R. Cheong, A. Rhee, C. J. Wang, I. Nemenman, and A. Levchenko, Information transduction capacity of noisy biochemical signaling networks, science 334, 354 (2011)
2011
-
[17]
J. O. Dubuis, G. Tkaˇ cik, E. F. Wieschaus, T. Gregor, and W. Bialek, Positional information, in bits, Proceedings of the National Academy of Sciences 110, 16301 (2013)
2013
-
[18]
S. E. Palmer, O. Marre, M. J. Berry, and W. Bialek, Pre- dictive information in a sensory population, Proceedings of the National Academy of Sciences 112, 6908 (2015)
2015
-
[19]
Chalk, O
M. Chalk, O. Marre, and G. Tkaˇ cik, Toward a unified theory of efficient, predictive, and sparse coding, Pro- ceedings of the National Academy of Sciences 115, 186 (2018)
2018
-
[20]
Bauer, M
M. Bauer, M. D. Petkova, T. Gregor, E. F. Wieschaus, and W. Bialek, Trading bits in the readout from a ge- netic network, Proceedings of the National Academy of Sciences 118, e2109011118 (2021)
2021
-
[21]
Sachdeva, T
V. Sachdeva, T. Mora, A. M. Walczak, and S. E. Palmer, Optimal prediction with resource constraints using the information bottleneck, PLoS computational biology 17, e1008743 (2021)
2021
-
[22]
A. J. Tjalma, V. Galstyan, J. Goedhart, L. Slim, N. B. Becker, and P. R. ten Wolde, Trade-offs between cost and information in cellular prediction, Proceedings of the National Academy of Sciences 120, e2303078120 (2023)
2023
-
[23]
A. J. Tjalma and P. R. ten Wolde, Predicting concen- tration changes via discrete receptor sampling, Physical Review Research 6, 033049 (2024)
2024
-
[24]
T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed. (John Wiley & Sons, 2006)
2006
-
[25]
Tkaˇ cik, C
G. Tkaˇ cik, C. G. Callan, Jr., and W. Bialek, Information flow and optimization in transcriptional regulation, Pro- ceedings of the National Academy of Sciences 105, 12265 (2008), 0705.0313
2008 arXiv
-
[26]
W. H. de Ronde, F. Tostevin, and P. R. ten Wolde, Effect of feedback on the fidelity of information transmission of time-varying signals, Physical Review E—Statistical, Nonlinear, and Soft Matter Physics 82, 031914 (2010)
2010
-
[27]
P. P. Mitra and J. B. Stark, Nonlinear limits to the infor- mation capacity of optical fibre communications, Nature 411, 1027 (2001)
2001
- [28]
-
[29]
Meijers, S
M. Meijers, S. Ito, and P. R. ten Wolde, Behavior of information flow near criticality, Physical Review E 103, L010102 (2021), 1906.00787
2021 arXiv
-
[30]
Fan and A
R. Fan and A. Hilfinger, Characterizing the nonmono- tonic behavior of mutual information along biochemical reaction cascades, Physical Review E110, 034309 (2024), 2309.10162
2024 arXiv
-
[31]
Van Kampen, Stochastic processes in physics and chemistry (North Holland, Amsterdam, 1992)
N. Van Kampen, Stochastic processes in physics and chemistry (North Holland, Amsterdam, 1992)
1992
-
[32]
Onsager and S
L. Onsager and S. Machlup, Fluctuations and Irreversible Processes, Physical Review 91, 1505 (1953)
1953
-
[33]
E. A. Nadaraya, On Estimating Regression, Theory of Probability & Its Applications 9, 141 (1964)
1964
-
[34]
G. S. Watson, Smooth Regression Analysis, Sankhy¯ a: The Indian Journal of Statistics, Series A (1961-2002) 26, 359 (1964)
1961
-
[35]
Malaguti and P
G. Malaguti and P. R. ten Wolde, Theory for the opti- mal detection of time-varying signals in cellular sensing systems, eLife 10, 1 (2021)
2021
-
[36]
A. V. Oppenheim and R. W. Schafer, Digital Signal Pro- cessing (Prentice-Hall, 1975)
1975
-
[37]
P. B. Warren, S. T˘ anase-Nicola, and P. R. ten Wolde, Ex- act results for noise power spectra in linear biochemical reaction networks, The Journal of Chemical Physics125, 144904 (2006), https://pubs.aip.org/aip/jcp/article- pdf/doi/10.1063/1.2356472/15393314/144904 1 online.pdf
2006 doi
-
[38]
S. A. Cepeda-Humerez, J. Ruess, and G. Tkaˇ cik, Esti- mating information in time-varying signals., PLoS com- putational biology 15, e1007290 (2019)
2019
-
[39]
The bound specifically re- quires the Gaussian model to use the covariance of the full, original system
Note that this argument does not apply to the Linear Noise Approximation (LNA). The bound specifically re- quires the Gaussian model to use the covariance of the full, original system. When the system is first linearized using the LNA, the resulting linear model does not retai...
-
[40]
T˘ anase-Nicola, P
S. T˘ anase-Nicola, P. B. Warren, and P. R. ten Wolde, Sig- nal detection, modularity, and the correlation between extrinsic and intrinsic noise in biochemical networks, Physical review letters 97, 068102 (2006)
2006
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.