REVIEW 4 major objections 6 minor 17 references
Decoding cell signaling via optimal transport and information theory
T0 review · 4 major / 6 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read The paper claims that signaling fidelity is two-dimensional—mutual information plus the inverse 2-Wasserstein distance—and that feedback circuits trade the first for the second.
desk verdict The dual-fidelity framework is a real idea worth engaging, but the TNF experiment doesn't establish the feedback claim: the unit rescaling looks like it manufactures the WT/A20 separation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the 2-Wasserstein distance (2-WD), the minimal 'transport cost' needed to reshape the input distribution into the output distribution; its inverse becomes geometric fidelity, while mutual information becomes informational fidelity. The two are combined in a Lagrangian, L = MI − λ·(2-WD)², optimized over the input noise level, with λ setting whether the cell prioritizes state resolution or distributional correspondence. Analytical Gaussian formulas express both metrics through means, squared coefficients of variation, and input-output covariance, and binding-affinity parameters θX and θZ act as tunable biochemical levers that move a circuit across the dual-fidelity
What would settle it
Compute the dual-fidelity coordinates using an independent calibration (e.g., a reporter whose output is measured in absolute molecule numbers, or single-molecule RNA counts) and see whether wild-type cells still show higher geometric fidelity than A20-knockout cells. Alternatively, recompute the 2-Wasserstein distance after rescaling output quantiles by the dose-response slope at several different TNF doses; if the ordering flips depending on the chosen dose, the gain-based mapping is driving the result.
Extended reading notes
Core claim
The central claim is that distributional correspondence—how faithfully the output distribution mirrors the input distribution—is a distinct, measurable dimension of signaling fidelity, alongside the usual informational dimension measured by mutual information. Using Gaussian-channel closed forms for MI and the 2-Wasserstein distance, the paper shows that six canonical regulatory motifs occupy different regions of the resulting dual-fidelity space as binding affinities and a trade-off parameter λ are varied. The motif-specific result is that feedback architectures (especially negative feedback) suppress noise and dynamic range, improving geometric fidelity while lowering informational fidelit
Load-bearing premise
The experimental part assumes that arbitrary fluorescence units can be converted to TNF-concentration units using the reciprocal slope of a fitted dose-response curve, and that wild-type and A20-knockout cells can be compared under different mapping rules (sigmoidal versus linear); if that rescaling is invalid, the observed separation between feedback and no-feedback cells could be an artifact.
Editorial extensions
If this is right
- If geometric fidelity is real, MI-only comparisons mis-rank circuits: a mutant with higher information transmission (A20-knockout) can be worse at faithful signaling, not better.
- Topology predetermines strategy: coherent type-1 feed-forward loops can reach the 'precise' regime (both fidelities high), while feedback motifs—especially negative feedback—are biased toward geometric fidelity and low information.
- Binding affinities provide a practical control knob: weak binding generally boosts informational fidelity, strong binding boosts geometric fidelity, with characteristic inversions in incoherent and feedback motifs.
- The framework yields a synthetic-design rule: tune promoter binding affinities (and, in principle, production/degradation rates) to place a circuit in the fidelity regime a task demands.
- The trade-off supplies a resolution to a paradox: wild-type cells can look information-suboptimal relative to a knockout, yet be better optimized once distributional correspondence is included.
Reading between the lines
- The paper leaves implicit that this two-dimensional view could reinterpret other evolution experiments where knockouts exceed wild-type in MI; those cases may be feedback-loss artifacts rather than improvements.
- A direct extension is to test the predicted binding-affinity lever in synthetic circuits: vary θX and θZ experimentally and see whether the dual-fidelity movement follows the computed direction for each motif, especially I1-FFL's inversion.
- With single-cell reporters calibrated to absolute molecule counts, the unit-rescaling step could be eliminated; such data would test whether the WT versus A20-knockout separation is a real biological trade-off or a product of the slope-based mapping.
- The Lagrangian resembles a rate-distortion problem with a geometric distortion; if general bounds on the 2-Wasserstein distance exist, it might be possible to derive universal inequalities between information rate and distributional distortion in signaling networks, which the paper only hints at.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a dual-fidelity framework for cell signaling, defining informational fidelity as mutual information I(X;Z) and geometric fidelity as the inverse 2-Wasserstein distance W(X;Z)^{-1}. A Lagrangian L = I - λW^2 is optimized over input noise for six gene-regulatory motifs under the LNA/Gaussian approximation, producing trade-off curves in the (I, W^{-1}) plane. The authors then apply the framework to TNF signaling data from Cheong et al. (2011), reporting that WT cells (with A20-mediated negative feedback) lie at higher geometric fidelity and lower informational fidelity than A20-deficient cells, and argue that this supports the theoretical prediction that negative feedback sacrifices information for distributional correspondence.
Significance. If the results were fully supported, the paper would offer a biologically meaningful second dimension of signaling fidelity and a practical computational framework. The analytic Gaussian formulas (Eqs. 8–10) are correct, and the LNA derivations in the SI are standard and clearly laid out. The conceptual distinction between state discrimination and distributional correspondence is valuable and likely of broad interest. The paper also includes explicit discussion of the limitations of the Gaussian approximation. However, the current experimental validation and the built-in character of the trade-off leave the central claims inadequately supported.
major comments (4)
- [SI S5, Eq. (S26), Figs. S3a–d] The unit rescaling used to compute the 2-WD is not unit-invariant and likely generates the WT/A20-/- separation in Fig. 6c. WT output is divided by Hill slopes at half-maximum of 215.0 and 156.9 a.u. per ng/mL, while A20-/- output is divided by linear slopes of 29.2 and 28.1 a.u. per ng/mL. This 5–7-fold difference in mapping factors compresses WT output quantiles toward the input scale and inflates A20-/- quantiles, mechanically increasing W for A20-/- and decreasing W for WT. The two genotypes are thus compared under different mapping rules, so Fig. 6c does not provide evidence for a feedback-dependent geometric-fidelity trade-off. A unit-invariant comparison (e.g., rank-transform both marginals to a common scale, or use a single calibration curve for both genotypes) is needed.
- [Abstract / Full text] The abstract states that 'RAS-MAPK data analysis further shows that jointly considering INF and GMF better characterizes intracellular signal relay than INF alone,' but no RAS-MAPK analysis appears anywhere in the main text, Methods, or SI. This is a claimed result with missing support. Either provide the missing analysis or revise the abstract so that it reflects the content of the paper.
- [Eq. (6) and Fig. 5] The trade-off between I and W^{-1} is largely built in by construction. Optimizing L = I - λW^2 over η_X^2 with increasing λ necessarily shifts the optimum toward smaller W (larger GMF) at the expense of I, for any pair of continuous functionals. Thus the qualitative statement that 'feedback architectures typically sacrifice informational fidelity to enhance geometric fidelity' is not an independent prediction of the model. The genuine content is in the motif-specific shape and location of the curves (e.g., which regimes are accessible at fixed λ, or the value of I at λ=0). The manuscript should be reframed accordingly, and the claims should be restricted to comparisons at fixed λ or to the predicted accessible regimes.
- [SI S5, 'MI-constrained reconstruction'] The experimental MI is not independently measured: the input marginal P_X is optimized so that the computed I matches the values reported by Cheong et al. Therefore the I-coordinate in Fig. 6c is constrained rather than estimated. The only freely estimated quantity is W, which is exactly the coordinate affected by the rescaling in the first major comment. Thus Fig. 6c cannot provide independent confirmation of the dual-fidelity trade-off. The analysis should either estimate both coordinates from a pre-specified P_X (e.g., the empirical TNF dose distribution) or demonstrate that the conclusions are robust to the MI-matching procedure.
minor comments (6)
- [Introduction, after Eq. (1)] The statement that 'the general form of the Lagrangian L can be considered the Sinkhorn distance' is inaccurate: the Sinkhorn distance refers to entropy-regularized optimal transport, not to an MI-minus-W^2 Lagrangian. Please correct or remove.
- [Fig. 6e] The theoretical and experimental curves are plotted on the same axes despite having different units (reciprocal copy number vs reciprocal ng/mL). The comparison is only qualitative; the figure should state this explicitly and ideally use normalized coordinates.
- [SI S5] The output variable is described as being in 'concentration units,' but fluorescence is in arbitrary units (a.u.). The text should distinguish these consistently, as the unit conversion is a central issue.
- [Results, 'Motif-specific patterns'] The regime thresholds (I ≥ 0.3 bits, W^{-1} ≥ 1) are arbitrary and unit-dependent. The paper should emphasize that they are illustrative and do not carry intrinsic biological meaning in experimental units.
- [Materials and methods / SI S3] The model uses Hill coefficient h=1, equal mean copy numbers, and fixed degradation timescales. A sensitivity analysis over these choices would strengthen the generality claims in the Discussion.
- [Fig. 6c / SI S5] The bootstrap error bars appear to include resampling of the reconstructed distributions, but not uncertainty in the slope-based mapping factors. This limitation should be stated.
Circularity Check
The experimental WT/A20−/− geometric-fidelity separation in Fig. 6c is generated by condition-specific a.u.-to-ng/mL rescaling (different fitted slopes for WT vs A20−/−), so the main experimental support for the paper's central claim is partly constructed rather than independently measured.
-
fitted input called prediction
[SI S5, Eq. (S26) and Figs. S3a–d; Results 'Experimental signature of dual-fidelity trade-offs' / Fig. 6c]
"The reciprocal of this slope, with units of ng/mL per a.u., is therefore used as the mapping factor to convert output quantiles into the same units as the input. ... Slope at half-max = 215.016 a.u./(ng/mL) ... Slope = 29.208 a.u./(ng/mL)"
Eq. (S26) defines W as the RMS difference between input and output quantile functions in ng/mL. Output quantiles are originally in a.u., so each condition's output quantiles are divided by a fitted condition-specific slope. WT is divided by 215.0/156.9 a.u. per ng/mL (Hill half-max slopes), while A20−/− is divided by 29.2/28.1 a.u. per ng/mL (linear slopes). The ~5–7× larger WT divisor compresses WT output quantiles toward the input scale, mechanically reducing W (raising geometric fidelity), while the small A20−/− divisor inflates W. Thus the WT/KO separation in Fig. 6c is not a unit-invariant property of the signaling distributions; it is imposed by applying different fitted mapping rules before computing the metric that is then presented as the new fidelity axis.
full rationale
The theoretical motif analysis is not substantially circular: the LNA supplies closed forms for output noise η_Z^2 and covariance ζ_XZ, and MI and 2-WD are evaluated from those expressions (Eqs. 8 and 10). Optimizing L = I − λ W² over input noise with a sweep of λ does parameterize a trade-off by construction, but the motif-specific ordering (e.g., C1-FFL high in both, NFL biased toward geometric fidelity) is a nontrivial output of the model, not written into the definition of L. Self-citations to Ito 2024 and Nandi et al. 2024 are methodological (OT bounds and LNA formulas) and are not load-bearing for the dual-fidelity dichotomy. The serious circularity is in the experimental validation. The Materials and methods state that the input marginal is optimized until MI matches Cheong et al.'s reported values, so the informational-fidelity coordinate is imported, not predicted. The genuinely new coordinate, geometric fidelity, is then computed from quantile functions after rescaling output a.u. quantiles by the reciprocal of a fitted, condition-specific dose-response slope. Because WT and A20−/− are rescaled with very different fitted slopes (Hill half-max for WT, linear for A20−/−), the resulting W values, and hence the Fig. 6c WT/KO separation, are largely determined by that arbitrary unit-conversion choice rather than by feedback biology. The paper's caveat that only trends and separations are compared does not neutralize this, since the separation itself is not invariant to the chosen rescaling. This makes the main experimental support for 'negative feedback increases geometric fidelity at the expense of informational fidelity' partially circular: the measured outcome is encoded in the condition-specific fitted mapping used to define it.
Assumptions & free parameters
free parameters (9)
- Lagrange multiplier lambda =
10^-4 to 1
- Binding affinity parameters theta_X, theta_Z =
{0.5, 1, 2}
- Input noise eta_X^2 =
optimized via Eq. (6)
- Mean copy numbers mu_X, mu_Y, mu_Z =
100
- Degradation rates tau_Y^-1, tau_Z^-1 =
1.0 and 10.0
- Hill coefficient h =
1
- Experimental input marginal P_X(x) =
optimized to match Cheong et al. MI
- Experimental unit mapping factors =
WT: 215.016, 156.863 a.u./(ng/mL); A20-/-: 29.208, 28.129 a.u./(ng/mL)
- Regime thresholds for informative/precise/geometric/poor =
0.3 bits and W^-1 = 1 (copy numbers)^-1
assumptions (8)
- domain assumption Linear noise approximation gives accurate steady-state noise for copy numbers ~100.
- domain assumption Input X is a static extrinsic Gaussian variable not regulated by Y or Z.
- domain assumption Regulation uses Hill functions with Hill coefficient h=1 and first-order degradation.
- standard math For one-dimensional Gaussian X and Z, W2^2 = (mu_X - mu_Z)^2 + (sigma_X - sigma_Z)^2.
- standard math For Gaussian joint distributions, MI is given by Eq. (8).
- domain assumption Reported mean and standard deviation from Cheong et al., with Gaussian conditionals, represent the true single-cell output distributions.
- ad hoc to paper Gain-based rescaling via fitted dose-response slope is a valid common scale for computing 2-WD between input (ng/mL) and output (a.u.).
- ad hoc to paper Cellular priorities are captured by maximizing L = I - lambda W^2 over input noise eta_X^2.
invented entities (1)
-
Geometric fidelity (GMF), the inverse 2-Wasserstein distance, as a second dimension of signaling fidelity
Cite this review
Pith. "Pith review of Decoding cell signaling via optimal transport and information theory." pith.science (2026). https://pith.science/paper/T7FIKQS6
@misc{pith2026260218028,
author = {Pith},
title = {Pith review of: Decoding cell signaling via optimal transport and information theory},
year = {2026},
howpublished = {\url{https://pith.science/paper/T7FIKQS6}},
note = {Machine review of arXiv:2602.18028}
}
read the original abstract
Cellular signal processing performs reliably despite molecular noise. Mutual information (MI) is widely used to quantify signaling fidelity, capturing how well outputs discriminate input states. However, it fails to capture whether the output preserves the statistical structure of the input, a property crucial in morphogen patterning and dose-dependent signaling. To address this gap, we introduce the 2-Wasserstein (2-WD) distance, which provides a geometric basis for comparing input and output distributions. We define MI as informational fidelity (INF) and the inverse of the 2-WD as geometric fidelity (GMF). Applying this dual-fidelity framework to canonical regulatory motifs under Gaussian channel approximation reveals topology-dependent trade-offs: coherent feed-forward loops can perform well in both dimensions, whereas feedback architectures reduce INF to enhance GMF. Experimental analysis of tumor necrosis factor signaling reveals dual-fidelity behavior qualitatively consistent with feedback regulation. RAS-MAPK data analysis further shows that jointly considering INF and GMF better characterizes intracellular signal relay than INF alone. Our results thus indicate that these signaling behaviors are not fully characterized by MI alone; instead, distributional correspondence provides a complementary dimension of signaling fidelity. Our study provides a practical framework for analyzing natural networks and guiding the design of task-specific synthetic circuits.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
van Kampen, N. G. Stochastic Processes in Physics and Chemistry, 3rd ed. (North-Holland, Amster- dam, 2007)
2007
-
[2]
Gardiner, C. W. Stochastic Methods: A Handbook for the Natural and Social Sc iences, 4th ed. (Springer, Berlin, 2009)
2009
-
[3]
& Ehrenberg, M
Elf, J. & Ehrenberg, M. Fast evaluation of fluctuations in bio chemical networks with the linear noise approximation. Genome Res. 13, 2475–2484 (2003)
2003
-
[4]
Summing up the noise in gene networks
Paulsson, J. Summing up the noise in gene networks. Nature 427, 415–418 (2004)
2004
-
[5]
Tănase-Nicola, S., Warren, P. B. & ten Wolde, P. R. Signal dete ction, modularity, and the correlation between extrinsic and intrinsic noise in biochemical netwo rks. Phys. Rev. Lett. 97, 068102 (2006)
2006
-
[6]
H., Tostevin, F
de Ronde, W. H., Tostevin, F. & ten Wolde, P. R. Effect of feedba ck on the fidelity of information transmission of time-varying signals. Phys. Rev. E 82, 031914 (2010)
2010
-
[7]
& Banik, S
Nandi, M., Chattopadhyay, S., Bandyopadhyay, S. & Banik, S. K. Channel assisted noise propagation in a two-step cascade. Chaos: An Interdiscip. J. Nonlinear Sci. 34, 083128 (2024)
2024
-
[8]
Wallace, E. W. J. A simplified derivation of the linear noise a pproximation. arXiv (2010). arXiv:1004.4280
arXiv 2010
Show all 17 references
-
[9]
Wallace, E. W. J., Gillespie, D. T., Sanft, K. R. & Petzold, L. R. Linear noise approximation is valid over limited times for any chemical system that is suffic iently large. IET Syst. Biol. 6, 102–115 (2012)
2012
-
[10]
& ten Wolde, P
Mugler, A., Tostevin, F. & ten Wolde, P. R. Spatial partition ing improves the reliability of biochemical signaling. Proc. Natl. Acad. Sci. U.S.A. 110, 5927–5932 (2013)
2013
-
[11]
Gillespie, D. T. The chemical langevin equation. J. Chem. Phys. 113, 297–306 (2000)
2000
-
[12]
& Chapman, S
Erban, R. & Chapman, S. J. Stochastic Differential Equations , 59–94 (Cambridge University Press, 2020)
2020
-
[13]
Swain, P. S. Efficient attenuation of stochasticity in gene ex pression through post-transcriptional control. J. Mol. Biol. 344, 965–976 (2004). 30/31
2004
-
[14]
An Introduction to Systems Biology: Design Principles of Bi ological Circuits (CRC Press, Boca Raton, FL, 2006)
Alon, U. An Introduction to Systems Biology: Design Principles of Bi ological Circuits (CRC Press, Boca Raton, FL, 2006)
2006
-
[15]
Mehta, P., Goyal, S., Long, T., Bassler, B. L. & Wingreen, N. S. I nformation processing and signal integration in bacterial quorum sensing. Mol. Syst. Biol. 5, 325 (2009)
2009
-
[16]
J., Nemenman, I
Cheong, R., Rhee, A., Wang, C. J., Nemenman, I. & Levchenko, A . Information transduction capacity of noisy biochemical signaling networks. Science 334, 354–358 (2011)
2011
-
[17]
& Cuturi, M
Peyré, G. & Cuturi, M. Computational optimal transport. Found. Trends Mach. Learn. 11, 355–607 (2019). 31/31
2019
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.