Pith. sign in

REVIEW 3 major objections 4 minor 13 references

Information-Theoretical Measures for Developmental Cell-Fate Proportioning Processes

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Counting statistics can be lifted to estimate reproducibility entropy, and the simulated embryo model scores far below the ideal benchmark.

desk verdict The counting-space mapping and W_REP are sound and worth keeping, but the uniform lift forces S_PAT = S_SCF, so PI = 0 by construction; the headline claim that the ITWT system is far from optimal is an artifact of the approximation. read the letter →

arxiv 2505.23659 v1 pith:YTJRLXA5 submitted 2025-05-29 q-bio.QM

classification q-bio.QM
keywords self-organizationpositionalinformationreproducibilityentropycell-fateproportioningcountingspacepatterningblastocysttheory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to make a recently proposed information-theoretic measure of developmental self-organization computable for real systems. The obstacle is the reproducibility entropy $S_{\mathrm{REP}}$, which requires probabilities over the space of all cell-fate patterns, a set of size $z^n$ that explodes with cell number and fate number. The authors map each fate-count vector to the set of patterns that produce it and, assuming all those patterns are equally likely, lift the small counting space to the large patterning space. Applying this lift to a simulated model of epiblast versus primitive endoderm proportioning gives a self-organization utility of roughly 0.001 to 0.007 bits, far below the ideal benchmark, and they propose a simpler counting-space entropy that decreases with system size. The payoff is a generic, tractable way to score self-organization in any cell-fate proportioning process.

What carries the argument

The load-bearing object is the pseudo-inverse transform $M$ (Definition 0.12), a map from a fate-count vector $\vec n$ to the set of all patterning vectors $\vec z$ with $T(\vec z)=\vec n$, equipped with the rule that the patterns in that set are equiprobable. Its cardinality is the multinomial coefficient $n!/(n_1!\cdots n_z!)$, which lets counting-space probabilities be spread uniformly over pattern space, turning the intractable sum over $z^n$ patterns into a sum over the $(n+1)(n+2)/2$ count vectors. Combined with perfect counting reproducibility, the same lift defines the ideal patterning space that serves as the benchmark, and the alternative $W_{\mathrm{REP}}$ is the Shannon entropy of the counting distribution itself, normalized by $n$.

What would settle it

Compute $\mathrm{PI}=S_{\mathrm{PAT}}-S_{\mathrm{SCF}}$ directly from the ITWT simulation patterns, using each cell's actual position rather than the uniform lift; if per-position fate probabilities differ across positions, $\mathrm{PI}$ will be positive and the low utility estimate will be shown to have missed real spatial structure.

Watch

Extended reading notes

Core claim

The central claim is that $S_{\mathrm{REP}}$ can be estimated from low-dimensional counting statistics through the pseudo-inverse transform $M$ (Definition 0.12), which assigns equal probability to every patterning vector $\vec z$ whose count vector satisfies $T(\vec z)=\vec n$; the size of each such set is the multinomial coefficient $n!/(n_1!\cdots n_z!)$. Under this uniform lift, an ideal proportioning process has zero positional information ($\mathrm{PI}=0$ always), so its utility $U$ equals its correlational information and decreases monotonically with the number of cells $n$. For the simulated ITWT system the three entropies $S_{\mathrm{PAT}}$, $S_{\mathrm{SCF}}$, and $S_{\mathrm{REP}}$ stay nearly equal across all embryo sizes, making the empirical utility very small and leading the authors to conclude that the current view of optimal behavior for cell-fate proportioning is incomplete. The paper also introduces a count-space entropy $W_{\mathrm{REP}}$ that improves (decreases) with $n$ and agrees with earlier findings that the ITWT system proportioning is accurate and precise.

Load-bearing premise

All cell-fate patterns with the same fate counts are treated as equally probable, so the lifting of counting statistics to pattern probabilities assumes away any dependence of fate on cell position.

Editorial extensions

If this is right

  • For a pure proportioning process, $\mathrm{PI}=0$ identically, so self-organization utility reduces to correlational information and falls monotonically as the number of cells grows.
  • The ITWT simulation's empirical utility of about 0.001 to 0.007 bits is only a fifth of the ideal value, so under this measure the simulated embryo is far from optimally self-organized.
  • The counting-space entropy $W_{\mathrm{REP}}$ decreases with system size and matches the reproducibility reported in earlier simulation studies, giving a cheap diagnostic when patterning data are unavailable.
  • The same counting-to-patterning lift applies to any proportioning process, experimental or simulated, because it needs only observed frequencies of fate-count vectors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the uniform lift forces the single-cell fate distribution to be position-independent, so the reported near-zero positional information is built into the assumption rather than measured from the simulation's spatial structure.
  • Editorial inference: if the ITWT simulation contains real spatial correlations in fate decisions, computing $\mathrm{PI}$ directly from single-cell marginals would yield a positive value, and the true utility could be substantially larger than 0.001 to 0.007 bits.
  • Editorial inference: the same approach could be applied to experimental embryos by counting fates per embryo, giving cross-species or cross-condition comparisons of self-organization scores.
  • Editorial inference: because $W_{\mathrm{REP}}$ discards spatial arrangement, it and the full utility measure answer different questions, so one should not be read as a substitute for the other.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper develops a mathematical strategy to estimate the reproducibility entropy S_REP of a cell-fate patterning ensemble from the lower-dimensional counting space. The strategy relies on a pseudo-inverse transformation M (Definition 0.12) that assigns equal probability to all patterning vectors sharing the same count vector. The authors derive closed-form expressions for the ideal cell-fate proportioning process (Propositions 0.1–0.7), apply the strategy to a spatially resolved simulation of EPI/PRE proportioning (the ITWT system), and report that the empirical utility is very low (0.001–0.007 bits), concluding that the system is far from optimal. They also propose a counting-space reproducibility entropy W_REP and show that it decreases with n, consistent with earlier findings in [8].

Significance. The combinatorial derivations in Propositions 0.1–0.7 are correct and the W_REP measure is a useful, tractable quantity for comparing count reproducibility across systems. However, the central empirical claim depends on the uniform lift, which destroys all spatial correlations: under M, S_PAT = S_SCF and PI = 0 identically, so the reported low utility is an artifact of the assumption, not a measured property of the simulation. The paper's contribution is therefore methodological, with a clear limitation that is not adequately flagged in the interpretation. The authors are careful to state the uniformity assumption explicitly, which is a strength, but they do not validate it against the spatially resolved simulation data, and the biological conclusions about optimality are not supported by the current analysis.

major comments (3)
  1. [Definition 0.12 and Fig 2] The uniform lift in Definition 0.12 makes P(Z=j, N=i) independent of position i, because all patterns with the same count vector are equiprobable. By Definitions 0.6 and 0.7, this forces E_S_PAT = E_S_SCF and hence E_PI = 0 identically for the empirical estimates in Fig 2. The reported near-zero positional information is therefore a direct consequence of the assumption, not a measurement of the ITWT simulation's spatial structure. Moreover, if the lift were applied consistently, PI would be exactly zero; the non-vanishing empirical PI shown in the Fig 2 inset suggests a numerical artifact or an undocumented deviation from the construction, which needs clarification. The utility values in Fig 2[D] are conditional on the exchangeability assumption, so the claim that the system is 'nowhere near optimal' is not supported. The authors should either compute the three entropies directly from the simulation's pattern ensemble (feasible for small n) or present the results as bounds derived under a stated prior.
  2. [Discussion, p. 18] The assertion that 'the current view of optimal behavior for cell-fate proportioning processes is incomplete' rests on the discrepancy between the low empirical utility and the high accuracy and precision reported in [8]. Since the low utility follows from the maximum-entropy lift (which can only overestimate S_REP given the count probabilities), the discrepancy is an artifact of the estimation method, not evidence about optimality. A concrete test would be to compare the lifted pattern distribution against the actual pattern distribution of the ITWT simulation for small n (e.g., n = 5) and report the true S_REP. If the lift is rejected, the biological conclusion should be withdrawn or substantially qualified.
  3. [Propositions 0.2–0.7 and Fig 3] The ideal process is itself defined using the same uniform lift (Definition 0.12), so both the ideal and empirical utilities are computed under exchangeable pattern distributions. The reported ratio of empirical to ideal utility (about 1/5) therefore reflects only the spread of the counting distribution relative to the delta distribution at the ideal count vector, not any spatial patterning performance. This framing should be stated explicitly throughout the paper and in the interpretation of Figs 2 and 3, and the wording 'self-organization potential' should be revised to indicate that the measure is conditional on the exchangeability prior.
minor comments (4)
  1. [Definition 0.17] In the paragraph following Definition 0.17, the text says 'max(SREP)' where it should say 'max(WREP)' when referring to the maximum of the counting-space entropy; this typo should be corrected.
  2. [Notation, p. 3] The use of '∨' and '∧' for logical OR and AND is unconventional and may confuse readers; standard notation (e.g., 'or' and 'and') is recommended in definitions.
  3. [Data Availability] The statement that 'All code files will be available from a GitHub repository' is not sufficient for review; a permanent repository link or an explicit statement of availability upon request should be provided.
  4. [Figs 2 and 3] The superscript labels 'E' on quantities such as E_SPAT, E_SSCF, and E_SREP are not defined in the figure captions; they should be defined in the captions or in the main text (they appear to denote 'empirical').

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the uniform lift is an explicit assumption used for SREP estimation, not a hidden derivation of the empirical PI.

full rationale

No circular step is established from the text. The pseudo-inverse M in Definition 0.12 does make the lifted pattern distribution exchangeable, so any ensemble generated by M would have S_PAT = S_SCF and PI = 0 by Definitions 0.6 and 0.7. However, the paper's procedure indicates that M is used specifically to estimate the reproducibility entropy: "For each ICM size, to compute the empirical reproducibility entropy, we created a list of all the realizable elements of the respective (sample) counting space, estimating their probabilities based on the observed counting vectors. Following this estimation of the counting probability space, we applied Definition 0.12 (pseudo-inverse transform M) for assigning probability estimates to patterning vector sets, thus obtaining an approximation of the respective patterning probability space." The empirical S_PAT and S_SCF are compared alongside S_REP, but the paper does not state that they were computed from the M-lift; direct estimation from the spatially resolved simulation is the natural reading, and the reported empirical PI is described as small but nonzero ("S_PAT ..., S_SCF ..., and S_REP ... are almost equal," and the utility "never fully vanishes"), which would be exactly zero rather than merely almost equal if the lift had been used for those quantities. The ideal result PI = 0 is a transparent consequence of Definition 0.14, which defines the ideal patterning space by combining the uniform lift with perfect counting reproducibility; deriving it in Propositions 0.2-0.4 is a definitional unpacking, not a hidden equivalence. The uniform lift is an acknowledged modeling assumption that could bias S_REP and hence U, but an untested or potentially biased assumption is a validity concern, not circularity. Reliance on the authors' prior simulation study [8] is ordinary reuse of one's own data and does not reduce the present derivation to a self-citation chain. The central mapping method is mathematically self-contained, and the empirical claims are not shown to be identical to the input assumptions by construction.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central mapping method rests on one ad hoc assumption (uniform lift) plus standard probability and combinatorics. The empirical illustration imports the ideal proportions and multinomial baselines from the authors' prior work, which are not independently verified here.

free parameters (2)
  • Ideal cell-fate proportions (EPI, PRE) = EPI 2/5, PRE 3/5, UND 0
    Taken from the ITWT wild-type simulation in [8]; used to define the ideal counting vector in Definition 0.13 and all ideal benchmarks. Not fitted in this paper, but chosen as the reference target.
  • Multinomial proportions for artificial baseline = Observed simulation proportions at each time
    In the W_REP comparison (Fig 4, case three), cell-fate counts are sampled from a multinomial with proportions taken from the simulated data, which are effectively fitted to the empirical counting distribution.
assumptions (4)
  • ad hoc to paper All cell-fate patterning vectors with the same counting vector are equiprobable (uniform lift, Definition 0.12).
    This is the key modeling assumption that makes S_REP tractable. It is stated as 'suitable' for proportioning processes but is not derived from the simulation dynamics. It forces P(Z=j,N=i) to be independent of i, hence S_PAT=S_SCF and PI=0 for the lifted empirical ensemble.
  • domain assumption The cell index N is uniformly distributed over cells (Definitions 0.3 and 0.4).
    The marginal P(Z=j) is defined as an average over cell positions with uniform weight. This is a natural exchangeability assumption but excludes position-dependent sampling.
  • standard math Shannon entropy and standard combinatorial identities.
    Used throughout for entropy definitions and for the cardinality of the counting space.
  • domain assumption For the ITWT system, the ideal proportions are (UND, EPI, PRE) = (0, 2/5, 3/5).
    Definition 0.13 and the examples use these as the perfect target counts, imported from [8].

how reviews work

0 comments
Cite this review

Pith. "Pith review of Information-Theoretical Measures for Developmental Cell-Fate Proportioning Processes." pith.science (2026). https://pith.science/paper/YTJRLXA5

@misc{pith2026250523659,
  author       = {Pith},
  title        = {Pith review of: Information-Theoretical Measures for Developmental Cell-Fate Proportioning Processes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YTJRLXA5}},
  note         = {Machine review of arXiv:2505.23659}
}
read the original abstract

Self-organization is a fundamental process of complex biological systems, particularly during the early stages of development. In the mammalian embryo, blastocyst formation exemplifies a self-organized system, involving the correct spatio-temporal segregation of three distinct cell fates: trophectoderm (TE), epiblast (EPI), and primitive endoderm (PRE). Despite the significance of this class of processes, quantifying the information content of self-organizing patterning systems remains challenging due to the complexity and the qualitative diversity of developmental mechanisms. In this study, we applied a recently proposed information-theoretical framework which quantifies the self-organization potential of cell-fate patterning systems, employing a utility function that integrates (local) positional information and (global) correlational information extracted from developmental pattern ensembles. Specifically, we examined a stochastic and spatially resolved simulation model of EPI-PRE lineage proportioning, evaluating its information content across various simulation scenarios with different numbers of system cells. To overcome the computational challenges hindering the application of this novel framework, we developed a mathematical strategy that indirectly maps the low-dimensional cell-fate counting probability space to the high-dimensional cell-fate patterning probability space, enabling the estimation of self-organization potential for general cell-fate proportioning processes. Overall, this novel information-theoretical framework provides a promising, universal approach for quantifying self-organization in developmental biology. By formalizing measures of self-organization, the employed quantification framework offers a valuable tool for uncovering insights into the underlying principles of cell-fate specification and the emergence of complexity in early developmental systems.

Figures

Figures reproduced from arXiv: 2505.23659 by the authors.

Figure 1
Figure 1. Illustrative example: n = 5 and z = 3. Facilitating the understanding of our notation, we show an illustrative example for a fixed number of system cells n = 5 and a fixed number of cell fates z = 3. [A] Patterning space Ω(Z⃗). [B] Probability heatmaps for uniform (left), empirical (center), and ideal (right) patterning spaces. [C] Counting space Ω(N⃗ ). [D] Probability heatmaps for images under T (Definition 0.11) … view at source ↗
Figure 2
Figure 2. Comparing empirical information-theoretical measures against their ideal [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Comparing empirical information-theoretical measures against their ideal [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Reproducibility entropy (counting space) with respect to ICM size: [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: Reproducibility entropy (counting space) with respect to simulation time [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: Comparison between patterning Ω(Z⃗) and counting Ω(N⃗ ) space cardinalities. Here, if n ≫ 1, then ((n + 1)(n + 2))/2 ≪ z n. Notation: number of system cells n; number of cell fates z. September 10, 2025 22/25 [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: Counting vector probability heatmaps: ITWT system. [PITH_FULL_IMAGE:figures/full_fig_p023_7.png]
Figure 8
Figure 8. Figure 8: Comparing empirical information-theoretical measures against their ideal [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: Reproducibility entropy (counting space) with respect to simulation time [PITH_FULL_IMAGE:figures/full_fig_p025_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

13 extracted references · 5 canonical work pages

  1. [8]

    AI-powered simulation-based inference of a gen- uinely spatial-stochastic gene regulation model of early mouse embryogenesis

    Ramirez Sierra MA, Sokolowski TR. AI-powered simulation-based inference of a gen- uinely spatial-stochastic gene regulation model of early mouse embryogenesis. PLOS Computational Biology. 2024;20(11):e1012473. doi:10.1371/journal.pcbi.1012473

  2. [1]

    Information content and optimization of self-organized developmen- tal systems

    Br¨ uckner DB, Tkaˇ cik G. Information content and optimization of self-organized developmen- tal systems. Proceedings of the National Academy of Sciences. 2024;121(23):e2322326121. doi:10.1073/pnas.2322326121

  3. [2]

    Common principles of early mammalian embryo self-organisation

    P lusa B, Piliszek A. Common principles of early mammalian embryo self-organisation. Development. 2020;147(dev183079). doi:10.1242/dev.183079

  4. [3]

    Local cellular interactions during the self-organization of stem cells

    Schr¨ oter C, Stapornwongkul KS, Trivedi V. Local cellular interactions during the self-organization of stem cells. Current Opinion in Cell Biology. 2023;85:102261. doi:10.1016/j.ceb.2023.102261

  5. [4]

    Principles of Self-Organization of the Mammalian Embryo

    Zhu M, Zernicka-Goetz M. Principles of Self-Organization of the Mammalian Embryo. Cell. 2020;183(6):1467–1478. doi:10.1016/j.cell.2020.11.003

  6. [5]

    From embryos to embryoids: How external signals and self-organization drive embryonic development

    Morales JS, Raspopovic J, Marcon L. From embryos to embryoids: How external signals and self-organization drive embryonic development. Stem Cell Reports. 2021;16(5):1039–1050. doi:10.1016/j.stemcr.2021.03.026

  7. [6]

    Understanding the cell: Future views of structural biology

    Beck M, Covino R, H¨ anelt I, M¨ uller-McNicoll M. Understanding the cell: Future views of structural biology. Cell. 2024;187(3):545–562. doi:10.1016/j.cell.2023.12.017

  8. [7]

    Lumen Expansion Facilitates Epiblast-Primitive Endoderm Fate Specification during Mouse Blastocyst Formation

    Ryan AQ, Chan CJ, Graner F, Hiiragi T. Lumen Expansion Facilitates Epiblast-Primitive Endoderm Fate Specification during Mouse Blastocyst Formation. Developmental Cell. 2019;51(6):684–697.e4. doi:10.1016/j.devcel.2019.10.011

Show all 13 references
  1. [9]

    Decoding of position in the developing neural tube from antiparallel morphogen gradients

    Zagorski M, Tabata Y, Brandenberg N, Lutolf MP, Tkaˇ cik G, Bollenbach T, et al. Decoding of position in the developing neural tube from antiparallel morphogen gradients. Science. 2017;356(6345):1379–1383. doi:10.1126/science.aam5887

  2. [10]

    The many bits of positional information

    Tkaˇ cik G, Gregor T. The many bits of positional information. Development. 2021;148(dev176065). doi:10.1242/dev.176065. September 10, 2025 19/25

  3. [11]

    Latent space of a small genetic network: Geometry of dynamics and information

    Seyboldt R, Lavoie J, Henry A, Vanaret J, Petkova MD, Gregor T, et al. Latent space of a small genetic network: Geometry of dynamics and information. Proceedings of the National Academy of Sciences. 2022;119(26):e2113651119. doi:10.1073/pnas.2113651119

  4. [12]

    Deriving a genetic regulatory network from an optimization principle; 2023

    Sokolowski TR, Gregor T, Bialek W, Tkaˇ cik G. Deriving a genetic regulatory network from an optimization principle; 2023. Available from:http://arxiv.org/abs/2302.05680

  5. [13]

    Center for Multiscale Modelling in Life Sciences

    Ramirez Sierra MA, Sokolowski TR. Comparing AI versus optimization workflows for simulation-based inference of spatial-stochastic systems. Machine Learning: Science and Technology. 2025;6(1):010502. doi:10.1088/2632-2153/ada0a3. September 10, 2025 20/25 Acknowledgments The suc...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.