REVIEW 3 major objections 5 minor 54 references
Surveying the space of descriptions of a composite system with machine learning
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that the continuous space of lossy compressions of a composite system—navigated by neural networks—reveals organizational structure that discrete-subsystem analyses miss.
desk verdict A genuinely new object of study—the continuous space of partial descriptions—with an honest, well-demonstrated ML framework; the main soft spot is the unvalidated InfoNCE lower bound for minimization, which the authors acknowledge. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the 'description' $U=(U_1,\dots,U_N)$: a set of probabilistic encodings, one per component, where each $U_i$ is generated from $X_i$ alone plus independent noise, so it carries information only about that component. This turns the choice of what to pay attention to in each component into a point in a continuous space, with the transmitted information $I(X_i;U_i)$ as coordinates. The identities that carry the argument are $\mathrm{TC}(U)=\sum_i I(X_i;U_i)-I(X;U)$ and $\Omega(U)=(N-2)I(U;X)+\sum_i[I(U_i;X_i)-I(U_{-i};X_{-i})]$, which express two standard multivariate-information summaries directly in terms of the description's channels. The optimization machinery uses a constraint on the total component information, InfoNCE (a contrastive estimator of mutual information) for joint terms, and an adversarial setup when the target quantity must be minimized; a post-hoc hardening step converts soft encodings to interpretable partitions of each component's outcomes.
What would settle it
Enumerate all hard compression schemes for a small case, such as the 4×4 sudoku with its $15^{16}$ possible hard descriptions reduced by symmetries, and compare the true extremal O-information or total correlation at a fixed component-information budget with the values found by the optimizer; any systematic gap would falsify the claim that the plotted curves bound the description space.
Extended reading notes
Core claim
The central discovery, on the paper's own terms, is that the space of descriptions—the set of possible entropy allocations to the components—is a meaningful organizational object, and that its extremal points are interpretable and computable. Concretely, for a system $X=(X_1,\dots,X_N)$, a description $U=(U_1,\dots,U_N)$ is any collection of channels $U_i=f_i(X_i,\epsilon_i)$ with $U_i$ conditionally independent of $X_j$ for $j\neq i$ given $X_i$; the space of such descriptions is equivalent to the space of lossy compression schemes of the components. The authors show that total correlation satisfies $\mathrm{TC}(U)=\sum_i I(X_i;U_i)-I(X;U)$ and that O-information has an analogous expression, so both can be optimized over the continuous space of encodings. In the cases studied, the optimized boundaries trace the outside of randomly sampled descriptions, and hardened versions of the extremal descriptions isolate compartments that natural subsystem accounts miss: the ferromagnetic chain versus the frustrated triangle in the spin system, diagonal-symmetric synergy patterns in sudoku, and word-initial 'th' clusters versus word-final 'ed'/'ing' clusters in English.
Load-bearing premise
The load-bearing premise is that the neural-network optimization actually finds descriptions close to the true extrema of the chosen information quantity; the authors state there is no guarantee of global optimality, so if the searches settle in biased local solutions the plotted boundaries and the qualitative redundancy/synergy findings could be artifacts of the search.
Editorial extensions
If this is right
- For systems too large for partial information decomposition—whose number of terms grows superexponentially and becomes impractical beyond about five components—the description-space approach offers a tractable, fine-grained alternative that can still identify redundancy- and synergy-dominated regimes.
- The extremal descriptions themselves function as a selection device: in the spin system they localize the ferromagnetic chain and the frustrated triangle, in sudoku they single out diagonal-symmetric partial-information patterns, and in English 4-grams they surface 'th', 'ed', and 'ing' as the groupings that most contribute to total correlation.
- Because every learned mapping defines a valid description, even suboptimal optimizations yield meaningful points in the space; the method's usefulness does not depend on guaranteed convergence to a global optimum.
- The framework is not restricted to discrete variables or to total correlation and O-information; it extends to continuous systems and can extremize any summary quantity built from mutual information terms, including binding entropy, S-information, $\Delta I$, TSE complexity, and specific PED atoms.
Reading between the lines
- A consequence not drawn in the paper: if the optimized boundaries are stable across random initializations, the shape of the description space could serve as a system-level fingerprint, allowing quantitative comparison of organizational similarity between systems of different sizes or types without aligning their components.
- The sudoku result that optimal descriptions lie beyond discrete subsystems suggests a general principle for constraint-satisfaction problems: the most informative coarse variables are soft mixtures spread across many variables, not subsets of variables. This could be tested by applying the same optimizer to other constraint problems such as graph coloring or SAT.
- A natural extension would be to certify the extrema, for example by comparing the InfoNCE-based estimates against exhaustive enumeration on small systems or against tighter variational bounds; if the boundaries were certified, the inferred redundancy/synergy labels would become testable claims rather than search artifacts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes to study a composite random system through the continuous space of 'descriptions,' where a description is a tuple of lossy per-component channels (U_i from X_i, conditionally independent across components). It derives expressions for total correlation and O-information of such descriptions and introduces a neural-network optimization framework that extremizes these quantities under a constraint on total component information. The method is demonstrated on a 5-spin Ising system, 4x4 sudoku boards, and 4-grams from English Wikipedia. The authors report that optimized descriptions trace the boundaries of randomly sampled descriptions for the spin system, exceed discrete-subsystem descriptions for sudoku, and reveal structure in natural-language 4-grams. The paper includes a released codebase and states that every learned map is a valid description, so suboptimal solutions are still interpretable points.
Significance. If the optimization reliably finds near-extremal descriptions, the framework offers a scalable alternative to coarse discrete-subsystem analyses and to PID/PED, whose term counts grow superexponentially. The formal identities in Eqs. (1)-(2) are correct given the conditional-independence structure, and the paper's honest caveat about global optimality is welcome. The case studies are well chosen and the comparisons against random sampling for the spin and sudoku systems provide some evidence that the optimizer finds extreme points. The released code and reproducible experimental setup are strengths. However, the empirical nature of the search means that the central visual and qualitative conclusions—the plotted boundaries and the redundant/synergistic regimes—depend on the reliability of the surrogate training objectives, and this reliability is not directly established.
major comments (3)
- [Methods, 'To find descriptions...' and 'For evaluation'] The mutual information terms that are minimized during training are estimated with the InfoNCE lower bound (Eq. 3), while the reported values are computed with the Monte Carlo likelihood-ratio estimator. No calibration between the two estimators is reported. Because InfoNCE is a lower bound whose tightness depends on the critic family (an MLP with squared-Euclidean similarity in 32 dimensions) and batch size, an encoder trained against the bound can drive the bound down while the true mutual information remains large. Since the plotted boundaries are evaluated with the MC estimator, the displayed curves could then correspond to points that are not near true extrema of the target quantity. I request a concrete calibration check: for at least one system, plot the InfoNCE bound estimate and the MC estimate along the optimization trajectory (or at final checkpoints) to show the two agree within tolerance, or otherwise demonstrate that the adversarial minimization is not exploiting a loose bound.
- [Sudoku section, 'The hardened compression schemes...'] The conversion from soft optimized descriptions to hard clustering via Bhattacharyya-coefficient gradient descent is a post-processing step that can change the values of the extremized quantity, yet no quantitative comparison is reported between the soft descriptions and their hardened counterparts. The stars in Figs. 2a and 3a are interpreted as extremal schemes, but if hardening shifts a star away from the boundary, the interpretive claims about specific digits or letters being clustered would not be supported by the extremization. Please report the target-quantity values before and after hardening for each displayed star, or state explicitly when the hardened scheme is re-evaluated and shown to lie on the same boundary.
- [Sudoku section, 'no guarantee of global optimality'] The paper acknowledges that there is no guarantee of global optimality, and for the sudoku and n-gram systems the only validation is that optimized descriptions lie several standard deviations from randomly sampled descriptions (Fig. 2b). This establishes that the optimizer finds unusual descriptions, but not that they are extremal or that the qualitative conclusions (e.g., that three-bit spin descriptions split into redundant and synergistic regimes, or that the most synergistic sudoku description is scheme vi) are robust to local optima and surrogate-objective bias. I recommend adding a sensitivity analysis: vary hyperparameters, random seeds, and the InfoNCE critic capacity for at least one system and report the spread of achieved objective values at fixed I_in. This would quantify how much of the claimed boundary shape depends on the specific search configuration.
minor comments (5)
- [Abstract] The phrase 'opens a new avenues' contains a subject-verb agreement error and should read 'opens new avenues' or 'opens a new avenue.'
- [Fig. 1b-d] The description of the random-sampling baseline for the spin system says the method 'closely trace[s] the bounds of the randomly sampled descriptions,' but the reader cannot verify this because the light blue trace is plotted on top of the gray dots; please add a version of the figure without the optimizing trace, or a quantitative measure of how close the optimized points are to the convex hull of the sampled points.
- [O-information definition, Eq. (2)] The notation I(U_{-i}; X_{-i}) is introduced in Eq. (2) but the definition of the slash subscript is first given in the following sentence; consider defining U_{-i} before the equation to avoid ambiguity.
- [Sudoku hardening paragraph] The claim that Bhattacharyya-coefficient gradient descent drives 'perfect distinguishability/indistinguishability' is not accompanied by a stopping criterion or a measure of how close the final schemes are to hard clusters; please specify the tolerance used in practice.
- [N-gram section, Fig. 3] The text says that '4-letter words contain variation that is more redundant' and that 'the most negative O-information occurs around 10–12 bits,' but the corresponding curves in Fig. 3b are not annotated with the full-information values; adding vertical markers for the full-information point would improve readability.
Circularity Check
No significant circularity: extremal descriptions are optimization outputs, not re-labeled inputs.
full rationale
The paper defines a description as a collection of channels p(u_i|x_i) with conditional independence, and the reported quantities TC(U) and Ω(U) are direct functions of those channels; Eqs. (1) and (2) are derived identities, not assumptions of the conclusions. The extremal descriptions are found by numerical search over channel parameters, and the reported values are evaluated with an independent Monte Carlo likelihood-ratio estimator (Eq. 6), which is distinct from the InfoNCE and likelihood-ratio bounds used during training. Thus the plotted quantities are not equal to the training losses by construction. The paper explicitly states that 'there is no guarantee of global optimality' (sudoku section), and this is an optimization-reliability caveat rather than evidence of circularity. The few self-citations ([28], [29], [30]) appear in methodological contexts (constrained feature information, information bottleneck, Bhattacharyya hardening) and are not load-bearing for the central claims; the core quantities are standard, and the random-sampling baselines provide independent external checks. No step in the derivation chain reduces by definition or by self-citation to its own input, so no circularity is present.
Assumptions & free parameters
assumptions (3)
- domain assumption The true distribution p(x) is known for each system (Boltzmann law for spins, uniform over valid sudoku boards, empirical n-gram frequencies).
- domain assumption InfoNCE provides a sufficiently accurate lower bound on mutual information for the multi-component terms used in the optimization.
- domain assumption The adversarial training for minimizing I(U;X) does not introduce systematic bias.
Cite this review
Pith. "Pith review of Surveying the space of descriptions of a composite system with machine learning." pith.science (2026). https://pith.science/paper/BCJAY4IN
@misc{pith2026241118579,
author = {Pith},
title = {Pith review of: Surveying the space of descriptions of a composite system with machine learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/BCJAY4IN}},
note = {Machine review of arXiv:2411.18579}
}
read the original abstract
Multivariate information theory provides a general and principled framework for understanding how the components of a complex system are connected. Existing analyses are coarse in nature -- built up from characterizations of discrete subsystems -- and can be computationally prohibitive. In this work, we propose to study the continuous space of possible descriptions of a composite system as a window into its organizational structure. A description consists of specific information conveyed about each of the components, and the space of possible descriptions is equivalent to the space of lossy compression schemes of the components. We introduce a machine learning framework to optimize descriptions that extremize key information theoretic quantities used to characterize organization, such as total correlation and O-information. Through case studies on spin systems, sudoku boards, and letter sequences from natural language, we identify extremal descriptions that reveal how system-wide variation emerges from individual components. By integrating machine learning into a fine-grained information theoretic analysis of composite random variables, our framework opens a new avenues for probing the structure of real-world complex systems.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
G. Nicolis and C. Nicolis,Foundations of complex sys- tems: emergence, information and predicition(World Scientific, 2012)
work page 2012
-
[3]
J. Ladyman, J. Lambert, and K. Wiesner, What is a complex system?, European Journal for Philosophy of Science3, 33 (2013)
work page 2013
-
[4]
B. C. Daniels, C. J. Ellison, D. C. Krakauer, and J. C. Flack, Quantifying collectivity, Current opinion in neu- robiology37, 106 (2016)
work page 2016
-
[5]
F. E. Rosas, P. A. M. Mediano, M. Gastpar, and H. J. Jensen, Quantifying high-order interdependencies via multivariate extensions of the mutual information, Phys. Rev. E100, 032305 (2019)
2019
-
[6]
D. A. Ehrlich, A. C. Schneider, V. Priesemann, M. Wibral, and A. Makkeh, A measure of the complex- ity of neural representations based on partial informa- tion decomposition, Transactions on Machine Learning Research (2023)
work page 2023
-
[7]
L. Martignon, G. Deco, K. Laskey, M. Diamond, W. Frei- wald, and E. Vaadia, Neural coding: higher-order tem- poral patterns in the neurostatistics of cell assemblies, Neural computation12, 2621 (2000)
work page 2000
- [8]
Show all 54 references
-
[9]
T. F. Varley, M. Pope, M. Grazia, Joshua, and O. Sporns, Partial entropy decomposition reveals higher-order infor- mation structures in human brain activity, Proceedings of the National Academy of Sciences120, e2300888120 6 (2023)
2023
-
[10]
A. I. Luppi, F. E. Rosas, P. A. Mediano, D. K. Menon, and E. A. Stamatakis, Information decomposition and the informational architecture of the brain, Trends in Cognitive Sciences (2024)
2024
-
[11]
M. Pope, T. F. Varley, and O. Sporns, Time-varying syn- ergy/redundancy dominance in the human cerebral cor- tex, bioRxiv , 2024 (2024)
2024
-
[12]
J. M. Miller, X. R. Wang, J. T. Lizier, M. Prokopenko, and L. F. Rossi, Measuring information dynamics in swarms, inGuided self-organization: Inception(Springer,
-
[13]
Pilkiewicz, B
K. Pilkiewicz, B. Lemasson, M. Rowland, A. Hein, J. Sun, A. Berdahl, M. Mayo, J. Moehlis, M. Porfiri, E. Fern´ andez-Juricic,et al., Decoding collective commu- nications using information theory tools, Journal of the Royal Society Interface17, 20190563 (2020)
2020
-
[14]
C. R. Twomey, A. T. Hartnett, M. M. Sosna, and P. Ro- manczuk, Searching for structure in collective systems, Theory in Biosciences140, 361 (2021)
2021
-
[15]
A. L. Burns, T. M. Schaerf, J. Lizier, S. Kawaguchi, M. Cox, R. King, J. Krause, and A. J. Ward, Self- organization and information transfer in antarctic krill swarms, Proceedings of the Royal Society B289, 20212361 (2022)
2022
-
[16]
T. E. Chan, M. P. Stumpf, and A. C. Babtie, Gene regula- tory network inference from single-cell data using multi- variate information measures, Cell systems5, 251 (2017)
2017
-
[17]
Sootla, D
S. Sootla, D. O. Theis, and R. Vicente, Analyzing infor- mation distribution in complex systems, Entropy19, 636 (2017)
2017
-
[18]
Scagliarini, D
T. Scagliarini, D. Nuzzi, Y. Antonacci, L. Faes, F. E. Rosas, D. Marinazzo, and S. Stramaglia, Gradients of o-information: Low-order descriptors of high-order de- pendencies, Phys. Rev. Res.5, 013025 (2023)
2023
-
[19]
Watanabe, Information theoretical analysis of multi- variate correlation, IBM Journal of research and devel- opment4, 66 (1960)
S. Watanabe, Information theoretical analysis of multi- variate correlation, IBM Journal of research and devel- opment4, 66 (1960)
1960
-
[20]
R. G. James, C. J. Ellison, and J. P. Crutchfield, Anatomy of a bit: Information in a time series obser- vation, Chaos: An Interdisciplinary Journal of Nonlinear Science21(2011)
2011
-
[21]
Marois and J
R. Marois and J. Ivanoff, Capacity limits of information processing in the brain, Trends in cognitive sciences9, 296 (2005)
2005
-
[22]
P. L. Williams and R. D. Beer, Nonnegative decom- position of multivariate information, arXiv preprint arXiv:1004.2515 (2010)
2010 arXiv
-
[23]
R. A. Ince, The partial entropy decomposition: De- composing multivariate entropy and mutual informa- tion via pointwise common surprisal, arXiv preprint arXiv:1702.01591 (2017)
2017 arXiv
-
[24]
Timme, W
N. Timme, W. Alford, B. Flecker, and J. M. Beggs, Synergy, redundancy, and multivariate information mea- sures: an experimentalist’s perspective, Journal of com- putational neuroscience36, 119 (2014)
2014
-
[25]
Kolchinsky, Partial information decomposition: Re- dundancy as information bottleneck, Entropy26, 546 (2024)
A. Kolchinsky, Partial information decomposition: Re- dundancy as information bottleneck, Entropy26, 546 (2024)
2024
-
[26]
A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, Deep variational information bottleneck, inInternational Conference on Learning Representations(2017)
2017
-
[27]
Koch-Janusz and Z
M. Koch-Janusz and Z. Ringel, Mutual information, neural networks and the renormalization group, Nature Physics14, 578 (2018)
2018
-
[28]
K. A. Murphy and D. S. Bassett, Interpretability with full complexity by constraining feature information, in International Conference on Learning Representations (ICLR)(2023)
2023
-
[29]
K. A. Murphy and D. S. Bassett, Information decom- position in complex systems via machine learning, Pro- ceedings of the National Academy of Sciences121, e2312988121 (2024)
2024
-
[30]
K. A. Murphy and D. S. Bassett, Machine-learning op- timized measurements of chaotic dynamical systems via the information bottleneck, Phys. Rev. Lett.132, 197201 (2024)
2024
-
[31]
T. M. Cover and J. A. Thomas,Elements of information theory(John Wiley & Sons, 1999)
1999
-
[32]
C. P. Burgess, I. Higgins, A. Pal, L. Matthey, N. Watters, G. Desjardins, and A. Lerchner, Understanding disentan- gling inβ-V AE, arXiv preprint arXiv:1804.03599 (2018)
2018 arXiv
-
[33]
Poole, S
B. Poole, S. Ozair, A. Van Den Oord, A. Alemi, and G. Tucker, On variational bounds of mutual informa- tion, inInternational Conference on Machine Learning (PMLR, 2019) pp. 5171–5180
2019
-
[34]
A. v. d. Oord, Y. Li, and O. Vinyals, Representa- tion learning with contrastive predictive coding, arXiv preprint arXiv:1807.03748 (2018)
2018 arXiv
-
[35]
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, A simple framework for contrastive learning of visual repre- sentations, inInternational conference on machine learn- ing(PMLR, 2020) pp. 1597–1607
2020
-
[36]
C. P. Gomes, B. Selman, N. Crato, and H. Kautz, Heavy- tailed phenomena in satisfiability and constraint satis- faction problems, Journal of automated reasoning24, 67 (2000)
2000
-
[37]
Ercsey-Ravasz and Z
M. Ercsey-Ravasz and Z. Toroczkai, The chaos within sudoku, Scientific reports2, 725 (2012)
2012
-
[38]
Varga, R
M. Varga, R. Sumi, Z. Toroczkai, and M. Ercsey-Ravasz, Order-to-chaos transition in the hardness of random boolean satisfiability problems, Physical Review E93, 052211 (2016)
2016
-
[39]
Kailath, The divergence and bhattacharyya distance measures in signal selection, IEEE transactions on com- munication technology15, 52 (1967)
T. Kailath, The divergence and bhattacharyya distance measures in signal selection, IEEE transactions on com- munication technology15, 52 (1967)
1967
-
[40]
K. A. Murphy, S. Dillavou, and D. S. Bassett, Compar- ing the information content of probabilistic representa- tion spaces, Transactions on Machine Learning Research (2025)
2025
-
[41]
C. E. Shannon, A mathematical theory of communica- tion, The Bell system technical journal27, 379 (1948)
1948
-
[42]
G. J. Stephens and W. Bialek, Statistical mechanics of letters in words, Physical Review E—Statistical, Nonlin- ear, and Soft Matter Physics81, 066119 (2010)
2010
-
[43]
Scagliarini, D
T. Scagliarini, D. Marinazzo, Y. Guo, S. Stramaglia, and F. E. Rosas, Quantifying high-order interdependencies on individual patterns via the local o-information: Theory and applications to music analysis, Physical Review Re- search4, 013184 (2022)
2022
-
[44]
Nirenberg, S
S. Nirenberg, S. M. Carcieri, A. L. Jacobs, and P. E. Latham, Retinal ganglion cells act largely as independent encoders, Nature411, 698 (2001)
2001
-
[45]
Maliniak, R
D. Maliniak, R. Powers, and B. F. Walter, The gender citation gap in international relations, International Or- ganization67, 889 (2013)
2013
-
[46]
Caplar, S
N. Caplar, S. Tacchella, and S. Birrer, Quantitative eval- uation of gender bias in astronomical publications from 7 citation counts, Nature Astronomy1, 1 (2017)
2017
-
[47]
Chakravartty, R
P. Chakravartty, R. Kuo, V. Grubbs, and C. McIlwain, #CommunicationSoWhite, Journal of Communication 68, 254 (2018)
2018
-
[48]
M. L. Dion, J. L. Sumner, and S. M. Mitchell, Gen- dered citation patterns across political science and so- cial science methodology fields, Political Analysis26, 312 (2018)
2018
-
[49]
J. D. Dworkin, K. A. Linn, E. G. Teich, P. Zurn, R. T. Shinohara, and D. S. Bassett, The extent and drivers of gender imbalance in neuroscience reference lists, Nature Neuroscience23, 918 (2020)
2020
-
[50]
E. G. Teich, J. Z. Kim, C. W. Lynn, S. C. Simon, A. A. Klishin, K. P. Szymula, P. Srivastava, L. C. Bassett, P. Zurn, J. D. Dworkin,et al., Citation inequity and gen- dered citation practices in contemporary physics, Nature Physics18, 1161 (2022)
2022
-
[51]
P. Zurn, D. S. Bassett, and N. C. Rust, The citation diversity statement: a practice of transparency, a way of life, Trends in Cognitive Sciences24, 669 (2020)
2020
-
[52]
Dworkin, P
J. Dworkin, P. Zurn, and D. S. Bassett, (In)citing action to realize an equitable future, Neuron106, 890 (2020)
2020
-
[53]
D. Zhou, E. J. Cornblath, J. Stiso, E. G. Teich, J. D. Dworkin, A. S. Blevins, and D. S. Bassett, Gender diver- sity statement and code notebook v1. 0, Zenodo (2020)
2020
-
[54]
Encoder MLP architecture
Z. Budrikis, Growing citation gender gap, Nature Re- views Physics2, 346 (2020). S1 Supplemental Material The code base has been released on Github at https://github.com/murphyka/description space. The analyses of the three systems from the main text can be repeated with the t...
2020
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.