REVIEW 3 major objections 5 minor 19 references
Compression, Regularity, Randomness and Emergent Structure: Rethinking Physical Complexity in the Data-Driven Era
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that every complexity measure can be placed on three axes—regularity, randomness, and complexity—and that this map reveals why the most structural measures resist computation.
desk verdict A plausible but under-supported taxonomy of complexity measures—useful as a conceptual map, but the central 'unified framework' claim rests on subjective table entries, so treat it as a well-organized survey rather than a systematic classification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the three-axis conceptual landscape—regularity, randomness, and complexity—together with the depth–accessibility plane that classifies each measure by theoretical richness versus ease of estimation. The formal anchor is the compression view: a sequence is regular when its Kolmogorov complexity $K(x)$ is much shorter than the sequence, random when $K(x)$ approaches its length, and complex when it is compressible only through non-trivial computational effort. The paper adopts a coarse-graining map $\pi:\Omega_{\text{micro}}\to\Omega_{\text{macro}}$ as the prerequisite for separating regularities $R(x)$ from noise, and reads latent-space models—encoder-decoder networks, latent ODEs, Koopman autoencoders, symbolic regression, and physics-informed networks—as operational implementations of these same compression and rule-discovery steps.
What would settle it
Recruit, say, ten researchers familiar with these measures and have them independently fill in Tables V and VI from the definitions alone, then compute inter-rater agreement such as Cohen's kappa; low agreement would falsify the claim that the taxonomy is a systematic organization rather than one author's perspective. A quantitative supplement would be to verify axis assignments on synthetic systems with known regularity and randomness—for instance, periodic, chaotic, and white-noise sequences—and check whether each measure behaves as its table placement predicts.
Extended reading notes
Core claim
The paper's central discovery, on its own terms, is that complexity measures are not interchangeable tools but probes with distinct sensitivities that can be systematically charted. Statistical entropies (Shannon, Rényi, Tsallis, approximate entropy, sample entropy, permutation entropy) mostly register randomness, with at most weak purchase on regularity. Algorithmic measures (Kolmogorov complexity, effective complexity, logical depth, sophistication, statistical complexity) capture regularity directly and define complexity as the boundary between structure and noise, but at the price of uncomputability or severe estimation difficulty. Dynamical measures (entropy rate, transfer entropy, active information storage, information modification) straddle the axes by tracking how information is stored, transferred, and transformed in time. The paper claims that placing these measures on the regularity–randomness–complexity triangle and on a depth–accessibility plane clarifies why two metrics can disagree on the same dataset, and that machine-learned latent representations operationalize the same trade-offs.
Load-bearing premise
The whole framework rests on the assumption that each measure has a definite place on the three axes, but those placements are the author's qualitative judgment, so another analyst might place them differently; if placements are not reproducible, the claimed organization is not systematic.
Editorial extensions
If this is right
- Measure selection becomes a deliberate act: a researcher studying noise should reach for entropy-type measures, one studying hidden structure needs algorithmic or dynamical measures, and one studying the order-disorder boundary needs a complexity measure plus a stated coarse-graining.
- The practical question for foundational measures shifts from 'is it computable?' to 'how well and with what bias can it be approximated?'
- Deep learning architectures acquire a principled interpretation: autoencoders approximate minimal description length, symbolic regression approximates effective complexity, and physics-informed networks anchor regularity extraction.
- The depth–accessibility trade-off predicts specific proxy failure modes—VAEs oversmooth rare structure, latent ODEs oversmooth sharp dynamics, and PINNs can enforce wrong priors—making synthetic benchmarking the recommended validation.
- Future learning systems could treat compressibility and regularity extraction as explicit objectives, tying training dynamics to the classical complexity axes.
Reading between the lines
- The axis placements in Tables V and VI are qualitative; a natural next step, not taken here, is to turn them into a rating instrument and measure inter-rater agreement, which would test whether the taxonomy is objective or one analyst's interpretation.
- The same three axes could organize model-selection criteria such as AIC, BIC, and MDL, connecting the taxonomy to mainstream statistical learning theory.
- The latent-space claim suggests a concrete diagnostic: measure the compression ratio of learned latent codes (for example, gzip size of $z$) across datasets and models; if strong reconstruction consistently coincides with low latent complexity, the 'operational arena' story gains empirical support.
- The decision guide in Table VIII can be read as falsifiable predictions—VAEs underestimate heavy-tailed structure and compression proxies miss dynamical correlations—that benchmark tests on the logistic map and Lorenz attractor could confirm or refute.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a conceptual taxonomy for complexity measures, organizing statistical, algorithmic, and dynamical measures along three axes—regularity, randomness, and complexity—and situating them in a common conceptual space. It reviews the mathematical definitions of these measures, places them in two classification tables, discusses their computational accessibility and approximability, and argues that modern deep-learning methods (autoencoders, latent ODEs, symbolic regression, and physics-informed neural networks) act as pragmatic approximations to classical uncomputable complexity ideals. The paper closes with a proxy-selection guide and a research outlook for linking complexity theory with data-driven discovery.
Significance. If the taxonomy were grounded in a reproducible, rule-based classification, it would be a genuinely useful synthesis: it would help practitioners choose among the many available complexity measures and clarify what each measure does and does not quantify. The manuscript is a broad, mostly accurate survey with a compression-centric framing that connects classical information/algorithmic theory to modern machine learning, and it candidly discusses the limitations of neural proxies. Its main value is organizational rather than novel mathematical contribution. The central weakness is that the axis placements in Tables V and VI are asserted through qualitative judgment rather than derived from the definitions; this directly affects the paper's central claim of providing a 'unified framework' and 'systematic conceptual organization.'
major comments (3)
- [§IV.A; Tables V and VI] The central claim of a systematic organization of measures is not supported by a reproducible rule for assigning axis emphases. The paper states in §IV.A that the placement in Table V 'serves as a conceptual guide rather than a precise coordinate system,' but the abstract and conclusion claim a unified framework; these two statements are in tension. The coarse-graining map π introduced in §II.C is never used to derive any entry in Tables V or VI, and the symbols 'weak,' 'some,' and 'approx.' have no thresholds or decision procedure. Without an operational criterion for what it means for a measure to 'capture randomness' or 'capture regularity,' the tables are not falsifiable, and a different analyst could plausibly assign different placements.
- [Table V, Statistical Complexity row; §III.B.d] The placement of Statistical Complexity as capturing regularity and complexity but not randomness is not derivable from its definition. The definition in §III.B.d is Cµ = H[S], the Shannon entropy of the distribution over causal states; entropy of a state distribution generally responds to the number of states and their probabilities, which is an unpredictability-related quantity. The table's '–' in the Randomness column is therefore a qualitative judgment rather than a consequence of the formal definition. This example illustrates why the placements in Tables V and VI need either a formal decision rule or an explicit repositioning of the paper's claims.
- [§III.A.1 vs. Table I] The prose and the table contradict each other. §III.A.1 states that statistical entropy measures 'weakly capture regularity' and 'capture complexity only in limited ways,' but Table I lists Shannon, Rényi, and Tsallis entropies as 'No' for both Regularity and Complexity. Since the tables are the paper's main deliverable, such internal inconsistencies in the central taxonomy need to be resolved before the framework can be considered reliable.
minor comments (5)
- [Abstract and §IV.A] The abstract calls the axes 'orthogonal,' while §IV.A describes them as 'three distinct yet intertwined' properties and §V.A says 'orthogonal, but inter-dependent.' Please align the terminology, since orthogonality is never formally defined in the paper.
- [§III.A.e] There is a duplicated word in the sentence 'a multiscale extension called, called Multiscale Sample Entropy'; please correct it.
- [§III.B.b] The notation for logical depth is inconsistent: both k(x) and K(x) appear for Kolmogorov complexity, and the program length is written as both |P| and |p| in the same passage. Please unify the notation.
- [References] Reference [3] is incomplete: the entry reads 'A. N. Kolmogorov and. Three approaches to the quantitative definition of information' with a missing co-author name.
- [Throughout] The equations in Sections III.A, III.B, and III.C are not numbered, which makes it unnecessarily difficult to refer to specific definitions in the comparative discussion.
Circularity Check
No significant circularity: the paper is a conceptual taxonomy whose classifications are qualitative and explicitly non-coordinate; no derivation reduces to its inputs.
full rationale
This paper is a review and conceptual taxonomy, not a derivation. It defines three axes (regularity, randomness, complexity) and then classifies existing measures by qualitative judgment. I found no step in which a result is derived from, or predicted by, a fitted parameter, a self-cited theorem, or a definition that secretly contains the conclusion. The only formal machinery introduced, the coarse-graining map π of Section II.C, is never used to compute any Table V or Table VI entry; the paper itself disclaims precision there: 'The placement of measures in Table V therefore serves as a conceptual guide rather than a precise coordinate system' (Section IV.A). That makes the classifications judgment-dependent, which is a reproducibility limitation but not circularity: the judgments do not reduce to the framework, and the framework does not reduce to the judgments. The author's prior papers (refs. 36, 37) are cited only as examples of neural-data-analysis applications in the introduction; they are not load-bearing for any central premise, uniqueness claim, or approximation argument. The modern-ML-as-approximation discussion is analogical and explicitly hedged (e.g., the Table VII footnote 'we speculate about this possible interpretation'), not a derivation from self-cited results. No 'prediction' is fitted from data, and no known result is renamed as a new derivation. Accordingly, no circularity is present.
Assumptions & free parameters
assumptions (3)
- domain assumption Complexity measures require a clearly defined coarse-graining map (pi) to distinguish regularity from randomness.
- ad hoc to paper The three axes (regularity, randomness, complexity) are independent and cover the fundamental properties of data.
- domain assumption Kolmogorov complexity is uncomputable; compression algorithms are valid proxies.
Cite this review
Pith. "Pith review of Compression, Regularity, Randomness and Emergent Structure: Rethinking Physical Complexity in the Data-Driven Era." pith.science (2026). https://pith.science/paper/XFBBN7J6
@misc{pith2026250507222,
author = {Pith},
title = {Pith review of: Compression, Regularity, Randomness and Emergent Structure: Rethinking Physical Complexity in the Data-Driven Era},
year = {2026},
howpublished = {\url{https://pith.science/paper/XFBBN7J6}},
note = {Machine review of arXiv:2505.07222}
}
read the original abstract
Complexity science offers a wide range of measures for quantifying unpredictability, structure, and information. Yet, a systematic conceptual organization of these measures is still missing. We present a unified framework that locates statistical, algorithmic, and dynamical measures along three axes (regularity, randomness, and complexity) and situates them in a common conceptual space. We map statistical, algorithmic, and dynamical measures into this conceptual space, discussing their computational accessibility and approximability. This taxonomy reveals the deep challenges posed by uncomputability and highlights the emergence of modern data-driven methods (including autoencoders, latent dynamical models, symbolic regression, and physics-informed neural networks) as pragmatic approximations to classical complexity ideals. Latent spaces emerge as operational arenas where regularity extraction, noise management, and structured compression converge, bridging theoretical foundations with practical modeling in high-dimensional systems. We close by outlining implications for physics-informed AI and AI-guided discovery in complex physical systems, arguing that classical questions of complexity remain central to next-generation scientific modeling.
Reference graph
Works this paper leans on
-
[1]
Overview Algorithmic complexity measures focus not on empirical frequencies or probability distributions but rather on the min- imal description length required to generate an object or dataset51,103. Rooted in algorithmic information theory , these measures seek to capture the intrinsic, deep structure within data, explicitly addressing (Table II): • The...
-
[2]
Detailed Measure Summaries a. Shannon Entropy Shannon entropy, as H(X) =−∑ i p(xi)log p(xi) quantifies the expected information content (or “surprise”) of an outcome drawn from a distribution 110,116. Shannon en- tropy is the foundational measure of statistical randomness. High value implies unpredictability, whereas low entropy in- dicates predictability...
-
[3]
Detailed Measure Summaries a. Kolmogorov Complexity Kolmogorov complexity K(x) of a string x is the length of the shortest program that outputs x on a universal Turing machine 3,72. It provides an absolute, model-free measure of an object’s intrinsic informa- tional content: • Highly regular data exhibit low K(x). • Truly random data approach K(x) close t...
-
[4]
Randomness, Structure, and Causality
Overview Dynamical information measures extend the analysis of complexity to explicitly account for temporal dependencies, causal relationships, and information flow over time. They address not only what information is present but also how it is processed, stored, and transmitted between components of a system. These measures (Table III): • Capture random...
-
[5]
Detailed Measure Summaries a. Entropy Rate Entropy rate hµ quantifies the average uncertainty per symbol conditioned on the entire past31,116: hµ = lim n→∞ 1 nH(X1,X2, . . . ,Xn) • A low entropy rate hµ implies strong regularity and memory. • A high entropy rate hµ implies weak temporal structure or high randomness. Entropy rate is fundamental in ergodic ...
-
[6]
Regularity–Randomness–Complexity triangle orga- nizes measures based on their emphasis
-
[7]
Depth-×-Accessibility plane highlights estimation dif- ficulty
-
[8]
Computability flow chart classifies exact, approxi- mate, and uncomputable regimes. Together these maps let researchers choose a metric that fits both their scientific question and their computational budget. Having established this foundation, we now turn to the rela- tionship between these classical complexity notions and the emerging field of data-driv...
Show all 19 references
-
[9]
Autoencoders perform explicit compression: • Encoder: x7→ z maps an input x to a lower-dimensional latent vector z
Compression and Structure Extraction via Autoencoders Autoencoders perform explicit compression by mapping observations into a lower-dimensional latent code and recon- structing them, capturing essential regularities while discard- ing noise54,124. Autoencoders perform explici...
-
[10]
Dynamics in Latent Spaces: Toward Predictive Regularity Modern latent compression extends beyond static data to modeling temporal sequences, enabling compact representa- tions of evolving structures. • Latent Ordinary Differential Equations (Latent ODEs)22,105 learn continuous...
-
[11]
In this context: • Latent variables become discoverable symbolic co- ordinates; recovered equations provide interpretable structure
Symbolic Latent Representations: Unfolding Logical Depth Symbolic-regression methods such as Eureqa106 and SINDy14,20 extract explicit algebraic rules for latent dynam- ics. In this context: • Latent variables become discoverable symbolic co- ordinates; recovered equations pro...
-
[12]
By embedding structural priors into training, these mod- els guide latent representations towardmeaningful regularities rather than mere observed pattern replication
Physics-Informed Latent Learning: Anchoring Regularity Physics-Informed Neural Networks (PINNs) and Neu- ral Operators such as DeepONet embed governing equations or conservation laws into the training loss, steering the latent space toward physically consistent manifolds60,78,...
-
[13]
Limitations of Neural Proxies While Autoencoders, V AEs, latent ODEs, and PINNs of- fer operational surrogates for uncomputable complexity mea- sures, they each introduce specific inductive biases and failure modes that affect their reliability. a. VAEs Variational autoencoder...
-
[14]
Measure Pluralism and Domain Alignment
Decision Guide for Proxy Selection a. Measure Pluralism and Domain Alignment. Com- plexity takes different forms across domains—from thermo- dynamic assembly in biophysics to algorithmic depth in com- putation—making it impossible for a single measure to cap- ture all aspects....
-
[15]
predictive
The Latent Space as an Operational Arena for Complexity Negotiation Latent spaces are more than mathematical conve- niences—they serve as operational arenas where learning systems structure data through three interconnected axes of complexity. • Regularity extraction : How muc...
1947
-
[562]
100Danilo Jimenez Rezende and Shakir Mohamed
University of California Press, 1961. 100Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on Machine Learning (ICML) , ICML’15, page 1530–1538. JMLR.org, 2015. 101Joshua S. Richman and ...
1961
-
[1988]
ran- domness, structure, and causality: Measures of complexity from theory to applications
Oxford University Press, Inc. 11Charles H. Bennett. Universal computation and physical dynamics. Phys- ica D: Nonlinear Phenomena , 86(1):268–273, 1995. Chaos, Order and Patterns: Aspects of Nonlinearity - @’The Gran Finale@’. 12William Bialek, Ilya Nemenman, and Naftali Tishb...
1995
-
[2000]
PMID: 10843903. 102M. Riedl, A. Müller, and N. Wessel. Practical considerations of permuta- tion entropy. The European Physical Journal Special Topics, 222(2):249– 262, Jun 2013. 103J. Rissanen. Modeling by shortest data description. Automatica, 14(5):465–471, 1978. 104Fernand...
2013
-
[2014]
78Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis
New Horizons for Neural Oscillations. 78Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via deeponet based on the uni- versal approximation theorem of operators. Nature Machine Intelligence, 3(3):218–229, Mar 2021. 79...
2021
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.