Pith. sign in

REVIEW 3 major objections 5 minor 19 references

Compression, Regularity, Randomness and Emergent Structure: Rethinking Physical Complexity in the Data-Driven Era

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper claims that every complexity measure can be placed on three axes—regularity, randomness, and complexity—and that this map reveals why the most structural measures resist computation.

desk verdict A plausible but under-supported taxonomy of complexity measures—useful as a conceptual map, but the central 'unified framework' claim rests on subjective table entries, so treat it as a well-organized survey rather than a systematic classification. read the letter →

arxiv 2505.07222 v1 pith:XFBBN7J6 submitted 2025-05-12 cs.LG cond-mat.stat-mechcs.ITmath.ITphysics.bio-phphysics.data-an

classification cs.LGcond-mat.stat-mechcs.ITmath.ITphysics.bio-phphysics.data-an
keywords complexitymeasuresregularityandrandomnessKolmogoroventropylatentspacemodelsphysics-informedneuralnetworksinformationdynamicscompression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to organize the sprawling field of complexity measurement. It claims that existing statistical, algorithmic, and dynamical measures can be placed in a common conceptual space defined by three axes: regularity (predictable, compressible structure), randomness (unpredictable, incompressible noise), and complexity (non-trivial structure between the two). The placement reveals a systematic trade-off: measures that richly capture structure, such as Kolmogorov complexity, logical depth, and effective complexity, are uncomputable or practically inaccessible, while accessible entropy measures mostly capture randomness. The paper then argues that modern data-driven methods—autoencoders, latent dynamical models, symbolic regression, and physics-informed networks—work as pragmatic approximations of these classical ideals, with latent spaces as the arena where compression, regularity extraction, and noise management meet. A sympathetic reader would care because this taxonomy offers a principled way to choose and interpret complexity measures in empirical science.

What carries the argument

The carrying object is the three-axis conceptual landscape—regularity, randomness, and complexity—together with the depth–accessibility plane that classifies each measure by theoretical richness versus ease of estimation. The formal anchor is the compression view: a sequence is regular when its Kolmogorov complexity $K(x)$ is much shorter than the sequence, random when $K(x)$ approaches its length, and complex when it is compressible only through non-trivial computational effort. The paper adopts a coarse-graining map $\pi:\Omega_{\text{micro}}\to\Omega_{\text{macro}}$ as the prerequisite for separating regularities $R(x)$ from noise, and reads latent-space models—encoder-decoder networks, latent ODEs, Koopman autoencoders, symbolic regression, and physics-informed networks—as operational implementations of these same compression and rule-discovery steps.

What would settle it

Recruit, say, ten researchers familiar with these measures and have them independently fill in Tables V and VI from the definitions alone, then compute inter-rater agreement such as Cohen's kappa; low agreement would falsify the claim that the taxonomy is a systematic organization rather than one author's perspective. A quantitative supplement would be to verify axis assignments on synthetic systems with known regularity and randomness—for instance, periodic, chaotic, and white-noise sequences—and check whether each measure behaves as its table placement predicts.

Watch

Extended reading notes

Core claim

The paper's central discovery, on its own terms, is that complexity measures are not interchangeable tools but probes with distinct sensitivities that can be systematically charted. Statistical entropies (Shannon, Rényi, Tsallis, approximate entropy, sample entropy, permutation entropy) mostly register randomness, with at most weak purchase on regularity. Algorithmic measures (Kolmogorov complexity, effective complexity, logical depth, sophistication, statistical complexity) capture regularity directly and define complexity as the boundary between structure and noise, but at the price of uncomputability or severe estimation difficulty. Dynamical measures (entropy rate, transfer entropy, active information storage, information modification) straddle the axes by tracking how information is stored, transferred, and transformed in time. The paper claims that placing these measures on the regularity–randomness–complexity triangle and on a depth–accessibility plane clarifies why two metrics can disagree on the same dataset, and that machine-learned latent representations operationalize the same trade-offs.

Load-bearing premise

The whole framework rests on the assumption that each measure has a definite place on the three axes, but those placements are the author's qualitative judgment, so another analyst might place them differently; if placements are not reproducible, the claimed organization is not systematic.

Editorial extensions

If this is right

  • Measure selection becomes a deliberate act: a researcher studying noise should reach for entropy-type measures, one studying hidden structure needs algorithmic or dynamical measures, and one studying the order-disorder boundary needs a complexity measure plus a stated coarse-graining.
  • The practical question for foundational measures shifts from 'is it computable?' to 'how well and with what bias can it be approximated?'
  • Deep learning architectures acquire a principled interpretation: autoencoders approximate minimal description length, symbolic regression approximates effective complexity, and physics-informed networks anchor regularity extraction.
  • The depth–accessibility trade-off predicts specific proxy failure modes—VAEs oversmooth rare structure, latent ODEs oversmooth sharp dynamics, and PINNs can enforce wrong priors—making synthetic benchmarking the recommended validation.
  • Future learning systems could treat compressibility and regularity extraction as explicit objectives, tying training dynamics to the classical complexity axes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The axis placements in Tables V and VI are qualitative; a natural next step, not taken here, is to turn them into a rating instrument and measure inter-rater agreement, which would test whether the taxonomy is objective or one analyst's interpretation.
  • The same three axes could organize model-selection criteria such as AIC, BIC, and MDL, connecting the taxonomy to mainstream statistical learning theory.
  • The latent-space claim suggests a concrete diagnostic: measure the compression ratio of learned latent codes (for example, gzip size of $z$) across datasets and models; if strong reconstruction consistently coincides with low latent complexity, the 'operational arena' story gains empirical support.
  • The decision guide in Table VIII can be read as falsifiable predictions—VAEs underestimate heavy-tailed structure and compression proxies miss dynamical correlations—that benchmark tests on the logistic map and Lorenz attractor could confirm or refute.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper proposes a conceptual taxonomy for complexity measures, organizing statistical, algorithmic, and dynamical measures along three axes—regularity, randomness, and complexity—and situating them in a common conceptual space. It reviews the mathematical definitions of these measures, places them in two classification tables, discusses their computational accessibility and approximability, and argues that modern deep-learning methods (autoencoders, latent ODEs, symbolic regression, and physics-informed neural networks) act as pragmatic approximations to classical uncomputable complexity ideals. The paper closes with a proxy-selection guide and a research outlook for linking complexity theory with data-driven discovery.

Significance. If the taxonomy were grounded in a reproducible, rule-based classification, it would be a genuinely useful synthesis: it would help practitioners choose among the many available complexity measures and clarify what each measure does and does not quantify. The manuscript is a broad, mostly accurate survey with a compression-centric framing that connects classical information/algorithmic theory to modern machine learning, and it candidly discusses the limitations of neural proxies. Its main value is organizational rather than novel mathematical contribution. The central weakness is that the axis placements in Tables V and VI are asserted through qualitative judgment rather than derived from the definitions; this directly affects the paper's central claim of providing a 'unified framework' and 'systematic conceptual organization.'

major comments (3)
  1. [§IV.A; Tables V and VI] The central claim of a systematic organization of measures is not supported by a reproducible rule for assigning axis emphases. The paper states in §IV.A that the placement in Table V 'serves as a conceptual guide rather than a precise coordinate system,' but the abstract and conclusion claim a unified framework; these two statements are in tension. The coarse-graining map π introduced in §II.C is never used to derive any entry in Tables V or VI, and the symbols 'weak,' 'some,' and 'approx.' have no thresholds or decision procedure. Without an operational criterion for what it means for a measure to 'capture randomness' or 'capture regularity,' the tables are not falsifiable, and a different analyst could plausibly assign different placements.
  2. [Table V, Statistical Complexity row; §III.B.d] The placement of Statistical Complexity as capturing regularity and complexity but not randomness is not derivable from its definition. The definition in §III.B.d is Cµ = H[S], the Shannon entropy of the distribution over causal states; entropy of a state distribution generally responds to the number of states and their probabilities, which is an unpredictability-related quantity. The table's '–' in the Randomness column is therefore a qualitative judgment rather than a consequence of the formal definition. This example illustrates why the placements in Tables V and VI need either a formal decision rule or an explicit repositioning of the paper's claims.
  3. [§III.A.1 vs. Table I] The prose and the table contradict each other. §III.A.1 states that statistical entropy measures 'weakly capture regularity' and 'capture complexity only in limited ways,' but Table I lists Shannon, Rényi, and Tsallis entropies as 'No' for both Regularity and Complexity. Since the tables are the paper's main deliverable, such internal inconsistencies in the central taxonomy need to be resolved before the framework can be considered reliable.
minor comments (5)
  1. [Abstract and §IV.A] The abstract calls the axes 'orthogonal,' while §IV.A describes them as 'three distinct yet intertwined' properties and §V.A says 'orthogonal, but inter-dependent.' Please align the terminology, since orthogonality is never formally defined in the paper.
  2. [§III.A.e] There is a duplicated word in the sentence 'a multiscale extension called, called Multiscale Sample Entropy'; please correct it.
  3. [§III.B.b] The notation for logical depth is inconsistent: both k(x) and K(x) appear for Kolmogorov complexity, and the program length is written as both |P| and |p| in the same passage. Please unify the notation.
  4. [References] Reference [3] is incomplete: the entry reads 'A. N. Kolmogorov and. Three approaches to the quantitative definition of information' with a missing co-author name.
  5. [Throughout] The equations in Sections III.A, III.B, and III.C are not numbered, which makes it unnecessarily difficult to refer to specific definitions in the comparative discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a conceptual taxonomy whose classifications are qualitative and explicitly non-coordinate; no derivation reduces to its inputs.

full rationale

This paper is a review and conceptual taxonomy, not a derivation. It defines three axes (regularity, randomness, complexity) and then classifies existing measures by qualitative judgment. I found no step in which a result is derived from, or predicted by, a fitted parameter, a self-cited theorem, or a definition that secretly contains the conclusion. The only formal machinery introduced, the coarse-graining map π of Section II.C, is never used to compute any Table V or Table VI entry; the paper itself disclaims precision there: 'The placement of measures in Table V therefore serves as a conceptual guide rather than a precise coordinate system' (Section IV.A). That makes the classifications judgment-dependent, which is a reproducibility limitation but not circularity: the judgments do not reduce to the framework, and the framework does not reduce to the judgments. The author's prior papers (refs. 36, 37) are cited only as examples of neural-data-analysis applications in the introduction; they are not load-bearing for any central premise, uniqueness claim, or approximation argument. The modern-ML-as-approximation discussion is analogical and explicitly hedged (e.g., the Table VII footnote 'we speculate about this possible interpretation'), not a derivation from self-cited results. No 'prediction' is fitted from data, and no known result is renamed as a new derivation. Accordingly, no circularity is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's framework introduces no formally defined new quantities. It relies on standard definitions from the literature, the assumption that three axes are sufficient, and the conventional partition of measures into three families.

assumptions (3)
  • domain assumption Complexity measures require a clearly defined coarse-graining map (pi) to distinguish regularity from randomness.
    Invoked in Section II.C following Lloyd and Pagels; the paper's framework assumes the user posits a coarse-graining.
  • ad hoc to paper The three axes (regularity, randomness, complexity) are independent and cover the fundamental properties of data.
    This is the paper's organizing premise, asserted in the introduction and Section IV without proof or external validation.
  • domain assumption Kolmogorov complexity is uncomputable; compression algorithms are valid proxies.
    Used throughout Section III.B and IV.C; this is a standard result and a common practical assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Compression, Regularity, Randomness and Emergent Structure: Rethinking Physical Complexity in the Data-Driven Era." pith.science (2026). https://pith.science/paper/XFBBN7J6

@misc{pith2026250507222,
  author       = {Pith},
  title        = {Pith review of: Compression, Regularity, Randomness and Emergent Structure: Rethinking Physical Complexity in the Data-Driven Era},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XFBBN7J6}},
  note         = {Machine review of arXiv:2505.07222}
}
read the original abstract

Complexity science offers a wide range of measures for quantifying unpredictability, structure, and information. Yet, a systematic conceptual organization of these measures is still missing. We present a unified framework that locates statistical, algorithmic, and dynamical measures along three axes (regularity, randomness, and complexity) and situates them in a common conceptual space. We map statistical, algorithmic, and dynamical measures into this conceptual space, discussing their computational accessibility and approximability. This taxonomy reveals the deep challenges posed by uncomputability and highlights the emergence of modern data-driven methods (including autoencoders, latent dynamical models, symbolic regression, and physics-informed neural networks) as pragmatic approximations to classical complexity ideals. Latent spaces emerge as operational arenas where regularity extraction, noise management, and structured compression converge, bridging theoretical foundations with practical modeling in high-dimensional systems. We close by outlining implications for physics-informed AI and AI-guided discovery in complex physical systems, arguing that classical questions of complexity remain central to next-generation scientific modeling.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 19 canonical work pages

  1. [1]

    Overview Algorithmic complexity measures focus not on empirical frequencies or probability distributions but rather on the min- imal description length required to generate an object or dataset51,103. Rooted in algorithmic information theory , these measures seek to capture the intrinsic, deep structure within data, explicitly addressing (Table II): • The...

  2. [2]

    surprise

    Detailed Measure Summaries a. Shannon Entropy Shannon entropy, as H(X) =−∑ i p(xi)log p(xi) quantifies the expected information content (or “surprise”) of an outcome drawn from a distribution 110,116. Shannon en- tropy is the foundational measure of statistical randomness. High value implies unpredictability, whereas low entropy in- dicates predictability...

  3. [3]

    macrostates

    Detailed Measure Summaries a. Kolmogorov Complexity Kolmogorov complexity K(x) of a string x is the length of the shortest program that outputs x on a universal Turing machine 3,72. It provides an absolute, model-free measure of an object’s intrinsic informa- tional content: • Highly regular data exhibit low K(x). • Truly random data approach K(x) close t...

  4. [4]

    Randomness, Structure, and Causality

    Overview Dynamical information measures extend the analysis of complexity to explicitly account for temporal dependencies, causal relationships, and information flow over time. They address not only what information is present but also how it is processed, stored, and transmitted between components of a system. These measures (Table III): • Capture random...

  5. [5]

    complexity

    Detailed Measure Summaries a. Entropy Rate Entropy rate hµ quantifies the average uncertainty per symbol conditioned on the entire past31,116: hµ = lim n→∞ 1 nH(X1,X2, . . . ,Xn) • A low entropy rate hµ implies strong regularity and memory. • A high entropy rate hµ implies weak temporal structure or high randomness. Entropy rate is fundamental in ergodic ...

  6. [6]

    Regularity–Randomness–Complexity triangle orga- nizes measures based on their emphasis

  7. [7]

    Depth-×-Accessibility plane highlights estimation dif- ficulty

  8. [8]

    Complexity

    Computability flow chart classifies exact, approxi- mate, and uncomputable regimes. Together these maps let researchers choose a metric that fits both their scientific question and their computational budget. Having established this foundation, we now turn to the rela- tionship between these classical complexity notions and the emerging field of data-driv...

Show all 19 references
  1. [9]

    Autoencoders perform explicit compression: • Encoder: x7→ z maps an input x to a lower-dimensional latent vector z

    Compression and Structure Extraction via Autoencoders Autoencoders perform explicit compression by mapping observations into a lower-dimensional latent code and recon- structing them, capturing essential regularities while discard- ing noise54,124. Autoencoders perform explici...

  2. [10]

    Dynamics in Latent Spaces: Toward Predictive Regularity Modern latent compression extends beyond static data to modeling temporal sequences, enabling compact representa- tions of evolving structures. • Latent Ordinary Differential Equations (Latent ODEs)22,105 learn continuous...

  3. [11]

    In this context: • Latent variables become discoverable symbolic co- ordinates; recovered equations provide interpretable structure

    Symbolic Latent Representations: Unfolding Logical Depth Symbolic-regression methods such as Eureqa106 and SINDy14,20 extract explicit algebraic rules for latent dynam- ics. In this context: • Latent variables become discoverable symbolic co- ordinates; recovered equations pro...

  4. [12]

    By embedding structural priors into training, these mod- els guide latent representations towardmeaningful regularities rather than mere observed pattern replication

    Physics-Informed Latent Learning: Anchoring Regularity Physics-Informed Neural Networks (PINNs) and Neu- ral Operators such as DeepONet embed governing equations or conservation laws into the training loss, steering the latent space toward physically consistent manifolds60,78,...

  5. [13]

    Limitations of Neural Proxies While Autoencoders, V AEs, latent ODEs, and PINNs of- fer operational surrogates for uncomputable complexity mea- sures, they each introduce specific inductive biases and failure modes that affect their reliability. a. VAEs Variational autoencoder...

  6. [14]

    Measure Pluralism and Domain Alignment

    Decision Guide for Proxy Selection a. Measure Pluralism and Domain Alignment. Com- plexity takes different forms across domains—from thermo- dynamic assembly in biophysics to algorithmic depth in com- putation—making it impossible for a single measure to cap- ture all aspects....

  7. [15]

    predictive

    The Latent Space as an Operational Arena for Complexity Negotiation Latent spaces are more than mathematical conve- niences—they serve as operational arenas where learning systems structure data through three interconnected axes of complexity. • Regularity extraction : How muc...

  8. [562]

    100Danilo Jimenez Rezende and Shakir Mohamed

    University of California Press, 1961. 100Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on Machine Learning (ICML) , ICML’15, page 1530–1538. JMLR.org, 2015. 101Joshua S. Richman and ...

  9. [1988]

    ran- domness, structure, and causality: Measures of complexity from theory to applications

    Oxford University Press, Inc. 11Charles H. Bennett. Universal computation and physical dynamics. Phys- ica D: Nonlinear Phenomena , 86(1):268–273, 1995. Chaos, Order and Patterns: Aspects of Nonlinearity - @’The Gran Finale@’. 12William Bialek, Ilya Nemenman, and Naftali Tishb...

  10. [2000]

    PMID: 10843903. 102M. Riedl, A. Müller, and N. Wessel. Practical considerations of permuta- tion entropy. The European Physical Journal Special Topics, 222(2):249– 262, Jun 2013. 103J. Rissanen. Modeling by shortest data description. Automatica, 14(5):465–471, 1978. 104Fernand...

  11. [2014]

    78Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis

    New Horizons for Neural Oscillations. 78Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via deeponet based on the uni- versal approximation theorem of operators. Nature Machine Intelligence, 3(3):218–229, Mar 2021. 79...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.