{"id":"53c34b19-e97f-4be5-97c0-9456104a4583","arxiv_id":"2505.07222","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A taxonomy places statistical, algorithmic, and dynamical complexity measures along three axes and frames machine-learning latent spaces as practical approximations of uncomputable complexity ideals.","lead":"This paper organizes dozens of complexity measures into a three-axis conceptual space (regularity, randomness, complexity) and links them to modern machine learning methods. It provides a map for choosing and interpreting complexity metrics in data-driven science.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table V–VI axis placements rely on unstated judgment; the framework's central claim needs a reproducible placement rule to be systematic.","rationale":"The reader's weakest assumption is that the three axes are sufficient and that placements in Tables V and VI are objective and reproducible. My stress-test identifies the same load-bearing concern: the paper never gives a formal rule connecting a measure's definition to its axis placement, so the taxonomy's systematicity is untested. I considered whether this is merely a style issue for a review paper, but the abstract and conclusion assert a unified framework with explanatory and selection value, so the reproducibility of the classification is part of the central claim. I also noted a supporting internal tension: Section V.A calls the axes 'orthogonal, but inter-dependent', while Section II.C defines complexity as an intermediate balance of regularity and randomness; without formal definitions, 'orthogonal axes' is unsupported and reinforces the need for a reproducible placement rule. The proposed inter-rater study would settle the concern directly: high agreement would show the taxonomy is robust despite its informality, while low agreement would require either a more formal framework or a more modest claim. No machine-checked proof or reproducible code is claimed, so this qualitative framework should be judged on the clarity and reproducibility of its organizing scheme. The paper is otherwise carefully written, with reasonable measure summaries and an interesting bridge to latent-space methods, but the central organizing step is the least secure part. I therefore keep the reader's CONDITIONAL verdict unchanged rather than moving to ACCEPT or REJECT, because the issue is testable and may be repairable with a formal placement rule or explicit hedging.","tokens_in":22392,"tokens_out":6760,"duration_ms":76111,"concrete_test":"Conduct a pre-registered inter-rater study: give five or more researchers familiar with complexity measures the Section II axis definitions, the measure descriptions in Section III, and Table IV's axis meanings, but withhold Tables V and VI. Ask each rater to classify every measure as capturing regularity, randomness, or complexity (yes/no/weak/approx.) and to state the coarse-graining map π they would assume for each measure. Compute Fleiss' kappa and record disagreements per measure. If kappa is below 0.6, or if raters cannot specify a unique coarse-graining for a majority of measures, the table placements are analyst-dependent and the paper needs either a formal placement rule or a softened central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that three axes—regularity, randomness, and complexity—provide a unified conceptual space that clarifies the operational meaning of complexity measures. For this to hold, a measure's placement in Tables V and VI must be derivable from the paper's definitions rather than from the author's qualitative judgment. The only formal apparatus introduced, the coarse-graining map π in Section II.C, is never used to derive any entry in Tables V or VI. The checkmarks are qualitative: 'weak', 'some', and 'approx.' have no thresholds, and no decision rule separates 'captures randomness' from 'does not capture randomness'. The paper itself says the map is a 'conceptual guide rather than a precise coordinate system' (Section IV.A), but the abstract and conclusion claim a unified framework, so the burden is to show the guide is not arbitrary. A different analyst could plausibly place measures differently: for example, Statistical Complexity is defined in Section III.B.d as the Shannon entropy of the distribution over causal states, and one could classify it as partially capturing randomness, whereas Table V marks it as capturing regularity and complexity only. Because the paper's payoff is selecting and interpreting measures, an unreproducible classification would undermine the central claim even if each individual measure description is accurate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a conceptual taxonomy for complexity measures, organizing statistical, algorithmic, and dynamical measures along three axes—regularity, randomness, and complexity—and situating them in a common conceptual space. It reviews the mathematical definitions of these measures, places them in two classification tables, discusses their computational accessibility and approximability, and argues that modern deep-learning methods (autoencoders, latent ODEs, symbolic regression, and physics-informed neural networks) act as pragmatic approximations to classical uncomputable complexity ideals. The paper closes with a proxy-selection guide and a research outlook for linking complexity theory with data-driven discovery.","tokens_in":22627,"tokens_out":4632,"duration_ms":49505,"significance":"If the taxonomy were grounded in a reproducible, rule-based classification, it would be a genuinely useful synthesis: it would help practitioners choose among the many available complexity measures and clarify what each measure does and does not quantify. The manuscript is a broad, mostly accurate survey with a compression-centric framing that connects classical information/algorithmic theory to modern machine learning, and it candidly discusses the limitations of neural proxies. Its main value is organizational rather than novel mathematical contribution. The central weakness is that the axis placements in Tables V and VI are asserted through qualitative judgment rather than derived from the definitions; this directly affects the paper's central claim of providing a 'unified framework' and 'systematic conceptual organization.'","major_comments":[{"comment":"The central claim of a systematic organization of measures is not supported by a reproducible rule for assigning axis emphases. The paper states in §IV.A that the placement in Table V 'serves as a conceptual guide rather than a precise coordinate system,' but the abstract and conclusion claim a unified framework; these two statements are in tension. The coarse-graining map π introduced in §II.C is never used to derive any entry in Tables V or VI, and the symbols 'weak,' 'some,' and 'approx.' have no thresholds or decision procedure. Without an operational criterion for what it means for a measure to 'capture randomness' or 'capture regularity,' the tables are not falsifiable, and a different analyst could plausibly assign different placements.","section":"§IV.A; Tables V and VI"},{"comment":"The placement of Statistical Complexity as capturing regularity and complexity but not randomness is not derivable from its definition. The definition in §III.B.d is Cµ = H[S], the Shannon entropy of the distribution over causal states; entropy of a state distribution generally responds to the number of states and their probabilities, which is an unpredictability-related quantity. The table's '–' in the Randomness column is therefore a qualitative judgment rather than a consequence of the formal definition. This example illustrates why the placements in Tables V and VI need either a formal decision rule or an explicit repositioning of the paper's claims.","section":"Table V, Statistical Complexity row; §III.B.d"},{"comment":"The prose and the table contradict each other. §III.A.1 states that statistical entropy measures 'weakly capture regularity' and 'capture complexity only in limited ways,' but Table I lists Shannon, Rényi, and Tsallis entropies as 'No' for both Regularity and Complexity. Since the tables are the paper's main deliverable, such internal inconsistencies in the central taxonomy need to be resolved before the framework can be considered reliable.","section":"§III.A.1 vs. Table I"}],"minor_comments":[{"comment":"The abstract calls the axes 'orthogonal,' while §IV.A describes them as 'three distinct yet intertwined' properties and §V.A says 'orthogonal, but inter-dependent.' Please align the terminology, since orthogonality is never formally defined in the paper.","section":"Abstract and §IV.A"},{"comment":"There is a duplicated word in the sentence 'a multiscale extension called, called Multiscale Sample Entropy'; please correct it.","section":"§III.A.e"},{"comment":"The notation for logical depth is inconsistent: both k(x) and K(x) appear for Kolmogorov complexity, and the program length is written as both |P| and |p| in the same passage. Please unify the notation.","section":"§III.B.b"},{"comment":"Reference [3] is incomplete: the entry reads 'A. N. Kolmogorov and. Three approaches to the quantitative definition of information' with a missing co-author name.","section":"References"},{"comment":"The equations in Sections III.A, III.B, and III.C are not numbered, which makes it unnecessarily difficult to refer to specific definitions in the comparative discussion.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is a position/review paper whose main contribution is the taxonomic organization. The paper is broad and mostly accurate, but the central classification is not yet reproducible; the author's own statement in §IV.A that the placement is 'a conceptual guide' partially concedes this. The weakness is fixable within the manuscript's scope by either adding a formal placement rule or reframing the contribution as an explicitly qualitative guide. The paper also contains a few internal inconsistencies between prose and tables that should be resolved in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a review, not a new-result paper, and it's best read as a well-organized survey with an attractive organizing scheme. The three-axis framing—regularity, randomness, complexity—is a reasonable way to compare classical complexity measures, and the connections to autoencoders, latent ODEs, and PINNs are apt. The paper does a solid job summarizing standard definitions, and the formulas are mostly correct.\n\nThe genuine contribution is the attempt to place all these measures in a common conceptual space and add a depth-accessibility plane. That's a useful teaching device and could help practitioners choose measures. The discussion of uncomputability and the need for pragmatic proxies is well-taken.\n\nThe soft spot is the central claim. The placements in Tables V and VI are qualitative judgments; the paper itself says Table V is a 'conceptual guide rather than a precise coordinate system' (Section IV.A), which is honest but undercuts the abstract's stronger language. The coarse-graining map π in Section II.C is never used to derive any entry. A different analyst could reasonably place statistical complexity—the entropy over causal states—as partially capturing randomness, whereas Table V marks it as regularity and complexity only. The paper also leans on earlier work (Crutchfield and Feldman 2003, Ay et al. 2006, Prokopenko et al. 2009) that already makes the non-duality point; the novelty is more in the layout than in the underlying concepts.\n\nMinor issues: the conclusion says 'three orthogonal, but inter-dependent, axes,' which is at least in tension; and Table VII includes an explicitly speculative row about attention mechanisms. These are minor, not load-bearing.\n\nWho this is for: readers new to complexity measures who want a map of the landscape, and ML researchers looking for a conceptual link between classical complexity and latent-space methods. It doesn't resolve open questions, and it doesn't give a decision rule for when one measure beats another. But as a survey it's competent and mostly accurate.\n\nI'd send it to peer review rather than desk-reject. It's a useful review that could be strengthened by a reproducible placement rule or at least an explicit acknowledgment that the placements are a proposed reading, not a derivation. If the venue wants rigorous new results this is borderline; if it publishes surveys, it's a reasonable fit.","headline":"A plausible but under-supported taxonomy of complexity measures—useful as a conceptual map, but the central 'unified framework' claim rests on subjective table entries, so treat it as a well-organized survey rather than a systematic classification.","tokens_in":23092,"tokens_out":2622,"would_cite":false,"duration_ms":23837,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that every complexity measure can be placed on three axes—regularity, randomness, and complexity—and that this map reveals why the most structural measures resist computation.","keywords":["complexity measures","regularity and randomness","Kolmogorov complexity","entropy","latent space models","physics-informed neural networks","information dynamics","compression"],"falsifier":"Recruit, say, ten researchers familiar with these measures and have them independently fill in Tables V and VI from the definitions alone, then compute inter-rater agreement such as Cohen's kappa; low agreement would falsify the claim that the taxonomy is a systematic organization rather than one author's perspective. A quantitative supplement would be to verify axis assignments on synthetic systems with known regularity and randomness—for instance, periodic, chaotic, and white-noise sequences—and check whether each measure behaves as its table placement predicts.","tokens_in":22199,"feed_emoji":"🧩","tokens_out":7962,"duration_ms":76631,"temperature":0.7,"pith_summary":"This paper attempts to organize the sprawling field of complexity measurement. It claims that existing statistical, algorithmic, and dynamical measures can be placed in a common conceptual space defined by three axes: regularity (predictable, compressible structure), randomness (unpredictable, incompressible noise), and complexity (non-trivial structure between the two). The placement reveals a systematic trade-off: measures that richly capture structure, such as Kolmogorov complexity, logical depth, and effective complexity, are uncomputable or practically inaccessible, while accessible entropy measures mostly capture randomness. The paper then argues that modern data-driven methods—autoencoders, latent dynamical models, symbolic regression, and physics-informed networks—work as pragmatic approximations of these classical ideals, with latent spaces as the arena where compression, regularity extraction, and noise management meet. A sympathetic reader would care because this taxonomy offers a principled way to choose and interpret complexity measures in empirical science.","feed_headline":"Three axes explain why complexity measures disagree","feed_subtitle":"A new map separates regularity from randomness and shows why the deepest measures resist computation.","key_machinery":"The carrying object is the three-axis conceptual landscape—regularity, randomness, and complexity—together with the depth–accessibility plane that classifies each measure by theoretical richness versus ease of estimation. The formal anchor is the compression view: a sequence is regular when its Kolmogorov complexity $K(x)$ is much shorter than the sequence, random when $K(x)$ approaches its length, and complex when it is compressible only through non-trivial computational effort. The paper adopts a coarse-graining map $\\pi:\\Omega_{\\text{micro}}\\to\\Omega_{\\text{macro}}$ as the prerequisite for separating regularities $R(x)$ from noise, and reads latent-space models—encoder-decoder networks, latent ODEs, Koopman autoencoders, symbolic regression, and physics-informed networks—as operational implementations of these same compression and rule-discovery steps.","core_discovery":"The paper's central discovery, on its own terms, is that complexity measures are not interchangeable tools but probes with distinct sensitivities that can be systematically charted. Statistical entropies (Shannon, Rényi, Tsallis, approximate entropy, sample entropy, permutation entropy) mostly register randomness, with at most weak purchase on regularity. Algorithmic measures (Kolmogorov complexity, effective complexity, logical depth, sophistication, statistical complexity) capture regularity directly and define complexity as the boundary between structure and noise, but at the price of uncomputability or severe estimation difficulty. Dynamical measures (entropy rate, transfer entropy, active information storage, information modification) straddle the axes by tracking how information is stored, transferred, and transformed in time. The paper claims that placing these measures on the regularity–randomness–complexity triangle and on a depth–accessibility plane clarifies why two metrics can disagree on the same dataset, and that machine-learned latent representations operationalize the same trade-offs.","pith_inferences":["The axis placements in Tables V and VI are qualitative; a natural next step, not taken here, is to turn them into a rating instrument and measure inter-rater agreement, which would test whether the taxonomy is objective or one analyst's interpretation.","The same three axes could organize model-selection criteria such as AIC, BIC, and MDL, connecting the taxonomy to mainstream statistical learning theory.","The latent-space claim suggests a concrete diagnostic: measure the compression ratio of learned latent codes (for example, gzip size of $z$) across datasets and models; if strong reconstruction consistently coincides with low latent complexity, the 'operational arena' story gains empirical support.","The decision guide in Table VIII can be read as falsifiable predictions—VAEs underestimate heavy-tailed structure and compression proxies miss dynamical correlations—that benchmark tests on the logistic map and Lorenz attractor could confirm or refute."],"forward_implications":["Measure selection becomes a deliberate act: a researcher studying noise should reach for entropy-type measures, one studying hidden structure needs algorithmic or dynamical measures, and one studying the order-disorder boundary needs a complexity measure plus a stated coarse-graining.","The practical question for foundational measures shifts from 'is it computable?' to 'how well and with what bias can it be approximated?'","Deep learning architectures acquire a principled interpretation: autoencoders approximate minimal description length, symbolic regression approximates effective complexity, and physics-informed networks anchor regularity extraction.","The depth–accessibility trade-off predicts specific proxy failure modes—VAEs oversmooth rare structure, latent ODEs oversmooth sharp dynamics, and PINNs can enforce wrong priors—making synthetic benchmarking the recommended validation.","Future learning systems could treat compressibility and regularity extraction as explicit objectives, tying training dynamics to the classical complexity axes."],"supporting_citations":[{"why":"Supplies the wide catalogue of complexity measures that the paper reorganizes.","marker":"[75]"},{"why":"Establishes that regularity and randomness are not opposites and that entropy convergence reveals hidden structure.","marker":"[31]"},{"why":"Defines statistical complexity via causal states, the prediction-focused measure the taxonomy places under complexity.","marker":"[34]"},{"why":"Defines effective complexity as the description length of structured regularities, separating structure from noise.","marker":"[44]"},{"why":"Defines logical depth as the computational effort needed to unfold a compressed description.","marker":"[10]"},{"why":"Provides the coarse-graining map and thermodynamic depth that the paper uses to separate regularities from random fluctuations.","marker":"[76]"},{"why":"Supplies the foundational Shannon entropy that anchors the randomness axis.","marker":"[110]"},{"why":"Offers variational autoencoders as the modern latent-compression approximation to minimal description length.","marker":"[62]"},{"why":"Provides physics-informed neural networks as the regularity-anchored proxy for structure discovery.","marker":"[98]"}],"fun_headline_variants":["One chart positions every complexity measure","Uncomputability meets machine-learned proxies","Regularity and randomness triangulated","Latent spaces operationalize classical complexity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole framework rests on the assumption that each measure has a definite place on the three axes, but those placements are the author's qualitative judgment, so another analyst might place them differently; if placements are not reproducible, the claimed organization is not systematic.","fun_headline_variants_meta":{"raw":{"variants":["One chart positions every complexity measure","Uncomputability meets machine-learned proxies","Regularity and randomness triangulated","Latent spaces operationalize classical complexity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000346,"raw_usage":{"total_tokens":1881,"prompt_tokens":914,"completion_tokens":967,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":917}},"tokens_in":530,"tokens_out":967,"duration_ms":10281,"temperature":1.0,"reasoning_tokens":917,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:20:57.585181+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit, say, ten researchers familiar with these measures and have them independently fill in Tables V and VI from the definitions alone, then compute inter-rater agreement such as Cohen's kappa; low agreement would falsify the claim that the taxonomy is a systematic organization rather than one author's perspective. A quantitative supplement would be to verify axis assignments on synthetic systems with known regularity and randomness—for instance, periodic, chaotic, and white-noise sequences—and check whether each measure behaves as its table placement predicts.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines logical depth as the computational effort needed to unfold a compressed description."}],"review_version":1}