Pith. sign in

REVIEW 4 major objections 4 minor 30 references

Dual-Valued Functions of Dual Matrices with Applications in Causal Emergence

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that extending matrix norms to dual matrices via the Gâteaux derivative yields a dual-valued Ky Fan p-k-norm whose infinitesimal part, maximized over k as p varies in [1,2), identifies the optimal macro-state count at…

desk verdict Solid dual-continuation theory; the causal-emergence k-selection claim is a single-example empirical conjecture that needs more support. read the letter →

arxiv 2411.08377 v1 pith:ILHMQ7QX submitted 2024-11-13 math.NA cs.NA

classification math.NAcs.NA MSC 15A6015B3330G3565C40
keywords dualnumberscontinuationGâteauxderivativedual-valuedmatrixnormsKyFanp-k-normcausalemergencetransitionalprobabilityeffectiveinformation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that real-valued matrix norms can be carried over to dual matrices—matrices whose entries are dual numbers $a + b\epsilon$ with $\epsilon^2=0$—by placing the Gâteaux derivative in the infinitesimal part, and that one resulting norm detects causal emergence. The authors define dual-valued vector and matrix norms, prove that the dual-valued Ky Fan $p$-$k$-norm equals the dual-valued vector $p$-norm of the first $k$ singular values, and introduce a dual transition probability matrix with dual effective information. On a dumbbell Markov chain, the $k$ that maximizes the infinitesimal part of the dual Ky Fan $p$-$k$-norm as $p$ ranges over $[1,2)$ matches the number of macro-states of the system. The payoff, if the claim generalizes, is a way to choose the optimal coarse-graining scale without enumerating all coarse-grainings.

What carries the argument

The load-bearing object is the dual continuation of the Ky Fan $p$-$k$-norm, whose infinitesimal part is obtained by maximizing $\langle G, A_i\rangle$ over $G$ in the subdifferential of the ordinary Ky Fan $p$-$k$-norm at $A_s$; this is what makes the norm sensitive to the infinitesimal structure of the transition matrix. The companion identity $\|A\|_{(k,p)}=\|\sigma^k\|_p$, where $\sigma^k$ is the dual vector of the first $k$ dual singular values, lets the norm be computed from a dual SVD. In the application, the DTPM $P=P_s+P_i\epsilon$ is fitted by solving two least-squares problems, and the $k$ that maximizes the infinitesimal part of $\|P\|_{(k,p)}$ over $p\in[1,2)$ is declared the optimal classification number.

What would settle it

Construct a Markov chain with a known number of metastable groups but no clear gap in its singular values; if the $k$ maximizing the infinitesimal part of $\|P\|_{(k,p)}$ over $p\in[1,2)$ does not equal that known number, or if the peak moves when $p$ is varied within $[1,2)$ or when the least-squares fitting is perturbed, the central claim would be falsified.

Watch

Extended reading notes

Core claim

The central claim is that the dual continuation of a real-valued function, sending $\varphi(a+b\epsilon)$ to $\varphi(a)+D_b\varphi(a)\epsilon$ with $D_b$ the Gâteaux derivative, produces valid dual-valued vector and matrix norms that keep the real-field properties, and that the resulting dual-valued Ky Fan $p$-$k$-norm carries usable information about causal emergence. In particular, for a dual transition probability matrix $P=P_s+P_i\epsilon$ fitted from time-series data, the infinitesimal part of $\|P\|_{(k,p)}$, maximized over $k$ with $p$ in $[1,2)$, identifies the optimal number of macro-states. The paper proves the norm identities, including $\|A\|_{(k,p)}=\|\sigma^k\|_p$ for the vector of the first $k$ dual singular values, and it shows that dual effective information, the dual Schatten $p$-norm, and dynamical reversibility of a DTPM all peak exactly at permutation matrices for $1\le p<2$. The numerical experiment on a dumbbell chain finds the peak at $k=5$, the known number of groups, supporting the claim.

Load-bearing premise

The method assumes that the dual transition matrix fitted from time-series data faithfully reflects the system's causal structure, so that the $k$ maximizing the infinitesimal part of the dual Ky Fan $p$-$k$-norm reliably marks the best coarse-graining scale; this is an empirical heuristic demonstrated on one dumbbell chain, not a proven theorem.

Editorial extensions

If this is right

  • Optimal coarse-graining can be read off directly from the dual Ky Fan $p$-$k$-norm, without enumerating candidate coarse-grainings or locating a singular-value cutoff.
  • The identity $\|A\|_{(k,p)}=\|\sigma^k\|_p$ gives a computable route: take a dual SVD of the DTPM, form the dual vector of its first $k$ singular values, and evaluate its dual vector $p$-norm.
  • For $1\le p<2$, a DTPM reaches its maximum dual Schatten $p$-norm exactly when it is a permutation matrix, and dual effective information reaches its maximum under the same condition, so the norm and effective information agree on the most reversible macro-state.
  • The procedure can be repeated on other time-series data: fit $P_s$ and $P_i$, compute the infinitesimal part of $\|P\|_{(k,p)}$, and take the maximizing $k$ as the system's optimal classification number.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if the maximizing-$k$ heuristic survives on more systems, dual-matrix norms could serve as a general model-order selection tool for Markov chains with slow mixing and no spectral gap, replacing subjective singular-value thresholds.
  • Beyond the paper: the same dual-continuation construction could be applied to other unitarily invariant norms or to complex dual matrices, with the Gâteaux-derivative formalism suggesting analogous identities for condition numbers or entropy-like quantities.
  • Beyond the paper: a direct testable extension is to compare the $k$ suggested by the infinitesimal part with metastability indicators on a family of random dumbbell-like chains with known cluster counts and no singular-value gap.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper develops a general framework for extending real-valued functions on vectors and matrices to dual-valued functions using the Gâteaux derivative, which it calls "dual continuation." It derives explicit formulas for dual-valued vector p-norms and dual-valued unitarily invariant matrix norms, including the Ky Fan p-k-norm, Schatten p-norm, nuclear norm, Frobenius norm, and operator norms, and proves several structural properties such as unitary invariance and consistency. The paper then introduces a dual transitional probability matrix (DTPM) and a dual-valued effective information EId, and studies their extremal properties in relation to dynamical reversibility. In a numerical experiment on a dumbbell Markov chain, the authors observe that the infinitesimal part of the dual-valued Ky Fan p-k-norm attains a maximum at k=5 for p=1.3,1.6,1.9, and they interpret this k as the optimal number of macro-states for causal emergence.

Significance. If the theoretical results are correct, the dual-continuation construction provides a systematic method for extending nonsmooth convex matrix functions to dual matrices, which could be useful in dual-number-based numerical analysis and optimization. The explicit formulas for dual-valued norms and the equivalence in Theorem 5.20 are potentially valuable reference results. The paper's strength is its use of standard convex-analysis tools (Borwein-Lewis, Watson, Nesterov) and the explicit nature of the derived formulas. However, the central applied claim concerning causal emergence rests on a single numerical example and an ad hoc construction of the DTPM, with no theoretical bridge from the Ky Fan p-k-norm maximizer to an optimal classification number. The theoretical gaps described in the major comments are load-bearing for the paper's main claims.

major comments (4)
  1. [§5.4, Eq. (5.2)] The operator-norm formula (5.2) is not well-posed for dual vectors x with zero standard part. In the dual-number algebra, division is defined only when the divisor has a nonzero standard part, yet the maximization in (5.2) is taken over all x ∈ DR^n with x ≠ 0, including x = x_i ε. The proof of Theorem 5.4 handles the As = O case by choosing x_i = 0 and x_s ≠ 0, but this does not justify the unrestricted definition. Either restrict the definition to vectors with nonzero standard part, or define and justify division by infinitesimal dual numbers, and then show the maximum is attained. As stated, the equivalence between the dual continuation and the induced operator norm is not established.
  2. [§5.3, Lemma 5.13 and Prop. 5.8] Lemma 5.13 states that any G ∈ ∂∥X∥_(k,p) has the form (5.9), but it does not establish the converse direction, namely that every symmetric positive semidefinite T with ∥T∥_2 ≤ 1 and ∥T∥_* = t gives an element of the subdifferential. Proposition 5.8 then maximizes ⟨G, A_i⟩ over the set of G of that form; if the converse fails, the maximum over the true subdifferential could be larger, and the formula (5.5) would be invalid. This is load-bearing because (5.5) underlies Corollary 5.16 and the Section 6 applications. Please provide a full characterization of ∂∥X∥_(k,p) (necessity and sufficiency) or cite a reference that contains it.
  3. [§5.3–5.4, Corollaries 5.16–5.18 and Prop. 5.22] The proofs of Corollaries 5.16, 5.17, 5.18, and Proposition 5.22 are omitted with the note "The detailed proofs are omitted." These results are used later: Corollary 5.16 is used in Proposition 6.6 to characterize maxima/minima of the dual-valued Schatten p-norm of a DTPM, and Proposition 5.22 concerns the dual-valued operator ∞-norm. The omissions are not merely cosmetic; they leave unverified the exact infinitesimal-part formulas that drive the causal-emergence analysis. Please include complete proofs or give precise references with theorem numbers.
  4. [§6.3] The headline claim—that argmax_k of the infinitesimal part of ∥P∥_(k,p) for p ∈ [1,2) identifies the optimal classification number—is supported only by one simulated dumbbell chain with a single random trajectory. No theorem bridges the maximizer of the Ky Fan p-k-norm infinitesimal part to causal emergence; the results in §6.2 concern extremal properties of EId and the Schatten p-norm for special DTPMs (permutation matrices and identical-column matrices), and do not address this maximizer. The construction of P via the two least-squares problems with Xi = x(2:T+1) − x(1:T−1) and Yi = x(3:T+2) − x(2:T+1) is ad hoc, and the trajectory length T, the random initial condition, and the block transition probabilities are not specified. No error bars, multiple seeds, null-model controls, or alternative systems are reported, so the observed peak at k=5 for p=1.3, 1.6, 1.9 (but not for p=1) may be an artifact of the fitting procedure and parameter selection. This needs either a theoretical justification or a systematic numerical study before the abstract's claim can be accepted.
minor comments (4)
  1. [Abstract and §6.3] The abstract states that p is adjusted in [1,2), but the reported experimental peaks occur for p=1.3, 1.6, 1.9, and the text in §6.3 uses the interval (1,2]. Please clarify whether p=1 also gives the peak and align the notation between the abstract and the body.
  2. [Throughout] The same double-bar notation ∥·∥ is used for both real norms and their dual continuations. Since the domain is usually clear from the argument, this is acceptable, but in statements like Theorem 4.2 and Theorem 5.2 where both appear in one formula, a superscript or subscript would improve readability.
  3. [§5.3, Eq. (5.5)] Formula (5.5) is stated for non-zero As with singular values satisfying (5.6), but the case σ_k = 0 is not addressed. When As has rank less than k, the strict inequality in (5.6) cannot hold if there are additional zero singular values beyond the k-th. Please state the assumptions precisely or add the zero-singular-value case.
  4. [§6.3] The description of the k-means step uses the already-determined k=5 to construct Q1 and Q2, which is a reasonable two-stage procedure but should be stated more explicitly so that the reader does not infer circularity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the core norm and EI_d derivations are external, and the causal-emergence k-selection is an empirical demonstration rather than a fitted prediction.

full rationale

The paper's theoretical chain is self-contained relative to external mathematical inputs. The dual continuation (Def. 3.3) is defined as phi(a) + D_b phi(a) epsilon, and the dual-valued vector/matrix norm formulas in Theorems 4.2 and 5.2 are proved from convexity, the max formula (Borwein-Lewis), and Watson's subdifferential characterizations (Lemmas 5.5, 5.7, 5.11, 5.14, 5.15). These are external results, not assumptions of the conclusions. Theorem 5.20, equating the dual Ky Fan p-k-norm to the dual vector p-norm of the first k singular values, invokes the authors' prior CDSVD theorem [26]; that is a published theorem with stated assumptions and a proof in prior work, and it does not presuppose the norm equality or the causal-emergence conclusion. The causal-emergence claim in Sec. 6.3 is not derived by setting the fitted parameters equal to the output: the DTPM P = Ps + Pi epsilon is obtained from two least-squares problems, and then the peak of the infinitesimal part of ||P||_(k,p) is found at k=5 for p=1.3,1.6,1.9. This peak is observed, not imposed; there is no equation in which the optimal k is defined as the maximizer and then reported as a prediction. The conclusion that this characterizes optimal classification is an empirical heuristic supported by one dumbbell-chain run, which is a limitation (no multiple seeds, null controls, or error bars) but not a circularity. No quoted step reduces a claimed prediction to its own input.

Assumptions & free parameters 2 free parameters · 4 assumptions · 2 invented entities

The mathematical framework removes the need for an explicit derivative of a norm by using subdifferentials, but it introduces a fitted matrix P_i as the model of infinitesimal dynamics. The causal-emergence claim also depends on choosing p in [1,2) and on the prior CDSVD result. These are the main borrowed or hand-chosen elements.

free parameters (2)
  • infinitesimal part P_i of the DTPM = solution of min ||Y_i - P_s X_i - P_i X_s||_F subject to 1^T P_i = 0 and nonnegativity on the support of P_s
    The dual-valued norm and EI_d depend on P_i, which is fitted to time-series data via two successive convex programs. A different fitting objective would change the computed k, so the central applied claim rests on this fitted quantity.
  • p values in [1,2) = 1.3, 1.6, 1.9 chosen for the experiment
    The paper reports the k=5 peak only for selected p values. No theory determines which p to use, and the robustness of the peak across p is not established.
assumptions (4)
  • standard math Subdifferential formulas for vector p-norms and Ky Fan p-k norms, as given by Watson and others.
    Used in Propositions 4.4, 5.8-5.10, and Lemma 5.13. These are standard results in convex analysis.
  • domain assumption Existence and properties of compact dual SVD (CDSVD) from the authors' prior work [26].
    Theorem 5.20 assumes the dual matrix has a CDSVD, and the clustering algorithm in Section 6.3 relies on it. This is the authors' own prior result, not re-derived here.
  • ad hoc to paper Dual continuation via Gateaux derivative is the correct extension for non-differentiable functions.
    Definition 3.3 postulates this extension for all real-valued functions. It is natural and reduces to the standard formula for differentiable functions, but it is not derived from a uniqueness principle.
  • ad hoc to paper The DTPM obtained from least-squares fitting encodes the system's causal structure.
    Section 6.3 constructs P_s and P_i by minimizing Frobenius residuals. There is no guarantee that the resulting dual matrix captures the causal information relevant to coarse-graining.
invented entities (2)
  • Dual Transitional Probability Matrix (DTPM)
    purpose: A dual matrix extending a TPM with an infinitesimal part, used to define EI_d and the Ky Fan norm heuristic for causal emergence.
    The DTPM is defined in this paper. It yields a falsifiable prediction about k, but the object itself is a new mathematical construct with no independent physical evidence.
  • Dual-valued effective information (EI_d)
    purpose: Extends the effective information of a TPM to a dual number, with standard part EI(P_s) and an infinitesimal correction from P_i.
    This is a new definition introduced by the paper. Its usefulness is asserted through the numerical experiment.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dual-Valued Functions of Dual Matrices with Applications in Causal Emergence." pith.science (2026). https://pith.science/paper/ILHMQ7QX

@misc{pith2026241108377,
  author       = {Pith},
  title        = {Pith review of: Dual-Valued Functions of Dual Matrices with Applications in Causal Emergence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ILHMQ7QX}},
  note         = {Machine review of arXiv:2411.08377}
}
abstract

Dual continuation, an innovative insight into extending the real-valued functions of real matrices to the dual-valued functions of dual matrices with a foundation of the G\^ateaux derivative, is proposed. Theoretically, the general forms of dual-valued vector and matrix norms, the remaining properties in the real field, are provided. In particular, we focus on the dual-valued vector $p$-norm $(1\!\leq\! p\!\leq\!\infty)$ and the unitarily invariant dual-valued Ky Fan $p$-$k$-norm $(1\!\leq\! p\!\leq\!\infty)$. The equivalence between the dual-valued Ky Fan $p$-$k$-norm and the dual-valued vector $p$-norm of the first $k$ singular values of the dual matrix is then demonstrated. Practically, we define the dual transitional probability matrix (DTPM), as well as its dual-valued effective information (${\rm{EI_d}}$). Additionally, we elucidate the correlation between the ${\rm{EI_d}}$, the dual-valued Schatten $p$-norm, and the dynamical reversibility of a DTPM. Through numerical experiments on a dumbbell Markov chain, our findings indicate that the value of $k$, corresponding to the maximum value of the infinitesimal part of the dual-valued Ky Fan $p$-$k$-norm by adjusting $p$ in the interval $[1,2)$, characterizes the optimal classification number of the system for the occurrence of the causal emergence.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 27 canonical work pages

  1. [1]

    J. M. Borwein and A. S. Lewis , Convex Analysis and Nonlinear Optimization: Theory and Examples, Springer New York, NY, 2 ed., 2006

  2. [2]

    W. K. Clifford, Preliminary Sketch of Biquaternions , Proc. London Math. Soc., s1-4 (1871), pp. 381–395

  3. [3]

    CVX Research , CVX: Matlab software for disciplined convex programming, version 2.0

    I. CVX Research , CVX: Matlab software for disciplined convex programming, version 2.0 . https://cvxr.com/cvx, 2012

  4. [4]

    Ding and N

    J. Ding and N. H. Rhee , Teaching Tip: When a Matrix and Its Inverse Are Stochastic , The College Mathematics Journal, 44 (2013), pp. 108–109

  5. [5]

    X. V. Doan and S. V avasis, Finding the Largest Low-Rank Clusters With Ky Fan 2- k-norm and ℓ1-Norm, SIAM J. Optim., 26 (2016), pp. 274–312

  6. [6]

    Gˆateaux, Sur les Fonctionnelles Continues et les Fonctionnelles Analytiques , C.R

    R. Gˆateaux, Sur les Fonctionnelles Continues et les Fonctionnelles Analytiques , C.R. Acad. Sci. Paris S´er. I Math., 157 (1913), pp. 325–327

  7. [7]

    Grant and S

    M. Grant and S. Boyd , Graph implementations for nonsmooth convex programs , in Recent Advances in Learning and Control, Lecture Notes in Control and Information Sciences, Springer-Verlag Limited, 2008, pp. 95–110

  8. [8]

    Gu and J

    Y. Gu and J. Y. S. Luh, Dual-Number Transformation and Its Applications to Robotics, IEEE Journal on Robotics and Automation, 3 (1987), pp. 615–623

Show all 30 references
  1. [9]

    M. A. G ¨ung¨or and ¨O. Tetik , De-Moivre and Euler Formulae for Dual-Complex Numbers , Universal Journal of Mathematics and Applications, 2 (2019), pp. 126–129

  2. [10]

    J. A. Hartigan and M. A. Wong, A K-Means Clustering Algorithm , J. R. Stat. Soc. C Appl. Stat., 28 (1979), pp. 100–108

  3. [11]

    E. P. Hoel , When the Map is Better Than the Territory , Entropy, 19 (2017), p. 188. 32 TONG WEI, WEIYANG DING, AND YIMIN WEI Table 2 Dual-Valued Functions of Dual Matrices Dual-valued matrix normsTheorems F ormulas Dual-valued Ky Fanp-k-norm(1< p <∞) Proposition 5.8∥A∥(k,p)= ...

  4. [12]

    E. P. Hoel, L. Albantakis, and G. Tononi , Quantifying Causal Emergence Shows That Macro Can Beat Micro , Proc. Natl. Acad. Sci. USA, 110 (2013), pp. 19790–19795

  5. [13]

    R. A. Horn and C. R. Johnson , Matrix Analysis , Cambridge University Press, 2012

  6. [14]

    D. J. Keˇcki´c, Orthogonality in O1 and O∞ Spaces and Normal Derivations, J. Operat. Theor., 51 (2004), pp. 89–104

  7. [15]

    E. E. Kramer, Polygenic Functions of the Dual Variable w = u + jv, Am. J. Math., 52 (1930), pp. 370–376

  8. [16]

    Messelmi , Dual-Complex Numbers and Their Holomorphic Functions , HAL Id: hal- 01114178, (2015)

    F. Messelmi , Dual-Complex Numbers and Their Holomorphic Functions , HAL Id: hal- 01114178, (2015)

  9. [17]

    Miao and Z

    X. Miao and Z. Huang , Norms of Dual Complex Vectors and Dual Complex Matrices , Com- mun. Appl. Math. Comput., 5 (2023), pp. 1484–1508

  10. [18]

    Nesterov, Lectures on Convex Optimization , Springer Cham, 2 ed., 2018

    Y. Nesterov, Lectures on Convex Optimization , Springer Cham, 2 ed., 2018

  11. [19]

    Pennestr `ı and R

    E. Pennestr `ı and R. Stefanelli , Linear Algebra and Numerical Algorithms Using Dual Numbers, Multibody Syst. Dyn., 18 (2007), pp. 323–344

  12. [20]

    L. Qi, D. M. Alexander, Z. Chen, C. Ling, and Z. Luo , Low Rank Approximation of Dual Complex Matrices, 2022, https://arxiv.org/abs/2201.12781

  13. [21]

    Qi and C

    L. Qi and C. Cui, Dual Markov Chain and Dual Number Matrices with Nonnegative Standard Parts, Commun. Appl. Math. Comput., (2024), pp. 1–20

  14. [22]

    L. Qi, C. Ling, and H. Yan, Dual Quaternions and Dual Quaternion Vectors , Commun. Appl. Math. Comput., 4 (2022), p. 1494–1508

  15. [23]

    Study, Von den Bewegungen und Umlegungen , Math

    E. Study, Von den Bewegungen und Umlegungen , Math. Ann., 39 (1891), pp. 441–565

  16. [24]

    W atson, Characterization of the Subdifferential of Some Matrix Norms , Linear Algebra and its Applications, 170 (1992), pp

    G. W atson, Characterization of the Subdifferential of Some Matrix Norms , Linear Algebra and its Applications, 170 (1992), pp. 33–45

  17. [25]

    W atson, On Matrix Approximation Problems with Ky Fan k Norms , Numer

    G. W atson, On Matrix Approximation Problems with Ky Fan k Norms , Numer. Algorithms, 5 (1993), pp. 263–272. DUAL-V ALUED FUNCTIONS 33

  18. [26]

    T. Wei, W. Ding, and Y. Wei , Singular Value Decomposition of Dual Matrices and Its Application to Traveling Wave Identification in the Brain , SIAM J. Matrix Anal. Appl., 45 (2024), pp. 634–660

  19. [27]

    M. Yang, Z. W ang, K. Liu, Y. Rong, B. Yuan, and J. Zhang , Finding Emergence in Data by Maximizing Effective Information , Natl. Sci. Rev., (2024), p. nwae279

  20. [28]

    B. Yuan, J. Zhang, A. Lyu, J. Wu, Z. W ang, M. Yang, K. Liu, M. Mou, and P. Cui , Emergence and Causality in Complex Systems: A Survey of Causal Emergence and Related Quantitative Studies , Entropy, 26 (2024), p. 108

  21. [29]

    Zhang, R

    J. Zhang, R. Tao, K. H. Leong, M. Yang, and B. Yuan , Dynamical Reversibility and A New Theory of Causal Emergence , 2024, https://arxiv.org/abs/2402.15054

  22. [30]

    Zhang, The Singular Value Decomposition, Applications and Beyond , 2015, https://arxiv

    Z. Zhang, The Singular Value Decomposition, Applications and Beyond , 2015, https://arxiv. org/abs/1510.08532

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.