REVIEW 4 major objections 4 minor 30 references
Dual-Valued Functions of Dual Matrices with Applications in Causal Emergence
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that extending matrix norms to dual matrices via the Gâteaux derivative yields a dual-valued Ky Fan p-k-norm whose infinitesimal part, maximized over k as p varies in [1,2), identifies the optimal macro-state count at…
desk verdict Solid dual-continuation theory; the causal-emergence k-selection claim is a single-example empirical conjecture that needs more support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the dual continuation of the Ky Fan $p$-$k$-norm, whose infinitesimal part is obtained by maximizing $\langle G, A_i\rangle$ over $G$ in the subdifferential of the ordinary Ky Fan $p$-$k$-norm at $A_s$; this is what makes the norm sensitive to the infinitesimal structure of the transition matrix. The companion identity $\|A\|_{(k,p)}=\|\sigma^k\|_p$, where $\sigma^k$ is the dual vector of the first $k$ dual singular values, lets the norm be computed from a dual SVD. In the application, the DTPM $P=P_s+P_i\epsilon$ is fitted by solving two least-squares problems, and the $k$ that maximizes the infinitesimal part of $\|P\|_{(k,p)}$ over $p\in[1,2)$ is declared the optimal classification number.
What would settle it
Construct a Markov chain with a known number of metastable groups but no clear gap in its singular values; if the $k$ maximizing the infinitesimal part of $\|P\|_{(k,p)}$ over $p\in[1,2)$ does not equal that known number, or if the peak moves when $p$ is varied within $[1,2)$ or when the least-squares fitting is perturbed, the central claim would be falsified.
Extended reading notes
Core claim
The central claim is that the dual continuation of a real-valued function, sending $\varphi(a+b\epsilon)$ to $\varphi(a)+D_b\varphi(a)\epsilon$ with $D_b$ the Gâteaux derivative, produces valid dual-valued vector and matrix norms that keep the real-field properties, and that the resulting dual-valued Ky Fan $p$-$k$-norm carries usable information about causal emergence. In particular, for a dual transition probability matrix $P=P_s+P_i\epsilon$ fitted from time-series data, the infinitesimal part of $\|P\|_{(k,p)}$, maximized over $k$ with $p$ in $[1,2)$, identifies the optimal number of macro-states. The paper proves the norm identities, including $\|A\|_{(k,p)}=\|\sigma^k\|_p$ for the vector of the first $k$ dual singular values, and it shows that dual effective information, the dual Schatten $p$-norm, and dynamical reversibility of a DTPM all peak exactly at permutation matrices for $1\le p<2$. The numerical experiment on a dumbbell chain finds the peak at $k=5$, the known number of groups, supporting the claim.
Load-bearing premise
The method assumes that the dual transition matrix fitted from time-series data faithfully reflects the system's causal structure, so that the $k$ maximizing the infinitesimal part of the dual Ky Fan $p$-$k$-norm reliably marks the best coarse-graining scale; this is an empirical heuristic demonstrated on one dumbbell chain, not a proven theorem.
Editorial extensions
If this is right
- Optimal coarse-graining can be read off directly from the dual Ky Fan $p$-$k$-norm, without enumerating candidate coarse-grainings or locating a singular-value cutoff.
- The identity $\|A\|_{(k,p)}=\|\sigma^k\|_p$ gives a computable route: take a dual SVD of the DTPM, form the dual vector of its first $k$ singular values, and evaluate its dual vector $p$-norm.
- For $1\le p<2$, a DTPM reaches its maximum dual Schatten $p$-norm exactly when it is a permutation matrix, and dual effective information reaches its maximum under the same condition, so the norm and effective information agree on the most reversible macro-state.
- The procedure can be repeated on other time-series data: fit $P_s$ and $P_i$, compute the infinitesimal part of $\|P\|_{(k,p)}$, and take the maximizing $k$ as the system's optimal classification number.
Reading between the lines
- Beyond the paper: if the maximizing-$k$ heuristic survives on more systems, dual-matrix norms could serve as a general model-order selection tool for Markov chains with slow mixing and no spectral gap, replacing subjective singular-value thresholds.
- Beyond the paper: the same dual-continuation construction could be applied to other unitarily invariant norms or to complex dual matrices, with the Gâteaux-derivative formalism suggesting analogous identities for condition numbers or entropy-like quantities.
- Beyond the paper: a direct testable extension is to compare the $k$ suggested by the infinitesimal part with metastability indicators on a family of random dumbbell-like chains with known cluster counts and no singular-value gap.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a general framework for extending real-valued functions on vectors and matrices to dual-valued functions using the Gâteaux derivative, which it calls "dual continuation." It derives explicit formulas for dual-valued vector p-norms and dual-valued unitarily invariant matrix norms, including the Ky Fan p-k-norm, Schatten p-norm, nuclear norm, Frobenius norm, and operator norms, and proves several structural properties such as unitary invariance and consistency. The paper then introduces a dual transitional probability matrix (DTPM) and a dual-valued effective information EId, and studies their extremal properties in relation to dynamical reversibility. In a numerical experiment on a dumbbell Markov chain, the authors observe that the infinitesimal part of the dual-valued Ky Fan p-k-norm attains a maximum at k=5 for p=1.3,1.6,1.9, and they interpret this k as the optimal number of macro-states for causal emergence.
Significance. If the theoretical results are correct, the dual-continuation construction provides a systematic method for extending nonsmooth convex matrix functions to dual matrices, which could be useful in dual-number-based numerical analysis and optimization. The explicit formulas for dual-valued norms and the equivalence in Theorem 5.20 are potentially valuable reference results. The paper's strength is its use of standard convex-analysis tools (Borwein-Lewis, Watson, Nesterov) and the explicit nature of the derived formulas. However, the central applied claim concerning causal emergence rests on a single numerical example and an ad hoc construction of the DTPM, with no theoretical bridge from the Ky Fan p-k-norm maximizer to an optimal classification number. The theoretical gaps described in the major comments are load-bearing for the paper's main claims.
major comments (4)
- [§5.4, Eq. (5.2)] The operator-norm formula (5.2) is not well-posed for dual vectors x with zero standard part. In the dual-number algebra, division is defined only when the divisor has a nonzero standard part, yet the maximization in (5.2) is taken over all x ∈ DR^n with x ≠ 0, including x = x_i ε. The proof of Theorem 5.4 handles the As = O case by choosing x_i = 0 and x_s ≠ 0, but this does not justify the unrestricted definition. Either restrict the definition to vectors with nonzero standard part, or define and justify division by infinitesimal dual numbers, and then show the maximum is attained. As stated, the equivalence between the dual continuation and the induced operator norm is not established.
- [§5.3, Lemma 5.13 and Prop. 5.8] Lemma 5.13 states that any G ∈ ∂∥X∥_(k,p) has the form (5.9), but it does not establish the converse direction, namely that every symmetric positive semidefinite T with ∥T∥_2 ≤ 1 and ∥T∥_* = t gives an element of the subdifferential. Proposition 5.8 then maximizes ⟨G, A_i⟩ over the set of G of that form; if the converse fails, the maximum over the true subdifferential could be larger, and the formula (5.5) would be invalid. This is load-bearing because (5.5) underlies Corollary 5.16 and the Section 6 applications. Please provide a full characterization of ∂∥X∥_(k,p) (necessity and sufficiency) or cite a reference that contains it.
- [§5.3–5.4, Corollaries 5.16–5.18 and Prop. 5.22] The proofs of Corollaries 5.16, 5.17, 5.18, and Proposition 5.22 are omitted with the note "The detailed proofs are omitted." These results are used later: Corollary 5.16 is used in Proposition 6.6 to characterize maxima/minima of the dual-valued Schatten p-norm of a DTPM, and Proposition 5.22 concerns the dual-valued operator ∞-norm. The omissions are not merely cosmetic; they leave unverified the exact infinitesimal-part formulas that drive the causal-emergence analysis. Please include complete proofs or give precise references with theorem numbers.
- [§6.3] The headline claim—that argmax_k of the infinitesimal part of ∥P∥_(k,p) for p ∈ [1,2) identifies the optimal classification number—is supported only by one simulated dumbbell chain with a single random trajectory. No theorem bridges the maximizer of the Ky Fan p-k-norm infinitesimal part to causal emergence; the results in §6.2 concern extremal properties of EId and the Schatten p-norm for special DTPMs (permutation matrices and identical-column matrices), and do not address this maximizer. The construction of P via the two least-squares problems with Xi = x(2:T+1) − x(1:T−1) and Yi = x(3:T+2) − x(2:T+1) is ad hoc, and the trajectory length T, the random initial condition, and the block transition probabilities are not specified. No error bars, multiple seeds, null-model controls, or alternative systems are reported, so the observed peak at k=5 for p=1.3, 1.6, 1.9 (but not for p=1) may be an artifact of the fitting procedure and parameter selection. This needs either a theoretical justification or a systematic numerical study before the abstract's claim can be accepted.
minor comments (4)
- [Abstract and §6.3] The abstract states that p is adjusted in [1,2), but the reported experimental peaks occur for p=1.3, 1.6, 1.9, and the text in §6.3 uses the interval (1,2]. Please clarify whether p=1 also gives the peak and align the notation between the abstract and the body.
- [Throughout] The same double-bar notation ∥·∥ is used for both real norms and their dual continuations. Since the domain is usually clear from the argument, this is acceptable, but in statements like Theorem 4.2 and Theorem 5.2 where both appear in one formula, a superscript or subscript would improve readability.
- [§5.3, Eq. (5.5)] Formula (5.5) is stated for non-zero As with singular values satisfying (5.6), but the case σ_k = 0 is not addressed. When As has rank less than k, the strict inequality in (5.6) cannot hold if there are additional zero singular values beyond the k-th. Please state the assumptions precisely or add the zero-singular-value case.
- [§6.3] The description of the k-means step uses the already-determined k=5 to construct Q1 and Q2, which is a reasonable two-stage procedure but should be stated more explicitly so that the reader does not infer circularity.
Circularity Check
No significant circularity; the core norm and EI_d derivations are external, and the causal-emergence k-selection is an empirical demonstration rather than a fitted prediction.
full rationale
The paper's theoretical chain is self-contained relative to external mathematical inputs. The dual continuation (Def. 3.3) is defined as phi(a) + D_b phi(a) epsilon, and the dual-valued vector/matrix norm formulas in Theorems 4.2 and 5.2 are proved from convexity, the max formula (Borwein-Lewis), and Watson's subdifferential characterizations (Lemmas 5.5, 5.7, 5.11, 5.14, 5.15). These are external results, not assumptions of the conclusions. Theorem 5.20, equating the dual Ky Fan p-k-norm to the dual vector p-norm of the first k singular values, invokes the authors' prior CDSVD theorem [26]; that is a published theorem with stated assumptions and a proof in prior work, and it does not presuppose the norm equality or the causal-emergence conclusion. The causal-emergence claim in Sec. 6.3 is not derived by setting the fitted parameters equal to the output: the DTPM P = Ps + Pi epsilon is obtained from two least-squares problems, and then the peak of the infinitesimal part of ||P||_(k,p) is found at k=5 for p=1.3,1.6,1.9. This peak is observed, not imposed; there is no equation in which the optimal k is defined as the maximizer and then reported as a prediction. The conclusion that this characterizes optimal classification is an empirical heuristic supported by one dumbbell-chain run, which is a limitation (no multiple seeds, null controls, or error bars) but not a circularity. No quoted step reduces a claimed prediction to its own input.
Assumptions & free parameters
free parameters (2)
- infinitesimal part P_i of the DTPM =
solution of min ||Y_i - P_s X_i - P_i X_s||_F subject to 1^T P_i = 0 and nonnegativity on the support of P_s
- p values in [1,2) =
1.3, 1.6, 1.9 chosen for the experiment
assumptions (4)
- standard math Subdifferential formulas for vector p-norms and Ky Fan p-k norms, as given by Watson and others.
- domain assumption Existence and properties of compact dual SVD (CDSVD) from the authors' prior work [26].
- ad hoc to paper Dual continuation via Gateaux derivative is the correct extension for non-differentiable functions.
- ad hoc to paper The DTPM obtained from least-squares fitting encodes the system's causal structure.
invented entities (2)
-
Dual Transitional Probability Matrix (DTPM)
-
Dual-valued effective information (EI_d)
Cite this review
Pith. "Pith review of Dual-Valued Functions of Dual Matrices with Applications in Causal Emergence." pith.science (2026). https://pith.science/paper/ILHMQ7QX
@misc{pith2026241108377,
author = {Pith},
title = {Pith review of: Dual-Valued Functions of Dual Matrices with Applications in Causal Emergence},
year = {2026},
howpublished = {\url{https://pith.science/paper/ILHMQ7QX}},
note = {Machine review of arXiv:2411.08377}
}
abstract
Dual continuation, an innovative insight into extending the real-valued functions of real matrices to the dual-valued functions of dual matrices with a foundation of the G\^ateaux derivative, is proposed. Theoretically, the general forms of dual-valued vector and matrix norms, the remaining properties in the real field, are provided. In particular, we focus on the dual-valued vector $p$-norm $(1\!\leq\! p\!\leq\!\infty)$ and the unitarily invariant dual-valued Ky Fan $p$-$k$-norm $(1\!\leq\! p\!\leq\!\infty)$. The equivalence between the dual-valued Ky Fan $p$-$k$-norm and the dual-valued vector $p$-norm of the first $k$ singular values of the dual matrix is then demonstrated. Practically, we define the dual transitional probability matrix (DTPM), as well as its dual-valued effective information (${\rm{EI_d}}$). Additionally, we elucidate the correlation between the ${\rm{EI_d}}$, the dual-valued Schatten $p$-norm, and the dynamical reversibility of a DTPM. Through numerical experiments on a dumbbell Markov chain, our findings indicate that the value of $k$, corresponding to the maximum value of the infinitesimal part of the dual-valued Ky Fan $p$-$k$-norm by adjusting $p$ in the interval $[1,2)$, characterizes the optimal classification number of the system for the occurrence of the causal emergence.
Reference graph
Works this paper leans on
-
[1]
J. M. Borwein and A. S. Lewis , Convex Analysis and Nonlinear Optimization: Theory and Examples, Springer New York, NY, 2 ed., 2006
work page 2006
-
[2]
W. K. Clifford, Preliminary Sketch of Biquaternions , Proc. London Math. Soc., s1-4 (1871), pp. 381–395
-
[3]
CVX Research , CVX: Matlab software for disciplined convex programming, version 2.0
I. CVX Research , CVX: Matlab software for disciplined convex programming, version 2.0 . https://cvxr.com/cvx, 2012
work page 2012
-
[4]
J. Ding and N. H. Rhee , Teaching Tip: When a Matrix and Its Inverse Are Stochastic , The College Mathematics Journal, 44 (2013), pp. 108–109
work page 2013
-
[5]
X. V. Doan and S. V avasis, Finding the Largest Low-Rank Clusters With Ky Fan 2- k-norm and ℓ1-Norm, SIAM J. Optim., 26 (2016), pp. 274–312
work page 2016
-
[6]
Gˆateaux, Sur les Fonctionnelles Continues et les Fonctionnelles Analytiques , C.R
R. Gˆateaux, Sur les Fonctionnelles Continues et les Fonctionnelles Analytiques , C.R. Acad. Sci. Paris S´er. I Math., 157 (1913), pp. 325–327
work page 1913
-
[7]
M. Grant and S. Boyd , Graph implementations for nonsmooth convex programs , in Recent Advances in Learning and Control, Lecture Notes in Control and Information Sciences, Springer-Verlag Limited, 2008, pp. 95–110
work page 2008
- [8]
Show all 30 references
-
[9]
M. A. G ¨ung¨or and ¨O. Tetik , De-Moivre and Euler Formulae for Dual-Complex Numbers , Universal Journal of Mathematics and Applications, 2 (2019), pp. 126–129
2019
-
[10]
J. A. Hartigan and M. A. Wong, A K-Means Clustering Algorithm , J. R. Stat. Soc. C Appl. Stat., 28 (1979), pp. 100–108
1979
-
[11]
E. P. Hoel , When the Map is Better Than the Territory , Entropy, 19 (2017), p. 188. 32 TONG WEI, WEIYANG DING, AND YIMIN WEI Table 2 Dual-Valued Functions of Dual Matrices Dual-valued matrix normsTheorems F ormulas Dual-valued Ky Fanp-k-norm(1< p <∞) Proposition 5.8∥A∥(k,p)= ...
2017
-
[12]
E. P. Hoel, L. Albantakis, and G. Tononi , Quantifying Causal Emergence Shows That Macro Can Beat Micro , Proc. Natl. Acad. Sci. USA, 110 (2013), pp. 19790–19795
2013
-
[13]
R. A. Horn and C. R. Johnson , Matrix Analysis , Cambridge University Press, 2012
2012
-
[14]
D. J. Keˇcki´c, Orthogonality in O1 and O∞ Spaces and Normal Derivations, J. Operat. Theor., 51 (2004), pp. 89–104
2004
-
[15]
E. E. Kramer, Polygenic Functions of the Dual Variable w = u + jv, Am. J. Math., 52 (1930), pp. 370–376
1930
-
[16]
Messelmi , Dual-Complex Numbers and Their Holomorphic Functions , HAL Id: hal- 01114178, (2015)
F. Messelmi , Dual-Complex Numbers and Their Holomorphic Functions , HAL Id: hal- 01114178, (2015)
2015
-
[17]
Miao and Z
X. Miao and Z. Huang , Norms of Dual Complex Vectors and Dual Complex Matrices , Com- mun. Appl. Math. Comput., 5 (2023), pp. 1484–1508
2023
-
[18]
Nesterov, Lectures on Convex Optimization , Springer Cham, 2 ed., 2018
Y. Nesterov, Lectures on Convex Optimization , Springer Cham, 2 ed., 2018
2018
-
[19]
Pennestr `ı and R
E. Pennestr `ı and R. Stefanelli , Linear Algebra and Numerical Algorithms Using Dual Numbers, Multibody Syst. Dyn., 18 (2007), pp. 323–344
2007
-
[20]
L. Qi, D. M. Alexander, Z. Chen, C. Ling, and Z. Luo , Low Rank Approximation of Dual Complex Matrices, 2022, https://arxiv.org/abs/2201.12781
2022 arXiv
-
[21]
Qi and C
L. Qi and C. Cui, Dual Markov Chain and Dual Number Matrices with Nonnegative Standard Parts, Commun. Appl. Math. Comput., (2024), pp. 1–20
2024
-
[22]
L. Qi, C. Ling, and H. Yan, Dual Quaternions and Dual Quaternion Vectors , Commun. Appl. Math. Comput., 4 (2022), p. 1494–1508
2022
-
[23]
Study, Von den Bewegungen und Umlegungen , Math
E. Study, Von den Bewegungen und Umlegungen , Math. Ann., 39 (1891), pp. 441–565
-
[24]
W atson, Characterization of the Subdifferential of Some Matrix Norms , Linear Algebra and its Applications, 170 (1992), pp
G. W atson, Characterization of the Subdifferential of Some Matrix Norms , Linear Algebra and its Applications, 170 (1992), pp. 33–45
1992
-
[25]
W atson, On Matrix Approximation Problems with Ky Fan k Norms , Numer
G. W atson, On Matrix Approximation Problems with Ky Fan k Norms , Numer. Algorithms, 5 (1993), pp. 263–272. DUAL-V ALUED FUNCTIONS 33
1993
-
[26]
T. Wei, W. Ding, and Y. Wei , Singular Value Decomposition of Dual Matrices and Its Application to Traveling Wave Identification in the Brain , SIAM J. Matrix Anal. Appl., 45 (2024), pp. 634–660
2024
-
[27]
M. Yang, Z. W ang, K. Liu, Y. Rong, B. Yuan, and J. Zhang , Finding Emergence in Data by Maximizing Effective Information , Natl. Sci. Rev., (2024), p. nwae279
2024
-
[28]
B. Yuan, J. Zhang, A. Lyu, J. Wu, Z. W ang, M. Yang, K. Liu, M. Mou, and P. Cui , Emergence and Causality in Complex Systems: A Survey of Causal Emergence and Related Quantitative Studies , Entropy, 26 (2024), p. 108
2024
-
[29]
Zhang, R
J. Zhang, R. Tao, K. H. Leong, M. Yang, and B. Yuan , Dynamical Reversibility and A New Theory of Causal Emergence , 2024, https://arxiv.org/abs/2402.15054
2024 arXiv
-
[30]
Zhang, The Singular Value Decomposition, Applications and Beyond , 2015, https://arxiv
Z. Zhang, The Singular Value Decomposition, Applications and Beyond , 2015, https://arxiv. org/abs/1510.08532
2015 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.