{"id":"01de4e80-be8b-4651-a765-1f489e4bd7c2","arxiv_id":"2411.08377","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper defines dual-valued norms via Gateaux derivatives and uses the infinitesimal part of a dual Ky Fan norm to estimate the optimal number of coarse-grained states in causal emergence.","lead":"This paper extends ordinary matrix functions to dual numbers, which have a small extra part that squares to zero, using Gateaux derivatives to define norms of dual matrices. It then applies these norms to a Markov chain example, claiming the method reveals how many coarse-grained groups best capture causal emergence.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The applied claim rests on a single fitted DTPM example; the peak at k=5 may be an artifact of the least-squares construction and p selection, with no theorem linking it to causal emergence.","rationale":"The reader's CONDITIONAL verdict is appropriate. The theoretical norm derivations appear coherent and self-contained, so the concern is not internal inconsistency but an unvalidated empirical bridge: the fitted DTPM's infinitesimal part and the Ky Fan p-k norm maximizer are asserted, without proof or robustness evidence, to reveal the optimal classification number for causal emergence. My stress-test agrees with the reader's weakest assumption and sharpens it by pointing to the arbitrary least-squares construction of Xi and Yi and the absence of any theoretical connection between the infinitesimal Ky Fan norm maximizer and EI or coarse-graining. The proposed concrete test would settle whether the peak is a genuine system property or a fitting artifact. Since the reader already assigned CONDITIONAL based on this same weak spot, no verdict adjustment is needed.","tokens_in":34886,"tokens_out":3716,"duration_ms":36637,"concrete_test":"Run the exact Section 6.3 pipeline on (a) a three-community Markov chain with known grouping, (b) a random TPM with no block structure, and (c) the dumbbell chain with at least 20 independently sampled initial conditions and trajectory lengths T. If the argmax_k of the infinitesimal part of ||P||_(k,p) does not reproducibly match the known number of communities in (a), or produces a spurious sharp peak in (b), or varies across seeds in (c), then the central characterization fails and the paper would need either a theoretical justification or substantially broader empirical validation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central applied claim depends on an empirical bridge that the paper does not establish. Section 6.3 constructs P = Ps + Pi*epsilon by solving two least-squares problems from one random trajectory of the dumbbell chain, with Xi = x(2:T+1) - x(1:T-1) and Yi = x(3:T+2) - x(2:T+1). This dual perturbation model is arbitrary, and the subsequent claim that argmax_k of the infinitesimal part of ||P||_(k,p) equals the optimal classification number is supported only by the observed peak at k=5 for p = 1.3, 1.6, 1.9. Sections 6.1-6.2 characterize extrema of EI_d and the dual Schatten norm only for special DTPMs such as permutation matrices; no proposition states that the Ky Fan p-k norm infinitesimal maximizer identifies a macro-state count. Because Pi is fit to the residual Yi - Ps*Xi, the peak may encode the chosen initial condition, trajectory length T, or the specific block sizes rather than a robust system property. No error bars, multiple seeds, or null-model controls are reported, and the abstract's claim of p in [1,2) is not fully consistent with the reported p values. The headline claim is therefore an empirical conjecture, not a demonstrated result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a general framework for extending real-valued functions on vectors and matrices to dual-valued functions using the Gâteaux derivative, which it calls \"dual continuation.\" It derives explicit formulas for dual-valued vector p-norms and dual-valued unitarily invariant matrix norms, including the Ky Fan p-k-norm, Schatten p-norm, nuclear norm, Frobenius norm, and operator norms, and proves several structural properties such as unitary invariance and consistency. The paper then introduces a dual transitional probability matrix (DTPM) and a dual-valued effective information EId, and studies their extremal properties in relation to dynamical reversibility. In a numerical experiment on a dumbbell Markov chain, the authors observe that the infinitesimal part of the dual-valued Ky Fan p-k-norm attains a maximum at k=5 for p=1.3,1.6,1.9, and they interpret this k as the optimal number of macro-states for causal emergence.","tokens_in":35168,"tokens_out":9867,"duration_ms":95103,"significance":"If the theoretical results are correct, the dual-continuation construction provides a systematic method for extending nonsmooth convex matrix functions to dual matrices, which could be useful in dual-number-based numerical analysis and optimization. The explicit formulas for dual-valued norms and the equivalence in Theorem 5.20 are potentially valuable reference results. The paper's strength is its use of standard convex-analysis tools (Borwein-Lewis, Watson, Nesterov) and the explicit nature of the derived formulas. However, the central applied claim concerning causal emergence rests on a single numerical example and an ad hoc construction of the DTPM, with no theoretical bridge from the Ky Fan p-k-norm maximizer to an optimal classification number. The theoretical gaps described in the major comments are load-bearing for the paper's main claims.","major_comments":[{"comment":"The operator-norm formula (5.2) is not well-posed for dual vectors x with zero standard part. In the dual-number algebra, division is defined only when the divisor has a nonzero standard part, yet the maximization in (5.2) is taken over all x ∈ DR^n with x ≠ 0, including x = x_i ε. The proof of Theorem 5.4 handles the As = O case by choosing x_i = 0 and x_s ≠ 0, but this does not justify the unrestricted definition. Either restrict the definition to vectors with nonzero standard part, or define and justify division by infinitesimal dual numbers, and then show the maximum is attained. As stated, the equivalence between the dual continuation and the induced operator norm is not established.","section":"§5.4, Eq. (5.2)"},{"comment":"Lemma 5.13 states that any G ∈ ∂∥X∥_(k,p) has the form (5.9), but it does not establish the converse direction, namely that every symmetric positive semidefinite T with ∥T∥_2 ≤ 1 and ∥T∥_* = t gives an element of the subdifferential. Proposition 5.8 then maximizes ⟨G, A_i⟩ over the set of G of that form; if the converse fails, the maximum over the true subdifferential could be larger, and the formula (5.5) would be invalid. This is load-bearing because (5.5) underlies Corollary 5.16 and the Section 6 applications. Please provide a full characterization of ∂∥X∥_(k,p) (necessity and sufficiency) or cite a reference that contains it.","section":"§5.3, Lemma 5.13 and Prop. 5.8"},{"comment":"The proofs of Corollaries 5.16, 5.17, 5.18, and Proposition 5.22 are omitted with the note \"The detailed proofs are omitted.\" These results are used later: Corollary 5.16 is used in Proposition 6.6 to characterize maxima/minima of the dual-valued Schatten p-norm of a DTPM, and Proposition 5.22 concerns the dual-valued operator ∞-norm. The omissions are not merely cosmetic; they leave unverified the exact infinitesimal-part formulas that drive the causal-emergence analysis. Please include complete proofs or give precise references with theorem numbers.","section":"§5.3–5.4, Corollaries 5.16–5.18 and Prop. 5.22"},{"comment":"The headline claim—that argmax_k of the infinitesimal part of ∥P∥_(k,p) for p ∈ [1,2) identifies the optimal classification number—is supported only by one simulated dumbbell chain with a single random trajectory. No theorem bridges the maximizer of the Ky Fan p-k-norm infinitesimal part to causal emergence; the results in §6.2 concern extremal properties of EId and the Schatten p-norm for special DTPMs (permutation matrices and identical-column matrices), and do not address this maximizer. The construction of P via the two least-squares problems with Xi = x(2:T+1) − x(1:T−1) and Yi = x(3:T+2) − x(2:T+1) is ad hoc, and the trajectory length T, the random initial condition, and the block transition probabilities are not specified. No error bars, multiple seeds, null-model controls, or alternative systems are reported, so the observed peak at k=5 for p=1.3, 1.6, 1.9 (but not for p=1) may be an artifact of the fitting procedure and parameter selection. This needs either a theoretical justification or a systematic numerical study before the abstract's claim can be accepted.","section":"§6.3"}],"minor_comments":[{"comment":"The abstract states that p is adjusted in [1,2), but the reported experimental peaks occur for p=1.3, 1.6, 1.9, and the text in §6.3 uses the interval (1,2]. Please clarify whether p=1 also gives the peak and align the notation between the abstract and the body.","section":"Abstract and §6.3"},{"comment":"The same double-bar notation ∥·∥ is used for both real norms and their dual continuations. Since the domain is usually clear from the argument, this is acceptable, but in statements like Theorem 4.2 and Theorem 5.2 where both appear in one formula, a superscript or subscript would improve readability.","section":"Throughout"},{"comment":"Formula (5.5) is stated for non-zero As with singular values satisfying (5.6), but the case σ_k = 0 is not addressed. When As has rank less than k, the strict inequality in (5.6) cannot hold if there are additional zero singular values beyond the k-th. Please state the assumptions precisely or add the zero-singular-value case.","section":"§5.3, Eq. (5.5)"},{"comment":"The description of the k-means step uses the already-determined k=5 to construct Q1 and Q2, which is a reasonable two-stage procedure but should be stated more explicitly so that the reader does not infer circularity.","section":"§6.3"}],"recommendation":"major_revision","confidential_remarks":"The theoretical framework in Sections 3–5 is plausible and likely of interest to readers of math.NA, though the gaps in the subdifferential characterization and the operator-norm definition need attention. The causal-ecosystem claim in Section 6 is the main selling point of the paper but is not yet supported at the level expected for the claimed conclusion. I would encourage the authors to either prove a theorem linking the maximizer of the dual Ky Fan p-k-norm infinitesimal part to a well-defined notion of optimal macro-state count, or substantially expand the numerical validation with multiple systems, seeds, and null models. The paper relies on its own prior work [26] for the CDSVD; this is legitimate but should be clearly flagged as a dependency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The theoretical core is genuinely new and mostly solid; the causal-emergence claim is a single-example heuristic wearing the abstract's confident clothes. If you read this for the dual continuation machinery, you'll get something worth keeping. If you read it for the 'optimal classification number' claim, you'll need to demand more evidence.\n\nWhat the paper does well: the Gâteaux-derivative dual continuation is a clean idea that covers non-differentiable functions (norms) in a unified way. The explicit formulas for dual-valued vector p-norms, Ky Fan p-k norms, Schatten, nuclear, spectral, and operator norms are derived from standard subdifferential calculus (Watson's theorems), and I didn't find a fatal error in the derivations. The equivalence theorem between the Ky Fan p-k norm and the vector p-norm of the first k dual singular values is conditional on the authors' CDSVD, but that's legitimate prior work, not circular. The propositions on EId and Schatten norms of DTPMs (extrema, permutation-matrix characterization) are correct as mathematical statements.\n\nSoft spots, in proportion: (1) The operator norm definition (5.2) divides by a dual number, which is not defined for vectors with zero standard part. The proof quietly avoids this by choosing xi=0, but the definition needs a qualifier. Minor and fixable. (2) Several corollaries (5.16-5.18) and the operator infinity-norm are stated without proof; again minor, since they follow from the same subdifferential lemmas. (3) The load-bearing applied claim is not a theorem. No proposition states that argmax_k of the infinitesimal part of the Ky Fan p-k norm equals the optimal macro-state count. The only support is one dumbbell chain, one random trajectory, no error bars, no null model, and a p-range that in the abstract is [1,2) but in the experiment includes p=1, which does not show the peak. The least-squares construction of Pi is one of many possible; the peak at k=5 could easily be an artifact of how the residual is encoded.\n\nVerdict: the theory deserves a serious referee; the application does not yet support the headline. I'd accept for review and push for a major revision: either add a theoretical justification for the k-selection heuristic or substantially broaden the empirical support (multiple systems, seeds, null controls) and soften the abstract.","headline":"Solid dual-continuation theory; the causal-emergence k-selection claim is a single-example empirical conjecture that needs more support.","tokens_in":35705,"tokens_out":2721,"would_cite":true,"duration_ms":27021,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["15A60","15B33","30G35","65C40"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that extending matrix norms to dual matrices via the Gâteaux derivative yields a dual-valued Ky Fan p-k-norm whose infinitesimal part, maximized over k as p varies in [1,2), identifies the optimal macro-state count at…","keywords":["dual numbers","dual continuation","Gâteaux derivative","dual-valued matrix norms","Ky Fan p-k-norm","causal emergence","dual transitional probability matrix","effective information"],"falsifier":"Construct a Markov chain with a known number of metastable groups but no clear gap in its singular values; if the $k$ maximizing the infinitesimal part of $\\|P\\|_{(k,p)}$ over $p\\in[1,2)$ does not equal that known number, or if the peak moves when $p$ is varied within $[1,2)$ or when the least-squares fitting is perturbed, the central claim would be falsified.","tokens_in":34622,"feed_emoji":"📐","tokens_out":9528,"duration_ms":76460,"temperature":0.7,"pith_summary":"The paper tries to establish that real-valued matrix norms can be carried over to dual matrices—matrices whose entries are dual numbers $a + b\\epsilon$ with $\\epsilon^2=0$—by placing the Gâteaux derivative in the infinitesimal part, and that one resulting norm detects causal emergence. The authors define dual-valued vector and matrix norms, prove that the dual-valued Ky Fan $p$-$k$-norm equals the dual-valued vector $p$-norm of the first $k$ singular values, and introduce a dual transition probability matrix with dual effective information. On a dumbbell Markov chain, the $k$ that maximizes the infinitesimal part of the dual Ky Fan $p$-$k$-norm as $p$ ranges over $[1,2)$ matches the number of macro-states of the system. The payoff, if the claim generalizes, is a way to choose the optimal coarse-graining scale without enumerating all coarse-grainings.","feed_headline":"Dual norm's infinitesimal term finds the number of macro-states","feed_subtitle":"Adjusting p in [1,2), the peak of the dual Ky Fan p-k-norm marks the optimal classification number for causal emergence.","key_machinery":"The load-bearing object is the dual continuation of the Ky Fan $p$-$k$-norm, whose infinitesimal part is obtained by maximizing $\\langle G, A_i\\rangle$ over $G$ in the subdifferential of the ordinary Ky Fan $p$-$k$-norm at $A_s$; this is what makes the norm sensitive to the infinitesimal structure of the transition matrix. The companion identity $\\|A\\|_{(k,p)}=\\|\\sigma^k\\|_p$, where $\\sigma^k$ is the dual vector of the first $k$ dual singular values, lets the norm be computed from a dual SVD. In the application, the DTPM $P=P_s+P_i\\epsilon$ is fitted by solving two least-squares problems, and the $k$ that maximizes the infinitesimal part of $\\|P\\|_{(k,p)}$ over $p\\in[1,2)$ is declared the optimal classification number.","core_discovery":"The central claim is that the dual continuation of a real-valued function, sending $\\varphi(a+b\\epsilon)$ to $\\varphi(a)+D_b\\varphi(a)\\epsilon$ with $D_b$ the Gâteaux derivative, produces valid dual-valued vector and matrix norms that keep the real-field properties, and that the resulting dual-valued Ky Fan $p$-$k$-norm carries usable information about causal emergence. In particular, for a dual transition probability matrix $P=P_s+P_i\\epsilon$ fitted from time-series data, the infinitesimal part of $\\|P\\|_{(k,p)}$, maximized over $k$ with $p$ in $[1,2)$, identifies the optimal number of macro-states. The paper proves the norm identities, including $\\|A\\|_{(k,p)}=\\|\\sigma^k\\|_p$ for the vector of the first $k$ dual singular values, and it shows that dual effective information, the dual Schatten $p$-norm, and dynamical reversibility of a DTPM all peak exactly at permutation matrices for $1\\le p<2$. The numerical experiment on a dumbbell chain finds the peak at $k=5$, the known number of groups, supporting the claim.","pith_inferences":["Beyond the paper: if the maximizing-$k$ heuristic survives on more systems, dual-matrix norms could serve as a general model-order selection tool for Markov chains with slow mixing and no spectral gap, replacing subjective singular-value thresholds.","Beyond the paper: the same dual-continuation construction could be applied to other unitarily invariant norms or to complex dual matrices, with the Gâteaux-derivative formalism suggesting analogous identities for condition numbers or entropy-like quantities.","Beyond the paper: a direct testable extension is to compare the $k$ suggested by the infinitesimal part with metastability indicators on a family of random dumbbell-like chains with known cluster counts and no singular-value gap."],"forward_implications":["Optimal coarse-graining can be read off directly from the dual Ky Fan $p$-$k$-norm, without enumerating candidate coarse-grainings or locating a singular-value cutoff.","The identity $\\|A\\|_{(k,p)}=\\|\\sigma^k\\|_p$ gives a computable route: take a dual SVD of the DTPM, form the dual vector of its first $k$ singular values, and evaluate its dual vector $p$-norm.","For $1\\le p<2$, a DTPM reaches its maximum dual Schatten $p$-norm exactly when it is a permutation matrix, and dual effective information reaches its maximum under the same condition, so the norm and effective information agree on the most reversible macro-state.","The procedure can be repeated on other time-series data: fit $P_s$ and $P_i$, compute the infinitesimal part of $\\|P\\|_{(k,p)}$, and take the maximizing $k$ as the system's optimal classification number."],"supporting_citations":[{"why":"Supplies the Gâteaux derivative whose limit defines the infinitesimal part of every dual continuation.","marker":"[6]"},{"why":"Max formula for subdifferentials of convex functions converts the Gâteaux derivative into a maximization over the subdifferential.","marker":"[1]"},{"why":"Provides the subdifferential and Gâteaux derivative formulas for matrix norms used to derive the dual-valued norms.","marker":"[24]"},{"why":"Compact dual SVD is used to prove the equivalence between the dual Ky Fan p-k-norm and the dual vector p-norm of the first k singular values.","marker":"[26]"},{"why":"The prior singular-value-cutoff and Schatten p-norm theory of causal emergence is the baseline this paper extends and contrasts with.","marker":"[29]"},{"why":"Defines causal emergence and effective information, the phenomenon and quantity the application targets.","marker":"[11,12]"},{"why":"Defines dual Markov chains, the setting in which the dual transition probability matrix and its effective information are introduced.","marker":"[21]"},{"why":"The fact that a stochastic matrix with stochastic inverse is a permutation matrix underpins the characterization of dynamically reversible DTPMs.","marker":"[4]"}],"fun_headline_variants":["Dual norm's infinitesimal term finds optimal macro-state count","Peak of dual Ky Fan norm reveals causal emergence groups","Infinitesimal dual norm pinpoints macro-state number","Dual matrix norm's tiny part signals causal emergence","Optimal macro-states found via dual p-k-norm peak"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that the dual transition matrix fitted from time-series data faithfully reflects the system's causal structure, so that the $k$ maximizing the infinitesimal part of the dual Ky Fan $p$-$k$-norm reliably marks the best coarse-graining scale; this is an empirical heuristic demonstrated on one dumbbell chain, not a proven theorem.","fun_headline_variants_meta":{"raw":{"variants":["Dual norm's infinitesimal term finds optimal macro-state count","Peak of dual Ky Fan norm reveals causal emergence groups","Infinitesimal dual norm pinpoints macro-state number","Dual matrix norm's tiny part signals causal emergence","Optimal macro-states found via dual p-k-norm peak"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1495,"prompt_tokens":1070,"completion_tokens":425,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":686,"completion_tokens_details":{"reasoning_tokens":343}},"tokens_in":686,"tokens_out":425,"duration_ms":4183,"temperature":1.0,"reasoning_tokens":343,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:38:58.038476+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a Markov chain with a known number of metastable groups but no clear gap in its singular values; if the $k$ maximizing the infinitesimal part of $\\|P\\|_{(k,p)}$ over $p\\in[1,2)$ does not equal that known number, or if the peak moves when $p$ is varied within $[1,2)$ or when the least-squares fitting is perturbed, the central claim would be falsified.","supporting_citations":[{"cited_title":"Gˆateaux, Sur les Fonctionnelles Continues et les Fonctionnelles Analytiques , C.R","cited_arxiv_id":null,"evidence_quote":"Supplies the Gâteaux derivative whose limit defines the infinitesimal part of every dual continuation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Max formula for subdifferentials of convex functions converts the Gâteaux derivative into a maximization over the subdifferential."},{"cited_title":"W atson, Characterization of the Subdifferential of Some Matrix Norms , Linear Algebra and its Applications, 170 (1992), pp","cited_arxiv_id":null,"evidence_quote":"Provides the subdifferential and Gâteaux derivative formulas for matrix norms used to derive the dual-valued norms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Compact dual SVD is used to prove the equivalence between the dual Ky Fan p-k-norm and the dual vector p-norm of the first k singular values."},{"cited_title":"Qi and C","cited_arxiv_id":null,"evidence_quote":"Defines dual Markov chains, the setting in which the dual transition probability matrix and its effective information are introduced."},{"cited_title":"Ding and N","cited_arxiv_id":null,"evidence_quote":"The fact that a stochastic matrix with stochastic inverse is a permutation matrix underpins the characterization of dynamically reversible DTPMs."}],"review_version":1}