{"id":"7d59e1a3-1e89-47a2-81f8-57595c782ed4","arxiv_id":"2504.13018","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Spatial-sign covariance based sparse CCA is consistent for elliptical symmetric distributions at a rate that depends on sparsity, with the dimension entering only through a bias term.","lead":"This paper introduces SSCCA, a sparse canonical correlation analysis that replaces the usual covariance matrix with a spatial-sign covariance to stay accurate when data have heavy tails. Tests on simulated and real gene and fatty acid data suggest it can be more reliable than standard sparse CCA in some non-Gaussian settings.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 1's assumption that r^{-1} is sub-Gaussian is violated by all three simulation distributions (normal, t_3, mixture normal), so the theorem's guaranteed scope excludes the heavy-tailed settings the paper claims to cover.","rationale":"The reader's conditional verdict points at the unproved Lemma 1 as the weakest assumption. I agree with that locus, but the concern can be sharpened: it is not merely that the proof is omitted; the lemma's sub-Gaussian reciprocal assumption is incompatible with the very distributions simulated in the paper. For any elliptical radial density that behaves like r^{p-1} near zero, the reciprocal has a polynomial tail, which is not sub-Gaussian. This includes the Gaussian, t_3, and mixture-normal settings in Section 3. Therefore Theorem 1, as stated, does not logically cover the empirical claims it is used to support. The concern is not that the estimator is inconsistent; the method may work, and the cited preprint may contain a proof valid under weaker moment assumptions. But the burden is on the authors to state and verify the correct condition. Since the reader already recommended conditional acceptance and this critique reinforces that recommendation without proving the central result false, I keep the verdict unchanged rather than moving to reject. A useful next step is to re-derive the concentration step from the cited preprint under the moment conditions only, which would settle whether the mismatch is a fixable assumption issue or a genuine gap in the theorem.","tokens_in":17184,"tokens_out":13443,"duration_ms":143516,"concrete_test":"Read Lemma 6 and its proof in Lu and Feng (2025, arXiv:2503.03575) and isolate every step that invokes sub-Gaussianity of r^{-1}. Then re-derive the sup-norm bound for p hat S - Sigma under only the stated moment conditions E|r|^{-k} <= zeta (E|r|^{-1})^k for k = 2, 3, 4. If the resulting bound acquires a p-polynomial factor or requires a separate high-probability lower bound on ||X - mu||, test whether the scaled t_3 and the 10/0.8 mixture-normal satisfy such a condition. If they do not, Theorem 1's assumptions exclude the main simulation settings and the paper must either restrict its consistency claim or supply a new Lemma 1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1 is conditional on Lemma 1, whose proof is not included but cited from Lu and Feng (2025). The load-bearing issue is not merely that the proof is absent: Lemma 1's Assumption (ii) requires r^{-1} to be sub-Gaussian, and this condition fails for every distribution simulated in Section 3. In the model X = mu + r Gamma u, a standard Gaussian, a scaled multivariate t_3, and the 10/0.8 mixture-normal all have radial densities with f_r(epsilon) approximately epsilon^{p-1} near zero, so P(r^{-1} > t) = P(r < 1/t) approximately t^{-p}. For sufficiently large t, the polynomial tail t^{-p} dominates any exp(-ct^2), so no finite sub-Gaussian constant exists. Hence Lemma 1, and with it Theorem 1, cannot be applied to the normal, t_3, or mixture-normal cases used in the simulations; in particular, consistency under heavy-tailed elliptical distributions is not established for the settings where SSCCA is argued to outperform KSCCA and SCCA. If the cited proof actually uses only the moment conditions E(|r|^{-k}) <= zeta E(|r|^{-1})^k, then the paper should state that corrected condition; as written, the theoretical guarantee and empirical scope are mismatched.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SSCCA, a sparse CCA method that replaces sample covariance matrices with the spatial-sign covariance matrix p-hat-S under an elliptical symmetric model X = μ + rΓu. The method solves the penalized problem (3) with L1 penalties. The main theoretical result, Theorem 1, asserts that under the error bound of Lemma 1 and sparsity/gap assumptions, a local maximizer of (3) is consistent for the leading canonical correlation and canonical directions, with a rate involving sqrt(log p/n) + p^{-1/2}. Simulations compare SSCCA with KSCCA and SCCA under normal, t_3, and mixture-normal distributions and report superior performance of SSCCA under heavy tails; a nutrimouse data application is also presented.","tokens_in":17488,"tokens_out":7311,"duration_ms":70322,"significance":"If the theoretical claims were fully established, the paper would offer a robust, heavy-tailed-friendly sparse CCA method with a natural use of spatial signs, and the numerical comparisons would be of practical interest. The core idea is reasonable and the appendix proof of Theorem 1 is coherent conditional on Lemma 1. However, the theoretical foundation is incomplete: Lemma 1 is not proved and is cited from an unpublished preprint, and its stated sub-Gaussian condition on r^{-1} appears to exclude all three simulation distributions. The abstract's claim of an 'optimal rate' is also unsupported by any lower-bound or minimax analysis. These gaps are load-bearing for the paper's central consistency claim and for the claimed empirical support.","major_comments":[{"comment":"The consistency result in Theorem 1 is entirely conditional on Lemma 1, but Lemma 1 is not proved in this paper; the paper only cites Lu and Feng (2025), an unpublished preprint. Because Theorem 1's proof begins 'By Lemma 1', the main theoretical contribution cannot be verified from the manuscript. The paper should either include a complete proof of Lemma 1 or clearly state that Theorem 1 is conditional on an external result and provide the full argument in an accessible form.","section":"Section 2, Lemma 1 and proof of Theorem 1"},{"comment":"The requirement that r^{-1} is sub-Gaussian is violated by all three distributions simulated in Section 3. Under model (2), the radial density of a standard multivariate normal, a scaled t_3, and the 10/0.8 mixture normal all behave as f_r(ε) ≈ C ε^{p-1} near zero, so P(r^{-1} > t) = P(r < 1/t) ≈ C' t^{-p}, which is a polynomial tail rather than a sub-Gaussian tail. Therefore Lemma 1, and with it Theorem 1, cannot be applied to the normal, t_3, or mixture-normal cases used in the simulations. If the proof of Lemma 1 actually only requires the moment conditions E(|r|^{-k}) ≤ ζ{E(|r|^{-1})}^k for k=2,3,4, the paper should state the weaker condition; as written, the theoretical guarantee and the empirical scope are mismatched, so the claimed robustness under heavy tails is not supported by the theory.","section":"Section 2, Lemma 1, Assumption (ii)"},{"comment":"The paper claims the estimator converges at an 'optimal rate' but provides no minimax lower bound or any comparison with known optimal rates for sparse CCA. The rate in Theorem 1 contains an additional p^{-1/2} bias term relative to the rates in Mai and Zhang (2019) and Gao et al. (2017), and the condition τ_1^2(√(log p/n) + p^{-1/2}) → 0 is not shown to be minimax optimal. The 'optimal' claim should be removed or replaced by a precise statement about the rate achieved under the stated assumptions.","section":"Abstract and Section 1"}],"minor_comments":[{"comment":"The inequality direction in Lemma 2 is reversed in both statements: it states max_j cos^2(...) ≤ 1 - 2δ/γ, but the proof shows that each cos^2(...) ≥ 1 - 2δ/γ, and the second bound should also be '≥' with the factor 4κδ/γ. The proof uses the correct direction, so this is a typographical error; please correct the lemma statements.","section":"Appendix, Lemma 2"},{"comment":"The probability bound in Lemma 1 is written as C√(log p + log(α^{-1/2}) / n + C/√p) with ambiguous parentheses; the intended bound appears to be C√({log p + log(1/α)}/n) + C/√p. Please clarify the notation.","section":"Section 2, Lemma 1 statement"},{"comment":"The consistency condition is written as τ_1^2 (√(log p)/n + p^{-1/2}) → 0, which is a typo: the first term should be √(log p / n) (or equivalently √(log p)/√n). Please correct the expression.","section":"After Theorem 1"},{"comment":"The definitions of FPR_g and FNR_g appear swapped: the formula for FPR_g is actually the false negative rate among nonzero coefficients, and the formula for FNR_g is the false positive rate among zero coefficients. Please correct the labels or the formulas.","section":"Section 3, definitions of FPR and FNR"},{"comment":"The reference 'Gonz´alez et al. (2008)' is listed twice with identical details; remove the duplicate.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's central lemma is cited from the authors' own unpublished preprint. If the proof of Lemma 1 is not supplied, the theoretical contribution cannot be evaluated. The mismatch between Lemma 1's assumptions and the simulation distributions is a serious correctness-risk concern. The 'optimal rate' claim in the abstract should be toned down unless a lower bound is added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper proposes a spatial-sign covariance-based sparse CCA (SSCCA) for elliptical data. The core idea is sensible and the writing is clear. What is genuinely new is the use of the spatial-sign covariance matrix inside a sparse CCA framework, together with a consistency theorem that adapts the proof strategy of Yoon et al. (2020) and Mai and Zhang (2019). That combination does not appear in the prior literature, and the proof of Theorem 1 itself is coherent and follows the established template.\n\nThe soft spot is load-bearing. Theorem 1 rests on Lemma 1, which is not proved in this paper but cited from Lu and Feng (2025), an unpublished preprint. That is already a concern, but the deeper problem is the lemma's condition that r^{-1} is sub-Gaussian. For the three simulation settings—multivariate normal, multivariate t_3, and mixture normal—the radial density behaves like epsilon^{p-1} near zero, so P(r^{-1} > t) = P(r < 1/t) ~ t^{-p}. That is a polynomial tail, not sub-Gaussian, so Lemma 1 cannot be applied to any of the distributions used in Section 3. The theorem's guaranteed scope therefore excludes exactly the heavy-tailed cases where the paper argues SSCCA outperforms KSCCA and SCCA. If the cited preprint actually proves Lemma 1 under only the moment conditions stated earlier in the lemma, the paper should say so; as written, the theory and simulations are mismatched.\n\nThe abstract's claim of an 'optimal rate' is also unsupported, since no lower bound is given. The simulation narrative is somewhat stronger than the tables. Under Model I, SSCCA does usually beat KSCCA in heavy tails, but under Model II (the approximately low-rank case) SSCCA is sometimes worse than KSCCA, for example in Table 4 for n=200, p=400 with t_3. That should be acknowledged. Minor: the reference list has a duplicated Gonz\\'alez et al. (2008) entry.\n\nAll that said, the paper deserves a serious referee. The proposed estimator is new, the theoretical framework is plausible, and the missing lemma is checkable in an external preprint. A good referee should push the authors to either repair the sub-Gaussian condition or present simulations that satisfy the assumptions, and to tone down the optimal-rate and dominance claims. With those fixes, the paper would be a useful contribution to robust sparse CCA.","headline":"The core idea is reasonable but the main theorem's key error bound assumes r^{-1} is sub-Gaussian, which fails for every distribution simulated, leaving the theoretical guarantee and the experiments mismatched.","tokens_in":17959,"tokens_out":2462,"would_cite":false,"duration_ms":25726,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H20","62H12","62F35"],"pacs":[],"model":"deepseek-v4-flash","headline":"Spatial-sign covariance keeps sparse CCA consistent under heavy tails.","keywords":["canonical correlation analysis","spatial-sign covariance","elliptical distributions","high-dimensional statistics","robust estimation","sparsity","heavy tails"],"falsifier":"Simulate elliptical data with the same covariance structure and check the empirical distribution of $\\|p\\widehat{S}-\\Sigma\\|_\\infty$ against the Lemma 1 bound; a violation at the claimed probability, or an inconsistent estimate of the leading direction as $n,p$ grow along paths where $\\tau_1^2(\\sqrt{\\log p}/n+p^{-1/2})\\to 0$, would refute Theorem 1.","tokens_in":16989,"feed_emoji":"","tokens_out":5515,"duration_ms":53890,"temperature":0.7,"pith_summary":"This paper claims that a sparse canonical correlation analysis built on the spatial-sign covariance matrix, rather than the sample covariance, remains consistent for elliptical symmetric distributions with heavy tails. Classical CCA requires finite second moments and breaks down in high dimensions; the proposed SSCCA replaces the sample covariance with a sign-based scatter matrix and adds an $\\ell^1$ penalty. If the main theorem is right, the leading canonical correlation and the direction vectors are recovered with high probability under a mild sparsity condition, with a convergence rate that differs from prior work only by a $p^{-1/2}$ bias term coming from the sign approximation. The practical upshot, supported by simulations, is that SSCCA matches standard methods under normality and beats them under $t$ or mixture distributions.","feed_headline":"Spatial-sign covariance keeps sparse CCA consistent under heavy tails","feed_subtitle":"Replacing sample covariance with a sign-based estimator preserves consistent direction recovery without finite moments.","key_machinery":"The carrying object is the spatial-sign covariance matrix $S = E[U(X-\\mu)U(X-\\mu)^\\top]$ with $U(x)=x/\\|x\\|$. For elliptical symmetric $X=\\mu+r\\Gamma u$, the population covariance obeys $\\Sigma \\approx \\operatorname{tr}(\\Sigma) S$ as dimension grows, and because scaling a linear combination does not change correlation, the unknown trace can be ignored; the method solves problem (3) with $p\\widehat{S}$ in place of the sample covariance. The proof then chains a uniform error bound for $p\\widehat{S}-\\Sigma$ through lower and upper bounds on the penalized objective at a normalized truth, producing an approximate maximizer on an $\\ell^1$ ball, and converts the canonical-correlation gap assumption into cosine-angle closeness. The $\\ell^1$ penalties enforce sparsity while the sign estimator absorbs heavy tails.","core_discovery":"The paper's central claim is Theorem 1: for data from an elliptical symmetric distribution, solving the $\\ell^1$-penalized problem with $p\\widehat{S}$ in place of the sample covariance produces a local maximizer whose cosine-squared angle with the true leading canonical direction tends to 1, and whose estimated canonical correlation approaches $\\rho_{1,*}$, provided $\\tau_1^2(\\sqrt{\\log p}/n + p^{-1/2}) \\to 0$. In other words, heavy tails need not be transformed away or modeled explicitly: the spatial-sign transformation alone stabilizes the cross-covariance enough for sparse CCA to be consistent. The paper also highlights the price of this robustness: a $p^{-1/2}$ approximation bias appears in the rate because $p\\widehat{S}$ approximates $\\Sigma$ only as the dimension grows.","pith_inferences":["A natural extension would replace the $\\ell^1$ penalty with group or fused penalties; the spatial-sign formulation is agnostic to the penalty choice, so the same proof template likely carries over when the sparsity assumption is replaced by a group-sparsity norm.","The $p^{-1/2}$ bias term implies SSCCA may be systematically less accurate when $p$ is small relative to $n$; a simulation scan with fixed $p$ and increasing $n$ could test where the sign approximation starts to hurt.","Because the spatial-sign covariance depends only on directions, the method may tolerate missing or clipped values that preserve sign, a robustness property worth testing explicitly but not claimed in the paper.","The consistency statement is for a local maximizer; an editorial guess is that the same proof route could be tightened to a global maximizer under stronger concavity or initialization conditions, though the paper does not address this."],"forward_implications":["If Theorem 1 is correct, SSCCA gives a consistency guarantee for sparse CCA without requiring Gaussian or sub-Gaussian tails, so high-dimensional genomics or financial data can be analyzed directly under an elliptical model.","The convergence rate matches earlier sparse CCA rates whenever $p\\log p/n \\to \\infty$, so in the usual high-dimensional regime the sign approximation bias disappears and no asymptotic efficiency is lost.","Simulations indicate that under heavy-tailed $t$ and mixture-normal distributions, SSCCA has lower estimation error and prediction loss than Kendall-tau and sample-covariance sparse CCA, and it selects sparser and more accurate support.","Because the method reuses the same convex optimization algorithm as existing sparse CCA solvers, the robustness improvement can be obtained without new computational machinery.","In the nutrimouse application, SSCCA achieves a higher out-of-sample canonical correlation while selecting fewer genes and fatty acids, suggesting a stronger per-variable predictive signal."],"supporting_citations":[{"why":"Supplies Lemma 1, the uniform error bound for $p\\widehat{S}-\\Sigma$, and the underlying approximation $\\Sigma\\approx\\operatorname{tr}(\\Sigma)S$; the theorem's consistency proof depends on this bound.","marker":"Lu and Feng (2025)"},{"why":"Provides the convex optimization algorithm, the BIC-type tuning criteria, and the Kendall-tau comparison method used in simulations.","marker":"Yoon et al. (2020)"},{"why":"Supplies Lemma 2, which converts a correlation gap into cosine-angle closeness between estimated and true directions, and gives a baseline sparse CCA approach.","marker":"Mai and Zhang (2019)"},{"why":"Defines the spatial-sign and spatial-sign covariance matrix that is the core estimator used throughout the paper.","marker":"Oja (2010)"},{"why":"Provides the prediction loss metric and a sparse CCA method used as context and comparison for estimation accuracy.","marker":"Gao et al. (2017)"},{"why":"Introduces the penalized matrix decomposition sparse CCA, which the paper cites as an early method that can be inconsistent when covariance is far from diagonal.","marker":"Witten et al. (2009)"}],"fun_headline_variants":["Sign-based sparse CCA stays consistent under heavy tails","Spatial-sign covariance: robust sparse CCA without finite moments","Heavy-tailed sparse CCA solved by sign covariance estimator","No moment assumptions: spatial-sign CCA achieves consistency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The theorem rests on Lemma 1's uniform error bound for $p\\widehat{S}-\\Sigma$, whose proof is not given in this paper but cited from a companion preprint; if that bound, or the underlying approximation $\\Sigma\\approx\\operatorname{tr}(\\Sigma)S$ with its $p^{-1/2}$ bias, fails for the actual data-generating process, the consistency claim collapses.","fun_headline_variants_meta":{"raw":{"variants":["Sign-based sparse CCA stays consistent under heavy tails","Spatial-sign covariance: robust sparse CCA without finite moments","Heavy-tailed sparse CCA solved by sign covariance estimator","No moment assumptions: spatial-sign CCA achieves consistency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1394,"prompt_tokens":865,"completion_tokens":529,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":462}},"tokens_in":481,"tokens_out":529,"duration_ms":5890,"temperature":1.0,"reasoning_tokens":462,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:17:12.557892+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate elliptical data with the same covariance structure and check the empirical distribution of $\\|p\\widehat{S}-\\Sigma\\|_\\infty$ against the Lemma 1 bound; a violation at the claimed probability, or an inconsistent estimate of the leading direction as $n,p$ grow along paths where $\\tau_1^2(\\sqrt{\\log p}/n+p^{-1/2})\\to 0$, would refute Theorem 1.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the convex optimization algorithm, the BIC-type tuning criteria, and the Kendall-tau comparison method used in simulations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Lemma 2, which converts a correlation gap into cosine-angle closeness between estimated and true directions, and gives a baseline sparse CCA approach."},{"cited_title":"Ma, and H","cited_arxiv_id":null,"evidence_quote":"Provides the prediction loss metric and a sparse CCA method used as context and comparison for estimation accuracy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the penalized matrix decomposition sparse CCA, which the paper cites as an early method that can be inconsistent when covariance is far from diagonal."}],"review_version":1}