Pith. sign in

REVIEW 3 major objections 4 minor 5 cited by

Estimating shared subspace with AJIVE: the power and limitation of multiple data matrices

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper establishes a two-regime theory for AJIVE shared-subspace estimation: high signal-to-noise gives a minimax-optimal $1/\sqrt K$ error decay, while low signal-to-noise leaves a non-diminishing error that also plagues an…

desk verdict Solid high-SNR minimax theory for AJIVE; the low-SNR 'fundamental barrier' is honestly labeled in the body but overclaimed in the abstract. read the letter →

arxiv 2501.09336 v2 pith:Z3I3XEZX submitted 2025-01-16 stat.ML cs.LGmath.STstat.TH

classification stat.MLcs.LGmath.STstat.TH MSC 62H2562C2062F12
keywords JIVEAJIVEsharedsubspaceestimationmulti-viewdataintegrationspectralmethodminimaxlowerboundmisalignmentincidentalparameterproblem
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Integrative data analysis frequently asks whether several noisy matrices share one low-dimensional signal subspace, with each matrix carrying its own private variation on top. This paper analyzes AJIVE, a two-stage spectral method for estimating that shared subspace, and establishes a two-regime picture. When the noise is weak enough, $\sigma\sqrt n/\sigma_{\min} \ll \min\{\sqrt\theta,\sqrt{K\theta}\}$, the estimation error decays like $(\sigma/\sigma_{\min})\sqrt{n/K + r/(K\theta)}$, matching a minimax lower bound, so pooling more matrices genuinely helps. When the noise is strong, a second-order bias term $\frac{1}{\theta(1\wedge K\theta)}\frac{\sigma^2 n}{\sigma_{\min}^2}$ does not vanish as $K\to\infty$, and the paper shows the same non-vanishing error occurs for an oracle-aided spectral estimator. The result matters because it circumscribes when multi-view integration can and cannot improve shared-subspace recovery.

What carries the argument

The load-bearing object is the averaged projection matrix $M := \frac{1}{K}\sum_{k=1}^K \widehat U_k \widehat U_k^\top$, formed from per-matrix estimates of the top $(r+r_k)$ left singular subspaces; AJIVE outputs the top-$r$ eigenspace of $M$. The analysis separates the noiseless leading term $U^\star U^{\star\top} + \frac{1}{K}\sum_k U_k^\star U_k^{\star\top}$ from a perturbation $\Delta$ built from first-order SVD expansions, and the misalignment parameter $\theta$ controls the eigengap of that leading term. For the oracle lower bound the machinery is a fourth-order polynomial approximation $Q$ of the oracle matrix $M$, whose expectation contains a cross term $\alpha_4 \frac{1}{K}\sum_k (U^\star U_k^{\star\top} + U_k^\star U^{\star\top})$ with $\alpha_4 \asymp \sigma^4 nd$; because $V^{\star\top}W^\star = 0.6I$ in the hard configuration, this cross term tilts the leading eigenspace away from $U^\star$ and the tilt does not shrink as $K$ grows.

What would settle it

Run a low-SNR sequence with $K$ from 25 to 10,000 on the paper's hard configuration ($V_k^{\star\top}W_k^\star = 0.6I$, $\sigma\sqrt n/\sigma_{\min}$ small, $r\le n/3$) and test a full joint nonconvex estimator, not the oracle-initialized one-step estimator: if its error converges to 0, the claimed fundamental barrier is false. Since the paper explicitly lacks an information-theoretic lower bound in this regime, a minimax calculation that forces positive error for all estimators would be the definitive confirmation.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a complete first-order characterization of AJIVE for the JIVE model $A_k = U^\star V_k^{\star\top} + U_k^\star W_k^{\star\top} + E_k$ with i.i.d. Gaussian noise of variance $\sigma^2$. Theorem 1 (simplified at equation (8)) says that under the high-SNR condition $\sigma\sqrt n/\sigma_{\min} \ll \min\{\sqrt\theta,\sqrt{K\theta}\}$ the estimation error $\|\widehat U\widehat U^\top - U^\star U^{\star\top}\|$ is bounded by $(\sigma/\sigma_{\min})(\sqrt{n/K} + \sqrt{r/(K\theta)})$ plus a second-order term; Theorem 2 shows the first-order rate is minimax optimal, with the rate separating a large-$\theta$ regime where unique subspaces barely interfere from a small-$\theta$ regime where their alignment degrades the rate. In the complementary low-SNR regime the second-order term dominates and converges to $\frac{1}{\theta}\frac{\sigma^2 n}{\sigma_{\min}^2}$ as $K\to\infty$. Theorem 3 constructs a worst-case loading configuration $V_k^{\star\top}W_k^\star = 0.6I$ on which even the oracle estimator, namely one alternating-minimization step initialized at the true $U^\star$, has error bounded below by $C\sigma^4 n^2/\sigma_{\min}^4$, so the non-diminishing error is presented as a structural barrier rather than an artifact of AJIVE.

Load-bearing premise

The load-bearing premise is that the oracle-aided spectral estimator, namely one alternating-minimization step initialized at the true shared subspace, represents the fundamental low-SNR limit; the paper states it cannot prove an information-theoretic lower bound, so any estimator outside this spectral class that drives the error to zero would overturn the barrier conclusion.

Editorial extensions

If this is right

  • In the high-SNR regime, averaging more matrices improves shared-subspace estimation at the minimax-optimal rate $1/\sqrt K$, so additional views give real accuracy gains.
  • The optimal rate worsens as the unique subspaces become more aligned: the $r/(K\theta)$ term shows that near-identifiability, not noise alone, controls difficulty.
  • In the low-SNR regime, AJIVE's error has a component $\frac{1}{\theta}\frac{\sigma^2 n}{\sigma_{\min}^2}$ that survives $K\to\infty$; beyond a threshold the number of matrices is irrelevant to the error floor.
  • An oracle-aided spectral estimator, given the true unique components in the first stage, still fails on a worst-case alignment, indicating the floor is shared by spectral methods rather than a quirk of AJIVE.
  • The persistent error fits the classical incidental-parameter phenomenon: the shared subspace is the structural parameter and the per-matrix unique subspaces are incidental parameters that contaminate likelihood-type estimation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A likely testable extension is that full alternating minimization, which re-estimates $U$ jointly with the unique components rather than stopping after one step from the truth, may behave differently from the oracle estimator in the low-SNR regime; the paper only analyzes the one-step version.
  • Because the theory assumes equal noise variance across matrices and known ranks, an immediate extension would allow per-matrix $\sigma_k^2$; the second-stage average would then need inverse-variance weighting, which may shift or remove the non-diminishing term.
  • The hard configuration uses adversarial loadings with $V^{\star\top}W^\star = 0.6I$; the numerical experiments show the low-SNR floor under shared loading while random loadings show the $1/\sqrt K$ decay, suggesting that alignment between $V_k^\star$ and $W_k^\star$, not just subspace misalignment $\theta$, controls whether the floor appears.
  • A practical diagnostic suggested by the theory: estimate $\theta$ and the SNR before deciding whether to pool more matrices; when $\theta$ and SNR are both small, the shared subspace should be treated as only partially identified.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies the AJIVE algorithm for estimating the shared column subspace in the JIVE model with K noisy data matrices. The main theoretical contribution is a two-regime characterization: in a high-SNR regime, the estimation error of AJIVE is shown to decay as σ/σmin times roughly sqrt(n/K) plus a misalignment-dependent term, and a matching minimax lower bound is established; in a low-SNR regime, a non-diminishing error term is identified. Since the paper does not prove an information-theoretic lower bound for this low-SNR barrier, it instead proves a lower bound for an oracle-aided spectral estimator (Theorem 3) and presents numerical evidence. The paper also contains a detailed proof appendix based on SVD perturbation expansions, concentration inequalities, and Fano/packing arguments.

Significance. The high-SNR results are a solid contribution: Theorem 1 provides a nearly complete rate characterization of AJIVE, and Theorem 2 shows the first-order rate is minimax-optimal. The dependence on the misalignment parameter θ and the explicit benefit of multiple matrices (1/sqrt(K) rate) are new and quantitatively useful. The paper also ships detailed, self-contained proofs, including a nontrivial Fano construction for the θ-dependent term that may be of independent interest. The low-SNR part is suggestive but not at the same level of rigor: Theorem 3 is an algorithm-specific lower bound for an oracle-initialized spectral estimator, not a minimax lower bound, and the paper's own caveats in Sections 5 and 7 are more accurate than the abstract. If the claims are appropriately qualified, the paper is a valuable contribution to the theory of multi-matrix subspace estimation.

major comments (3)
  1. [Section 3, Eqs. (6)-(8)] The abstract and Section 1 state that in low-SNR settings AJIVE exhibits a non-diminishing error as K grows, but Theorem 1 does not prove this. Theorem 1 is an upper bound that holds only under the high-SNR conditions (6a)-(6b); when σ√n/σmin is not ≪ √θ, those conditions fail and the bound, including the simplified expression (8), is not a valid consequence. The discussion immediately after (8) interprets the second-order term E2 as a low-SNR non-diminishing error, but this is an extrapolation outside the theorem's assumptions. Please either prove a low-SNR lower bound for AJIVE itself or explicitly label this statement as a conjecture supported by the oracle-estimator analysis and the simulations.
  2. [Section 5, Theorem 3] The phrase 'fundamental barrier' in the abstract and in the contribution list is stronger than what Theorem 3 establishes. Theorem 3 is a lower bound for the oracle-aided spectral estimator (Algorithm 2) on a specially constructed configuration with V⋆ᵀW⋆ = 0.6I; it is not a minimax lower bound over all estimators, and it does not rule out non-spectral methods that could have vanishing error as K grows. The paper itself concedes this in Section 5 ('we are unable to deliver an information-theoretic limit') and Section 7 ('might be a fundamental limit'). The abstract, Section 1, and the concluding remarks should be aligned with these qualifications; the current wording overstates the strength of Theorem 3.
  3. [Section 5.3] The connection to Neyman and Scott's problem is presented as if the oracle estimator being 'one step of alternating minimization starting from the ground truth' makes it representative of the MLE. This is an analogy, not a proof: Algorithm 2 is neither the nonconvex least-squares estimator nor a full alternating-minimization run, and the paper does not show inconsistency of the MLE itself. The text should clearly state that the Neyman-Scott connection is heuristic, or provide an analysis of the actual nonconvex estimator.
minor comments (4)
  1. [Abstract] The abstract's phrase 'a fundamental barrier' is inconsistent with Section 7's accurate wording 'might be a fundamental limit'; please use the qualified version throughout.
  2. [Section D.3.1] The statement 'Without loss of generality, assume n is divisible by 3 and K is even' is not a true WLOG; rounding or a padding argument should be supplied, or the assumption should be stated as a technical condition.
  3. [Section C.1.1] In the proof of (22) and (23), the sentence 'Moreover, this truncation is tight...' is followed by a probability statement for a deterministic expectation; the phrase 'with probability at least 1 - O(N^-11)' around (23b) is misleading since (23b) is an exact identity, not a high-probability event.
  4. [Section 6, Figure 3] The figure caption states the y-axis is in log scale but does not specify whether the x-axis is √K or log √K; adding the axis scaling in the caption would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the derivation chains are independent; the low-SNR 'fundamental barrier' is an overclaim, but not a circular step.

full rationale

This paper is a proof-based analysis and does not derive its conclusions from its own inputs by construction. The central upper-bound result (Theorem 1) is established through a first-order SVD perturbation expansion attributed to the external result [Xia21], followed by matrix concentration inequalities and a spectral perturbation argument. The minimax lower bound (Theorem 2) uses Fano's method with explicit packing constructions, and the constants and rates are obtained from the constructions, not from the target upper bound. Theorem 3 is explicitly an algorithm-specific lower bound for the oracle-aided spectral estimator; the paper honestly states that it cannot prove an information-theoretic limit ('we are unable to deliver an information-theoretic limit') and only conjectures that the non-diminishing error 'might be a fundamental limit.' That the abstract uses the stronger phrase 'fundamental barrier' is a scope/correctness concern about overclaiming, not a circularity. The self-citation to [CCFM21] (a monograph coauthored by Cong Ma) is used for standard truncated matrix Bernstein inequalities and a geometric lemma; these are established published results with stated assumptions, not special premises whose validity depends on the present paper's claims. The simulations generate θ-misaligned subspaces via the randomized construction in Section 2.2, but that construction is a direct implementation of Definition 1, not a fitted parameter renamed as a prediction. No fitted quantity is later called a prediction, and no uniqueness theorem from the authors' prior work is used to force the choice of estimator. Therefore the derivation chain is self-contained against its declared inputs, and no circular step can be exhibited with the specificity required by the review rules.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The central claims rest on the Gaussian homoscedastic noise model, exact rank knowledge, and the θ-misalignment identifiability condition; the SVD perturbation machinery of Xia (2021) is imported as background. The only hand-chosen constants are construction constants in the lower bound (0.6) and a simulation knob (γ = 0.5); no parameters are fitted to data in the theorems.

free parameters (2)
  • cross-loading constant 0.6 in the Theorem 3 configuration = 0.6
    The oracle lower bound picks V^TW = 0.6I to make the shared-unique cross term survive in E[Q] (Appendix E, 'Specifying the configuration'). Any value in (0,1) would work, so it is a hand-chosen construction constant, not fitted to data.
  • unique-component signal scale gamma = 0.5
    In the simulations, W⋆k is scaled by 1/γ with γ = 0.5 to control the strength of the unique components relative to the shared component (Section 6). This is a simulation knob, not used in the theorems.
assumptions (6)
  • domain assumption Noise entries are i.i.d. Gaussian with common variance σ² across all K matrices
    Section 2.1, equation (3). All rates and the averaging step in AJIVE stage 2 assume equal noise variance; heteroscedastic noise is not covered and would change the averaging argument.
  • domain assumption Ranks r and rk are known exactly
    Algorithm 1 takes r and {rk} as inputs, and the first-stage SVD retains exactly r + rk singular vectors. Rank misspecification is outside the analysis.
  • domain assumption The unique subspaces are θ-misaligned with θ > 0, so ∩k col(U⋆k) = ∅
    Definition 1 and Section 2.2. θ = 0 makes the shared subspace unidentifiable, and the rates carry 1/θ or 1/(Kθ) factors, so the bounds degrade as misalignment decreases.
  • standard math The SVD perturbation expansion of Xia (2021), Theorem 1, is valid and its remainder is controlled by conditions (6a)-(6b)
    Appendix C, equations (25)-(26). The first-order approximations of the subspace errors rest on this expansion with ∥Δk∥ ≤ σ²min/8.
  • standard math Standard concentration and information-theoretic tools: matrix Bernstein, Vershynin bounds, Gilbert-Varshamov, Fano's method
    Used throughout Appendices C, D, and E; standard results, cited ([Ver18], [CCFM21], [JJ12], [Yu97]).
  • ad hoc to paper The oracle spectral estimator (one step of alternating minimization from the truth) is representative of achievable low-SNR performance
    Section 5, Algorithm 2, Theorem 3. This is the load-bearing premise for the 'fundamental barrier' interpretation; the paper states it cannot prove the information-theoretic version of this claim.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Estimating shared subspace with AJIVE: the power and limitation of multiple data matrices." pith.science (2026). https://pith.science/paper/Z3I3XEZX

@misc{pith2026250109336,
  author       = {Pith},
  title        = {Pith review of: Estimating shared subspace with AJIVE: the power and limitation of multiple data matrices},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z3I3XEZX}},
  note         = {Machine review of arXiv:2501.09336}
}
read the original abstract

Integrative data analysis often requires disentangling joint and individual variations across multiple datasets, a challenge commonly addressed by the Joint and Individual Variation Explained (JIVE) model. While numerous methods have been developed to estimate the shared subspace under JIVE, the theoretical understanding of their performance remains limited, particularly in the context of multiple matrices and varying degrees of subspace misalignment. This paper bridges this gap by providing a systematic analysis of shared subspace estimation in multi-matrix settings. We focus on the Angle-based Joint and Individual Variation Explained (AJIVE) method, a two-stage spectral approach, and establish new performance guarantees that uncover its strengths and limitations. Specifically, we show that in high signal-to-noise ratio (SNR) regimes, AJIVE's estimation error decreases with the number of matrices, demonstrating the power of multi-matrix integration. Conversely, in low-SNR settings, AJIVE exhibits a non-diminishing error, highlighting fundamental limitations. To complement these results, we derive minimax lower bounds, showing that AJIVE achieves optimal rates in high-SNR regimes. Furthermore, we analyze an oracle-aided spectral estimator to demonstrate that the non-diminishing error in low-SNR scenarios is a fundamental barrier. Extensive numerical experiments corroborate our theoretical findings, providing insights into the interplay between SNR, the number of matrices, and subspace misalignment.

Figures

Figures reproduced from arXiv: 2501.09336 by the authors.

Figure 1
Figure 1. Estimation error when θ = 1/2. The parameters are chosen to be d = 20, r = 2, σ = 0.001, and γ = 0.5. The loading matrices use the random generation scheme. Each point is an average of 100 trials. Results. Theorem 1 shows that when r ≍ ravg ≍ 1, AJIVE achieves estimation error [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Estimation error when θ is small. The loading matrices use the random generation scheme. For all three subfigures, the parameters are d = 20, r = 2, σ = 10−6 , and γ = 0.5. In (a), K = 100, θ = 0.0001 and n ranges from 16 to 400. In (b), n = 20, θ = 0.0001, and K ranges from 25 to 10000. In (c), n = 20, K = 100, and θ ranges from 0.01 to 0.0001. Each point is an average of 100 trials. 20 40 √ 60 80 100 K 0.016 0.008… view at source ↗
Figure 3
Figure 3. Estimation error ∥Ub Ub ⊤ − U⋆U⋆⊤∥ vs √ K using random vs shared loading generation scheme. The parameters are chosen to be n = d = 20, r = 2, γ = 0.5, and K ranges from 25 to 10000. Each point is an average of 100 trials. The y-axis is in log scale. 12 [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Estimation error when θe = 1/2. The parameters are chosen to be n = 20, r = 2, K = 100, and γ = 0.5. The loading matrices use the random generation scheme. Each point is an average of 100 trials. Lemma 1. Let E be a n1 × n2 matrix where each entry Eij is a zero-mean Ga…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stacked SVD or SVD stacked? A Random Matrix Theory perspective on data integration

    stat.ML 2025-07 accept novelty 8.0 of 10

    In the proportional high-dimensional limit, the paper derives exact squared-overlap formulas and phase transitions for Stack-SVD and SVD-Stack, and proves optimally weighted Stack-SVD always beats optimally weighted S...

  2. Transfer Learning in High-Dimensional Clustering: Minimax Thresholds and Applications in Single-Cell Data

    math.ST 2026-07 conditional novelty 7.5 of 10

    In high-d two-community GMMs, consistent target clustering via transfer is possible iff either the target SNR clears the usual (d/n)^{1/4} barrier or the source is strong and aligned enough that µ∆_T, ∆_S, and µ∆_S∆_T...

  3. Unveiling Invariant and Transferable Latent Factors Across Heterogeneous Environments via ATLAS

    math.ST 2026-07 conditional novelty 7.0 of 10

    ATLAS is a multi-environment factor model estimator that separates invariant (shared-loading) factors from environment-specific factors and uses auxiliary labels to align and transfer prediction-relevant signals.

  4. Spectral Joint Subspace Estimation for Heterogeneous Multi-View Data: Geometry and Reweighting

    math.ST 2025-12 conditional novelty 7.0 of 10

    Geometry of the individual components, not the number of views, determines whether AJIVE's error vanishes at the K^{-1/2} rate; a weighted variant handles heterogeneous views.

  5. A functional tensor model for dynamic multilayer networks with common invariant subspaces and the RKHS estimation

    stat.ME 2025-09 unverdicted novelty 6.0 of 10

    A functional tensor model with common invariant subspaces and RKHS estimation for analyzing dynamic multilayer networks, applied to bike-share and food-trade data.

Reference graph

Works this paper leans on

29 extracted references · 24 canonical work pages · cited by 5 Pith papers

  1. [1]

    Spectral methods for data science: A statistical perspective

    Yuxin Chen, Yuejie Chi, Jianqing Fan, and Cong Ma. Spectral methods for data science: A statistical perspective. Foundations and Trends in Machine Learning , 14(5):566--806, 2021

  2. [2]

    Vincent Poor, and Yuxin Chen

    Changxiao Cai, Gen Li, Yuejie Chi, H. Vincent Poor, and Yuxin Chen. Subspace estimation from unbalanced and incomplete data matrices: _ 2, statistical guarantees. Ann. Statist. , 49(2):944--967, 2021

  3. [3]

    Tony Cai, Zongming Ma, and Yihong Wu

    T. Tony Cai, Zongming Ma, and Yihong Wu. Sparse PCA : optimal rates and adaptive estimation. Ann. Statist. , 41(6):3074--3110, 2013

  4. [4]

    Tony Cai and Anru Zhang

    T. Tony Cai and Anru Zhang. Rate-optimal perturbation bounds for singular subspaces with applications to high-dimensional statistics. Ann. Statist. , 46(1):60--89, 2018

  5. [5]

    Detection limits in the high-dimensional spiked rectangular model

    Ahmed El Alaoui and Michael I Jordan. Detection limits in the high-dimensional spiked rectangular model. In Conference On Learning Theory , pages 410--438. PMLR, 2018

  6. [6]

    Angle-based joint and individual variation explained

    Qing Feng, Meilei Jiang, Jan Hannig, and JS Marron. Angle-based joint and individual variation explained. Journal of multivariate analysis , 166:241--265, 2018

  7. [7]

    Distributed estimation of principal eigenspaces

    Jianqing Fan, Dong Wang, Kaizheng Wang, and Ziwei Zhu. Distributed estimation of principal eigenspaces. Ann. Statist. , 47(6):3009--3031, 2019

  8. [8]

    Structural learning and integrative decomposition of multi-view data

    Irina Gaynanova and Gen Li. Structural learning and integrative decomposition of multi-view data. Biometrics , 75(4):1121--1132, 2019

Show all 29 references
  1. [9]

    Covariate-driven factorization by thresholding for multiblock data

    Xing Gao, Sungwon Lee, Gen Li, and Sungkyu Jung. Covariate-driven factorization by thresholding for multiblock data. Biometrics , 77(3):1011--1023, 2021

  2. [10]

    On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables

    Leon Isserlis. On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables. Biometrika , 12(1/2):134--139, 1918

  3. [11]

    Information and coding theory

    Gareth A Jones and J Mary Jones. Information and coding theory . Springer Science & Business Media, 2012

  4. [12]

    The incidental parameter problem since 1948

    Tony Lancaster. The incidental parameter problem since 1948. Journal of econometrics , 95(2):391--413, 2000

  5. [13]

    Joint and individual variation explained ( JIVE ) for integrated analysis of multiple data types

    Eric F Lock, Katherine A Hoadley, James Stephen Marron, and Andrew B Nobel. Joint and individual variation explained ( JIVE ) for integrated analysis of multiple data types. The annals of applied statistics , 7(1):523, 2013

  6. [14]

    P. W. MacDonald, E. Levina, and J. Zhu. Latent space models for multiplex networks with shared structure. Biometrika , 109(3):683--706, 2022

  7. [15]

    Optimal estimation of shared singular subspaces across multiple noisy matrices

    Zhengchi Ma and Rong Ma. Optimal estimation of shared singular subspaces across multiple noisy matrices. arXiv preprint arXiv:2411.17054 , 2024

  8. [16]

    Consistent estimates based on partially consistent observations

    Jerzy Neyman and Elizabeth L Scott. Consistent estimates based on partially consistent observations. Econometrica: journal of the Econometric Society , pages 1--32, 1948

  9. [17]

    Fourth-order properties of normally distributed random matrices

    Heinz Neudecker and Tom Wansbeek. Fourth-order properties of normally distributed random matrices. Linear Algebra and its Applications , 97:13--21, 1987

  10. [18]

    Data integration via analysis of subspaces ( DIVAS )

    Jack Prothero, Meilei Jiang, Jan Hannig, Quoc Tran-Dinh, Andrew Ackerman, and JS Marron. Data integration via analysis of subspaces ( DIVAS ). TEST , pages 1--42, 2024

  11. [19]

    RaJIVE : Robust angle based JIVE for integrating noisy multi-source data

    Erica Ponzi, Magne Thoresen, and Abhik Ghosh. RaJIVE : Robust angle based JIVE for integrating noisy multi-source data. arXiv preprint arXiv:2101.09110 , 2021

  12. [20]

    Triple component matrix factorization: Untangling global, local, and noisy components

    Naichen Shi, Salar Fattahi, and Raed Al Kontar. Triple component matrix factorization: Untangling global, local, and noisy components. Journal of Machine Learning Research , 25(332):1--76, 2024

  13. [21]

    Personalized PCA : Decoupling shared and unique features

    Naichen Shi and Raed Al Kontar. Personalized PCA : Decoupling shared and unique features. Journal of machine learning research , 25:1--82, 2024

  14. [22]

    Heterogeneous matrix factorization: When features differ by datasets

    Naichen Shi, Raed Al Kontar, and Salar Fattahi. Heterogeneous matrix factorization: When features differ by datasets. arXiv preprint arXiv:2305.17744 , 2023

  15. [23]

    A spectral method for multi-view subspace learning using the product of projections

    Renat Sergazinov, Armeen Taeb, and Irina Gaynanova. A spectral method for multi-view subspace learning using the product of projections. arXiv preprint arXiv:2410.19125 , 2024

  16. [24]

    High-dimensional probability: An introduction with applications in data science , volume 47

    Roman Vershynin. High-dimensional probability: An introduction with applications in data science , volume 47. Cambridge university press, 2018

  17. [25]

    Perturbation theory for pseudo-inverses

    Per- ke Wedin. Perturbation theory for pseudo-inverses. BIT Numerical Mathematics , 13:217--232, 1973

  18. [26]

    Normal approximation and confidence region of singular subspaces

    Dong Xia. Normal approximation and confidence region of singular subspaces. Electronic Journal of Statistics , 15(2):3798--3851, 2021

  19. [27]

    Assouad, F ano, and L e C am

    Bin Yu. Assouad, F ano, and L e C am. In Festschrift for L ucien L e C am , pages 423--435. Springer, New York, 1997

  20. [28]

    Group component analysis for multiblock data: Common and individual feature extraction

    Guoxu Zhou, Andrzej Cichocki, Yu Zhang, and Danilo P Mandic. Group component analysis for multiblock data: Common and individual feature extraction. IEEE transactions on neural networks and learning systems , 27(11):2426--2439, 2015

  21. [29]

    Limit results for distributed estimation of invariant subspaces in multiple networks inference and PCA

    Runbing Zheng and Minh Tang. Limit results for distributed estimation of invariant subspaces in multiple networks inference and PCA . arXiv preprint arXiv:2206.04306 , 2022

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.