{"id":"cb8fa685-a78d-402f-9840-9bb21f1c51df","arxiv_id":"2411.17054","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Stack-SVD is minimax optimal for fully shared singular subspaces; with partial sharing, rate-optimal estimation requires locating shared vectors, and the proposed tracing algorithm does so under orthogonality and strong-signal conditions.","lead":"This paper determines when the common practice of stacking noisy data tables and running SVD gives the best possible estimate of shared hidden structure. It also identifies when that practice fails and provides an algorithm that still works under partial sharing.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper claims the proposed algorithms attain minimax optimality under partial sharing, but no theorem bounds the risk of the estimator that actually selects the index set J; the practical algorithm is only analyzed under an orthogonal-unshared assumption that the paper itself leaves for future…","rationale":"The reader's weakest-assumption analysis identified the orthogonality of unshared subspaces as load-bearing for the algorithm, and my reading agrees. The paper is honest about this limitation in Section 7, but the abstract and contribution list claim minimax optimality for methods under partial sharing without that caveat. What would be needed is either an end-to-end theorem bounding the risk of the estimator using bJ, or an explicit qualification that the practical algorithm's optimality is proven only in the orthogonal, well-separated regime. The simulations in Section 5.3 support the algorithm's behavior in the orthogonal setting, and the oracle theorems are plausible, but the missing bridge between oracle and practical estimator is precisely where the central claim is least secure. I did not find evidence of internal inconsistency in the stated theorems themselves, so the appropriate verdict remains CONDITIONAL: the paper is a strong contribution whose headline claim needs tightening or additional analysis.","tokens_in":26693,"tokens_out":14956,"duration_ms":149155,"concrete_test":"Simulate the Section 5.3 Setting 2 with unshared vectors U1*, U2* rotated to have inner product rho = 0.5 (and rho = 0.8), keeping all singular values and gaps at the Table 3 row where the orthogonal algorithm has success rate 1. Over 1000 trials, record the fraction of runs with bJ = J and the average sin-Theta distance between Ur and the subspace spanned by the selected vectors. If exact recovery fails for rho > 0, or if bJ strictly contains J in a nontrivial fraction of runs, then the practical algorithm does not deliver the oracle rates of Theorem 3.5 outside the orthogonal submodel, confirming that the minimax-optimality claim is conditional on an assumption the paper has not analyzed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the proposed methods attain minimax rate-optimality under partial sharing rests on an unproved bridge between the oracle results and the practical tracing algorithm. Theorem 3.1 gives minimax rates for the oracle estimator Uhat_J^r assuming the true index set J is known, and Theorem 3.5 extends this to non-orthogonal unshared subspaces. But the actual estimator uses the index set bJ produced by Algorithms 1 and 2. The consistency of bJ is established only in Theorems 4.1 and 4.2, which assume the unshared singular vectors are mutually orthogonal and that all singular values of (X1 X2) are separated, a condition stronger than the switch-gap conditions in Hr,t and Sr,t. Section 7 explicitly states that the non-orthogonal case is future work and that in the presence of pronounced non-orthogonality the algorithm may return an over-inclusive set J subset of bJ. Thus no theorem proves minimax optimality for the estimator that uses estimated J, either in the non-orthogonal Sr,t model or even over the full orthogonal parameter space used by the oracle theorems. The abstract's assertion that the methods are proven minimax rate-optimal under partial sharing is therefore stronger than the results in the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the problem of estimating a shared left singular subspace Ur from multiple noisy matrices Yi = Xi + Zi in a low-rank matrix denoising framework. The authors analyze Stack-SVD, which takes the top left singular vectors of the stacked matrix (Y1 Y2), and establish minimax upper and lower bounds for sin-Theta risk when the signal matrices share an identical singular subspace (Theorem 2.1), with extensions to k matrices (Corollary 2.3). For partially shared subspaces, they introduce parameter spaces with orthogonal unshared subspaces (H_{r,t}) and with non-orthogonal unshared subspaces (S_{r,t}), and prove matching upper and lower bounds for an oracle estimator that knows the true index set J of shared vectors in the stacked SVD (Theorems 3.1-3.5). They then propose Algorithms 1 and 2 to identify J, prove their consistency under an additional orthogonality and separation condition (Theorems 4.1 and 4.2), and support the theory with simulations and a single-cell data application.","tokens_in":26881,"tokens_out":8501,"duration_ms":74602,"significance":"Should the deferred proofs be correct, the paper makes a substantial theoretical contribution: it provides minimax rates for shared singular subspace estimation across multiple noisy matrices, identifies dimension- and SNR-dependent phase transitions, and demonstrates that Stack-SVD is rate-optimal under full sharing while a popular Average-SVD alternative can be suboptimal. The matching rate forms in Theorems 2.1 and 3.1, together with simulation results that track the predicted phase boundaries, are encouraging internal evidence. The proposed tracing algorithm is simple and shows promising empirical performance on simulated and single-cell data. However, the central minimax-optimality claim under partial sharing is established only for an oracle estimator that knows J, not for the algorithm that estimates J, and the consistency guarantees for the algorithm are proven under assumptions stronger than those defining the minimax parameter spaces. These gaps are acknowledged in Section 7 but not resolved, and they materially limit the force of the abstract's claims.","major_comments":[{"comment":"The minimax optimality results for the partial-sharing models, Theorems 3.1 and 3.5, are stated for the oracle estimator \\hat U_r^J, which uses the true index set J of shared singular vectors in the stacked SVD. The practical procedure of Section 4 replaces J by the output \\hat J of Algorithm 1, but no theorem in the paper bounds the risk of \\hat U_r^{\\hat J}. Theorem 4.1 only gives P(\\hat J = J) → 1 under conditions that are stronger than the defining conditions of H_{r,t} and S_{r,t}, and it does not combine this consistency event with the risk bounds (9)-(12). Consequently, the abstract's claim that the proposed methods are proven minimax rate-optimal under partial sharing is not supported for the estimator a user would actually run.","section":"Section 4 / Theorem 3.1 / Theorem 4.1"},{"comment":"The assumptions of Theorem 4.1 require every singular value of the stacked signal matrix to be separated (σ_k ≥ (1+δ)σ_{k+1}) and require the minimum matrix signal strengths to satisfy (α^{c1} ∧ β^{c1}) ≥ Cn and (α∧β)^2 ≥ C2(p1∨p2). These conditions are global, while the minimax parameter spaces H_{r,t} and S_{r,t} impose only local gap conditions at type-switch positions and a lower bound on the minimum gap over switches. The theorem therefore proves consistency of the tracing algorithm only on a sub-regime of the parameter spaces used for the minimax results, leaving open whether the algorithm's success probability is high enough to preserve the rate (9) when the signal is near the boundary of H_{r,t} or S_{r,t}.","section":"Section 4, Theorem 4.1"},{"comment":"Theorem 3.5 extends the oracle minimax bounds to non-orthogonal unshared subspaces (parameter space S_{r,t}), but the tracing algorithm's consistency (Theorems 4.1 and 4.2) is proved only under U_{1*}^T U_{2*} = 0. The paper's own Section 7 states that for pronounced non-orthogonality the algorithm may return an over-inclusive set with J ⊂ \\hat J and that extending the algorithm is future work. Thus the paper does not establish minimax optimality for the practical estimator outside the orthogonal-unshared setting, and the abstract's broad phrase 'under partial sharing' overstates the proven scope.","section":"Section 3.3 / Section 7"},{"comment":"The lower bounds (7)-(8) are stated under the additional conditions p1 ≍ p2 or γ ≳ τ^2(p1+p2), whereas the upper bounds (5)-(6) hold on the full class F_{r,γ}. The manuscript does not provide a lower bound for the regime with p1 and p2 of different orders and moderate γ, so the phrase 'Stack-SVD achieves minimax rate-optimality when the true singular subspaces are identical' in the abstract is not true on the full parameter space defined in Section 2.1. Section 7 acknowledges the gap, but the abstract and the summary bullet points should be qualified accordingly.","section":"Section 2.1, Theorem 2.1"},{"comment":"All proofs are deferred to a supplementary file that is not included with the submission; the main text refers to 'Section S1.4' after Theorem 2.1, to 'Section S1.1' after Theorem 3.1, and Theorem 4.1 is stated without a proof sketch. Because the contributions are primarily theoretical, the absence of the supplement leaves the central lower-bound and consistency arguments unverifiable in the submitted version. The authors should provide the supplement for review.","section":"Supplementary Material / Throughout"}],"minor_comments":[{"comment":"In step 8, the phrase 'Take the last r1−k1 and r2−k2 index sets' is inconsistent with the subsequent description that these are the indices of the smallest values of d1i and d2j; please clarify whether the selection is based on the smallest or the largest distances.","section":"Section 4, Algorithm 1"},{"comment":"The symbols α, β, γ are reused with different meanings: here α = σ_min(X1), β = σ_min(X2), and γ = α ∧ β, whereas γ is the signal-strength parameter in F_{r,γ} and in the discussion following Theorem 2.1; this notational collision should be fixed.","section":"Section 4, Theorem 4.1"},{"comment":"The term 'column singular matrix' is not standard and is not defined in the text; also, the displayed SVD expressions would be much easier to check if the dimensions of S, U*, Σ*, and V* were stated explicitly.","section":"Section 3.3, Theorem 3.4"},{"comment":"The text announces a novel lower-bound argument but gives no outline of it in the main text; given that the lower bound is a central technical novelty, the authors should include an informal description of the construction in Section 2 or in the Introduction.","section":"Section 2.1, after Theorem 2.1"},{"comment":"There are several typos: 'staked' and 'the staked matrix' in Example 3 and Section 4; 'A interesting direction' in Section 7; 'the the singular value' in the paragraph after Theorem 3.1; and in the references, [25] contains a stray '6' in 'The Annals of statistics 47 6 3009-3031'.","section":"Throughout"},{"comment":"The axes and constants in the phase diagrams are not defined in the caption or the main text; please specify what is plotted (e.g., log SNR versus dimension ratios) so that the claimed regions can be checked.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for the Annals of Statistics in topic and technical depth. My main reservations concern the gap between the oracle results and the practical tracing algorithm, and the stronger conditions under which the algorithm is proven consistent; both are acknowledged by the authors in Section 7 but are not reflected in the abstract. I recommend requesting a revision that either proves a risk bound for the estimator using the estimated index set \\hat J (at least under the H_{r,t} conditions), or explicitly limits the abstract and contribution claims to the oracle setting and to the orthogonal-unshared case. Please also ask the authors to supply the supplementary material before final review, since all proofs are deferred to it and the current submission cannot be fully verified without it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a real contribution, not a repackaging. It gives minimax rates for estimating a shared left singular subspace from stacked noisy matrices, shows Stack-SVD achieves those rates under complete sharing, and maps out partial-sharing regimes with phase transitions. As far as I can tell, no previous work had minimax lower bounds for this problem; the Fan–Zheng–Tang and JIVE lines are asymptotic or algorithmic. The lower-bound construction looks genuinely new, and the upper bounds are tied to external perturbation bounds rather than to the paper's own results. On those grounds the paper deserves referee time.\n\nWhere it is solid: Theorem 2.1 has matching upper and lower bounds on the same parameter space, the rates have no fitted constants, and the simulations line up with the predicted phase diagram. Proposition 1 cleanly explains why unshared vectors can be pulled into the stacked SVD. The citation pattern is honest—Cai and Zhang (2018) is used as a black box in the natural way, and the JIVE-related literature is covered\nfairly.\n\nThe soft spot is not in the oracle results but in the bridge to the algorithm. The minimax theorems for the shared estimator assume the index set J is known. The actual algorithm estimates J, and its consistency is proven only under mutually orthogonal unshared subspaces plus complete singular-value separation in the stacked matrix. The non-orthogonal case is explicitly left to future work, and the paper itself says the algorithm may over-include vectors. So no theorem currently proves minimax optimality for the fully practical procedure under the general partial-sharing model. That is a genuine gap, and the abstract overstates the claim when it says the methods are proven minimax rate-optimal under partial sharing. The gap is addressable—prove a risk bound for the estimator using the estimated J under the orthogonality condition, or state the theorem with the extra assumptions up front—but it should be fixed before publication.\n\nA second limitation: all proofs are in a Supplement that is not included, so I could not verify the lower-bound machinery. That is common for this genre, but it matters here because the paper's central claim rests on proofs I have not seen.\n\nWho is this for? Anyone working on multi-view PCA, JIVE-type methods, or subspace estimation. My recommendation: send it to a knowledgeable referee with the Supplement, and require the authors to either add a theorem for the full procedure or soften the claims in the abstract.","headline":"A serious minimax-theory contribution on Stack-SVD with a real but bridgeable gap between the oracle results and the claimed guarantees for the practical algorithm.","tokens_in":27452,"tokens_out":2260,"would_cite":true,"duration_ms":22093,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H12","62C20","62H25"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that when several noisy low-rank matrices share a left singular subspace, the Stack-SVD estimator—the SVD of the concatenated matrices—is minimax rate-optimal, and that when the shared subspace is only partial…","keywords":["shared singular subspace","low-rank matrix denoising","Stack-SVD","minimax optimality","singular vector tracing","multi-view data integration","spectral methods"],"falsifier":"Simulate the fully shared two-matrix model over a grid of $n$, $p_1=p_2$, and per-direction signal $\\gamma^2$ chosen so the predicted minimax rate $n(\\gamma^2+p_1+p_2)/\\gamma^4$ lies strictly between 0 and 1, and compare the empirical worst-case $\\|\\sin\\Theta(U,\\hat U)\\|^2$ of Stack-SVD with that rate; if the observed errors decay at a strictly faster rate as $n$ grows, the matching lower bound cannot be correct.","tokens_in":26437,"feed_emoji":"📊","tokens_out":9797,"duration_ms":87236,"temperature":0.7,"pith_summary":"The paper asks how accurately one can recover the shared left singular subspace of several low-rank signal matrices from their noisy observations. Its central answer is that Stack-SVD, which computes the top singular vectors of the concatenated noisy matrix, is minimax rate-optimal when the matrices share the same singular subspace, meaning no estimator can have smaller worst-case error up to constants, with matching upper and lower bounds of order $n(\\gamma^2 + p_1 + p_2)/\\gamma^4$ in squared spectral distance. The paper then characterizes when Stack-SVD stays optimal under partial sharing: it remains rate-optimal when unshared signals are weak, while a shifted version that selects the shared directions of the stacked SVD is needed when unshared signals dominate. It also proposes an algorithm that traces which singular vectors of the stacked matrix are shared, and proves that under orthogonality of the unshared vectors it recovers the correct set with probability tending to one. The result is a benchmark for multi-view data integration: simple stacking is provably as good as any method in the fully shared case, and a modest tracing step restores optimality when parts of the subspace are matrix-specific.","feed_headline":"Stacking noisy matrices is the statistically optimal move","feed_subtitle":"A proof that the standard stacking trick already reaches the best possible error rate.","key_machinery":"The argument is carried by two structural facts and a separation index. Proposition 1 says that when all singular vectors of the two signal matrices are pairwise orthogonal, every singular vector of $X_1$ and $X_2$ appears in the SVD of the stacked matrix $(X_1\\, X_2)$, possibly reordered, so estimating the shared subspace reduces to locating the correct columns of the stacked SVD. Proposition 2 is a one-sided perturbation bound for a selected block of $r$ singular vectors, giving $\\mathbb{E}\\|\\sin\\Theta(\\hat U_r, U_r)\\|^2 \\le c p_1(\\sigma_r^2(X)+p_2)/\\sigma_r^4(X)\\wedge 1$; it replaces the uniform two-sided perturbation bound that would make the rate depend on the wrong side of the matrix. Around these sit the parameter spaces $H_{r,t}$ and $S_{r,t}$, defined by requiring an eigen-gap $g^2 > c\\sigma_{s+1}^2$ at each vector-type switch and a minimum switch gap $t^2$; the index set $J$ of shared singular vectors in the stacked matrix is what the oracle estimator and the tracing algorithm are built to recover.","core_discovery":"The core discovery is that the minimax risk for estimating a fully shared left singular subspace from noisy matrices is, up to constants, $n(\\gamma^2 + p_1 + p_2)/\\gamma^4$ in squared spectral distance, with an extra factor $r$ in Frobenius distance, where $\\gamma$ is the minimum over shared directions of the sum of squared singular values across matrices; Stack-SVD attains this rate and no estimator can improve on it when the matrix dimensions are comparable or the combined signal is large. In the partial-sharing model where each matrix has shared vectors $U_r$ plus unshared vectors $U_{1*}$ and $U_{2*}$ with $U_{1*}^\\top U_{2*}=0$, the optimal rate is governed by the minimum eigen-gap $t$ at the points where the stacked singular vectors switch type: an oracle that selects the shared positions achieves $n(t^2 + p_1 + p_2)/t^4$, with matching lower bounds. The naive top-$r$ Stack-SVD selector becomes inconsistent when unshared signals are strong, but the paper shows that selecting the singular vectors at the correct positions, for example the $(d+1)$-th through $(d+r)$-th when $d$ unshared vectors lead in the stacked spectrum, restores the minimax rate. Non-orthogonality of the unshared subspaces is shown not to change this picture for the oracle estimator, because the stacking SVD rotates the unshared block while leaving the shared subspace and its singular values intact.","pith_inferences":["Editorial inference: a natural extension is to adapt the tracing algorithm to quantify mild non-orthogonality of unshared vectors, using the gap between within-matrix and cross-matrix $\\sin\\Theta$ distances; the paper's over-inclusive-set observation suggests the size of that gap carries information about the angle between unshared subspaces.","Editorial inference: the phase-transition threshold $\\min\\{\\sigma^2_{(i)}(X_1)+\\sigma^2_{(i)}(X_2)\\}/\\tau^2 \\asymp \\sqrt{n(n+p_1+p_2)}$ gives a practical diagnostic for whether stacking will help, and suggests that adaptive procedures trading off individual SVD and stacked SVD could interpolate smoothly across the critical region.","Editorial inference: the one-sided perturbation bound should extend to the right singular subspace by swapping the roles of $n$ and $p_i$, and to higher-order analogues such as stacked tensors, yielding similar minimax benchmarks for multi-view problems outside the matrix case."],"forward_implications":["When all $k$ matrices share the same left singular subspace, Stack-SVD is minimax rate-optimal, so any averaging or principal-angle alternative can do no better under comparable dimensions or strong signals.","Stacking can identify shared directions that are individually non-identifiable in every single matrix, because the squared singular value of a shared direction in the stacked matrix is the sum of its squared signals across matrices.","When unshared signals dominate the shared ones, taking the top $r$ singular vectors of the stacked matrix is inconsistent, but selecting the singular vectors at the shared positions restores the minimax rate.","The tracing algorithm separating shared from unshared singular vectors is consistent when the unshared vectors are mutually orthogonal and the singular values are well separated, making the oracle estimator practically implementable.","For the oracle estimator, non-orthogonal unshared subspaces do not change the minimax rate, since the stacked SVD preserves the shared subspace through the rotation of the unshared block."],"supporting_citations":[{"why":"supplies the one-sided rate-optimal perturbation bound for singular subspaces that Proposition 2 generalizes to partial subspace selection.","marker":"[11]"},{"why":"provides the single-matrix subspace estimation lower bounds and parameter-space setup on which the stacked lower-bound argument builds.","marker":"[9]"},{"why":"defines the angle-based joint-and-individual model for partially shared subspaces that the paper's partial-sharing analysis benchmarks against.","marker":"[28]"},{"why":"introduces the joint-and-individual variation framework whose structural assumption of shared plus orthogonal individual components the paper adopts and sharpens.","marker":"[39]"},{"why":"represents the averaging and principal-angle alternative that the paper analyzes and shows can be sub-optimal compared with Stack-SVD.","marker":"[25]"},{"why":"gives the uniform perturbation bound for singular subspaces that the paper contrasts with its sharper one-sided bound.","marker":"[61]"}],"fun_headline_variants":["Stack-SVD provably optimal for shared singular subspaces","Stacked SVD hits the minimax rate for shared subspaces","Don't average, stack: optimal shared subspace recovery","When stacking noisy matrices is provably minimax optimal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's fast algorithm for telling shared from unshared singular vectors is proven to work only when the unshared vectors point in mutually orthogonal directions; if they do not, the bookkeeping can include too many vectors, and the full procedure's optimality is not established.","fun_headline_variants_meta":{"raw":{"variants":["Stack-SVD provably optimal for shared singular subspaces","Stacked SVD hits the minimax rate for shared subspaces","Don't average, stack: optimal shared subspace recovery","When stacking noisy matrices is provably minimax optimal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00101,"raw_usage":{"total_tokens":4342,"prompt_tokens":1092,"completion_tokens":3250,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":3183}},"tokens_in":708,"tokens_out":3250,"duration_ms":23904,"temperature":1.0,"reasoning_tokens":3183,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:36:43.992581+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the fully shared two-matrix model over a grid of $n$, $p_1=p_2$, and per-direction signal $\\gamma^2$ chosen so the predicted minimax rate $n(\\gamma^2+p_1+p_2)/\\gamma^4$ lies strictly between 0 and 1, and compare the empirical worst-case $\\|\\sin\\Theta(U,\\hat U)\\|^2$ of Stack-SVD with that rate; if the observed errors decay at a strictly faster rate as $n$ grows, the matching lower bound cannot be correct.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the one-sided rate-optimal perturbation bound for singular subspaces that Proposition 2 generalizes to partial subspace selection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the single-matrix subspace estimation lower bounds and parameter-space setup on which the stacked lower-bound argument builds."},{"cited_title":"and M ARRON , J","cited_arxiv_id":null,"evidence_quote":"defines the angle-based joint-and-individual model for partially shared subspaces that the paper's partial-sharing analysis benchmarks against."},{"cited_title":"and K RISHNASWAMY , S","cited_arxiv_id":null,"evidence_quote":"introduces the joint-and-individual variation framework whose structural assumption of shared plus orthogonal individual components the paper adopts and sharpens."},{"cited_title":"and Z HU, Z","cited_arxiv_id":null,"evidence_quote":"represents the averaging and principal-angle alternative that the paper analyzes and shows can be sub-optimal compared with Stack-SVD."},{"cited_title":"and C HUA, T.-S","cited_arxiv_id":null,"evidence_quote":"gives the uniform perturbation bound for singular subspaces that the paper contrasts with its sharper one-sided bound."}],"review_version":1}