{"id":"03385b11-e022-477b-833c-a8e8ec069d35","arxiv_id":"2512.02866","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Geometry of the individual components, not the number of views, determines whether AJIVE's error vanishes at the K^{-1/2} rate; a weighted variant handles heterogeneous views.","lead":"This paper shows that a common method for finding shared structure across several data tables (AJIVE) performs far better than recent theory suggested, as long as the views' hidden components are not all aligned in the same direction. The authors add a weighting scheme, HeteroJIVE, that emphasizes informative views and prove when it is optimal.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2 omits an independence-across-views requirement: sign symmetry alone cannot yield the K^{-1/2} cancellation; the paper's own shared-loading simulation contradicts the theorem as stated.","rationale":"The most load-bearing issue is not the θ condition (which the paper explicitly acknowledges as an identifiability gap), but the missing independence requirement in the cancellation argument. The proof's matrix-Hoeffding step is necessary: without independence across k, the 'averaging out' of the sign-symmetric bias term does not occur. The paper's own simulation of the shared scheme provides empirical evidence that the theorem as stated is too broad, so this is not a merely cosmetic gap. I recommend CONDITIONAL rather than REJECT because the paper's random-scheme simulations and the independent-loading version of Theorem 2 are credible, and adding 'independent across k' to Assumption 1 (plus adjusting the abstract and theorem statements) would fix the issue. The reader's selected weakest assumption (θ) is important but secondary; it is a limitation of the method rather than an inconsistency in the stated theorem.","tokens_in":32488,"tokens_out":18236,"duration_ms":157959,"concrete_test":"Run the equal-weight AJIVE on the shared scheme of Section 5.1 with V_k=V, W_k=W drawn once from O(d,r) with V^T W≠0, and U_k generated with θ=0.5, for K=8,16,32,64 (n=d=20, σ=0.1, γ=0.5, 100 replicates). If the averaged error does not decay as K^{-1/2} (matching the flattening shared curve in Figure 2), then Theorem 2 as stated—which permits shared V_k under Assumption 1—is contradicted. A minimal analytical companion: compute Var(Σ_{k=1}^K (1/K) c [Γ^{-1}]_{12}) under perfect dependence; it is O(1), not O(1/K), invalidating the matrix-Hoeffding step in Supplement B.2.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Assumption 1 states only that each V_k is sign-symmetric conditional on W_k (a marginal, per-view condition). It does not require the V_k to be independent across k. The proof of the O(K^{-1/2}) rate (Supplement B.2, after Eq. (14)) relies on showing E[Γ_k^{-r}]_{12}=0 and then applying matrix Hoeffding to Σ_k w_k Σ_r c_{k,r} U^T \\bar U_k Γ_k^{-r} \\bar U_k^T U_⊥ Λ^{-1}. Matrix Hoeffding requires independence of the summands. If V_k=V for all k (the paper's 'shared scheme' in Section 5.1), the summands are perfectly correlated; the centered bias is a single random matrix of norm ~ ε² θ^{-1} δ(1-δ²)^{-1}, not a quantity that shrinks as K^{-1/2}. Figure 2 (left) empirically confirms this: the shared-scheme curve flattens as K grows even though the loadings are sign-symmetric. Thus Theorem 2's statement—that Assumption 1 alone yields a vanishing K^{-1/2} rate—is false as written; the valid claim requires independent sign-symmetric loadings across views. This is the mechanism that actually delivers the averaging.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HeteroJIVE, a weighted two-stage spectral estimator for the JIVE multi-view model, and provides non-asymptotic error bounds for fixed weights. For equal weights it revisits the Yang–Ma (2025) non-diminishing barrier, claiming that under sign-symmetric random loadings the second-order bias is centered and vanishes at an O(K^{-1/2}) rate without iterative refinement. For general weights it gives bounds separating statistical and structural heterogeneity, derives an oracle reweighting scheme minimizing an upper bound, and implements a data-driven plug-in version. The claims are supported by a detailed high-order spectral-projector expansion in the supplement, simulations with several loading geometries, and a TCGA-BRCA application.","tokens_in":32785,"tokens_out":7431,"duration_ms":82299,"significance":"If correct, the paper would substantially refine the theory of AJIVE: it would show that the known O(1) low-SNR barrier is geometry-dependent, that simple spectral aggregation is minimax-optimal in a broader regime than previously known, and that reweighting can address both SNR heterogeneity and individual-component interference. The supplement contains a genuinely detailed expansion of the two-stage spectral error, and the explicit dependence of the second-order term on the loading geometry parameter δ_k is a useful contribution. The oracle-weight derivation also correctly reproduces the known optimal weights of Baharav et al. in the individual-free benchmark. However, one of the two central theoretical claims—the O(K^{-1/2}) rate under Assumption 1—is stated with an assumption that is too weak; the proof silently requires cross-view independence, and the paper's own shared-loading simulation contradicts the theorem as stated. The abstract also promises a rank-one majority sign-alignment result that does not appear in the body.","major_comments":[{"comment":"Theorem 2 is false as stated because Assumption 1 is a per-view marginal condition and does not require independence of (V_k, W_k) across k. The proof of Theorem 4 (Supplement B.2) shows E[Γ_k^{-r}]_{12}=0 and then applies matrix Hoeffding to the sum over k. Matrix Hoeffding requires independence of the summands. If V_k=V and W_k=W for all k, as in the paper's 'shared scheme' of §5.1, Assumption 1 holds but the centered bias does not average out; it is a single random matrix of norm ~ ε^2 θ^{-1} δ(1-δ^2)^{-1}. Figure 2 (left) empirically confirms this: the shared-scheme curve flattens as K grows. The theorem needs an explicit independence or conditional-independence condition across views, and the discussion should distinguish 'sign-symmetric loadings' from 'independent sign-symmetric loadings'.","section":"§3.2, Theorem 2; Supplement B.2"},{"comment":"The arXiv abstract states: 'Under a majority sign-alignment condition in rank-one setting, a bias at the squared single-view perturbation scale can persist.' I could not find this result anywhere in the main text or the supplement. No theorem, proposition, or section addresses majority sign-alignment or the rank-one persistent-bias claim. This is either a missing contribution or an unsupported abstract claim; it must be added or removed. The abstract of the full-text PDF also differs from the arXiv metadata, so the two should be reconciled.","section":"Abstract vs. main text"},{"comment":"The 'oracle-optimal' weighting is not actually shown to be optimal for the estimation error. The criterion J(w) in §4.1 is an upper bound, not the exact risk, and the oracle weight minimizes J(w), not the true subspace error. Proposition 2 only proves that a positive fixed point is an approximate stationary point of J, not of the risk. The only setting where genuine optimality is established is the individual-free case (U_k=0), where the weights reduce to w_k ∝ ε_k^{-2} and match Baharav et al. (2025). The abstract's phrase 'explicit weight that is optimal whenever its identifiability gap is constant' is therefore unsupported and should be replaced with a statement about minimizing the proved upper bound.","section":"§4.1 and Abstract"},{"comment":"The data-driven weights bw are computed from the same data used to evaluate the final estimator, but all theorems (Theorems 1–4) treat w as fixed and independent of the data. No result controls the additional error from plug-in estimation of ε_k, θ, and M_k, or the randomness of bw. The simulation section demonstrates empirical success, but the practical recommendation would be much stronger if the paper either provided a stability/sample-splitting argument or explicitly stated that the data-driven procedure has no current theoretical guarantee. As written, the theory and the implemented method are separated by an unquantified plug-in gap.","section":"§4.2 and §5"}],"minor_comments":[{"comment":"The parity argument in the proof of Theorem 4 is misphrased: after defining f_ij, the text says 'Each summand is even number of even functions in V_k'. The valid statement is that every path from index 1 to index 2 uses an odd number of odd f's (f_12 or f_21), so the product is odd; the conclusion E[Γ_k^{-r}]_{12}=0 still follows. Please correct the wording.","section":"Supplement B.2"},{"comment":"The sentence 'As soon as either randomness across views ... or orthogonality ... is introduced, the error decays roughly at the K^{-1/2} rate' is imprecise: 'randomness across views' must mean independent randomness, not mere sign-symmetry, since the shared scheme is also random but does not decay. Please align the prose with the corrected assumption.","section":"§5.1"},{"comment":"The comparison with Tang et al. (2025) says the bound 'merely requiring K≳log n' whereas Tang et al. require K≳ε^{-4} log n. This is a meaningful point, but the sentence 'our bound vanishes as either K→∞ or SNR→∞' should be stated as 'vanishes as K→∞ under the maintained SNR condition', since the constants depend on the SNR condition (12).","section":"§1.1 and Table 1"},{"comment":"The real-data section reports SWISS scores and misclassification rates without any measure of variability. Given that the reported differences are small (0.5547 vs 0.5415; 0.3103 vs 0.2759), a bootstrap or repeated-split assessment would help. This is a presentation issue, not a blocker.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The core issue is fixable but substantive: Theorem 2's assumption is too weak, and the paper's own Figure 2 (left) demonstrates the failure mode. I recommend requiring the authors to add an independence assumption across views (or a valid dependence structure) and to re-state Theorem 2 and Theorem 4 accordingly. The abstract's missing majority sign-alignment claim and the 'optimal' weighting overstatement should also be corrected. The high-order expansion and the δ_k refinement are valuable, so I would not recommend rejection; the paper can likely be made sound with focused revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: the paper has a real and non-obvious idea, and the deterministic part holds up; but Theorem 2 as stated is false, because it omits the independence across views that the proof needs.\n\nI read the manuscript with the stress-test note in hand. The reader's instinct was right. Assumption 1 is per-view sign symmetry, but the K^{-1/2} cancellation in the proof (Supplement B.2) comes from matrix Hoeffding on the sum over k of the centered cross terms. Matrix Hoeffding needs the summands independent. The paper's own 'shared scheme' simulation is the counterexample: V_k and W_k are shared across views but sign-symmetric, so Assumption 1 holds, and Figure 2 (left) shows the error flattening rather than decaying. That's not a minor technicality; it's the mechanism that delivers the averaging. The fix is simple—assume independence of the views or at least of the summands—and then the rate is plausible. But as written, Theorem 2 (and the random case of Theorem 4) is wrong, and the title claim 'without iterative refinement' rests on it.\n\nThe deterministic contributions are solid and genuinely new. The identification of δ_k, the angle between the joint and individual loadings, as the control parameter for the second-order bias is a nice refinement of Yang and Ma. The general-weight theorem and the oracle weighting scheme are well-motivated, and the reduction to Baharav et al. in the no-individual case is clean. The use of Xia's spectral projector expansion is appropriate and the appendix is substantive.\n\nOther soft spots: the abstract advertises a rank-one 'majority sign-alignment' bias result that I could not find in the main text or supplement. That should be either included or deleted. The real-data analysis is a single TCGA run with no uncertainty quantification; fine for illustration. No code or data, which makes the simulations hard to verify. The spectral-gap condition (13) is strong, but the paper is transparent about it and it is probably unavoidable for identifiability.\n\nThe parity argument in B.2 is worded confusingly but the conclusion is correct, so that one is not a real defect.\n\nBottom line: the core geometric insight is likely correct for independent random loadings, and the deterministic analysis is self-contained. A careful referee would request the independence condition, a corrected abstract, and at least simulation code. That is a cheap fix for what could be a solid paper. I would send it to peer review, but I would not cite the current version.","headline":"The geometry-dependent barrier is a real idea, but Theorem 2 as stated is false: it needs an independence-across-views condition the paper never states, and the abstract promises a rank-one result that never appears.","tokens_in":33263,"tokens_out":7381,"would_cite":false,"duration_ms":68483,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","62H12"],"pacs":[],"model":"deepseek-v4-flash","headline":"A weighted two-stage spectral estimator recovers a shared subspace across heterogeneous views at the optimal O(K^{−1/2}) rate, showing the AJIVE error barrier is geometry dependent rather than universal.","keywords":["multi-view data","joint subspace estimation","JIVE","AJIVE","spectral aggregation","heterogeneity","perturbation bounds","reweighting"],"falsifier":"Simulate JIVE data with sign-symmetric random loadings, constant SNR across views, and θ bounded below by a constant; if the joint subspace error does not decay at the K^{−1/2} rate as K grows, the central rate claim fails. Equivalently, compute the empirical mean of the off-diagonal block of the inverse Gram matrix Γ_k^{−1}; a nonzero mean would directly contradict the bias-cancellation mechanism behind the theorem.","tokens_in":32369,"feed_emoji":"📊","tokens_out":4795,"duration_ms":53764,"temperature":0.7,"pith_summary":"This paper addresses a practical problem: when several data matrices measured on the same samples share a low-dimensional signal subspace, how should one aggregate per-view singular subspaces when views differ in noise level, dimension, and strength of view-specific structure? It shows that the equal-weight AJIVE rule is not as limited as recent theory suggested: a previously reported error barrier that does not shrink as the number of views K grows is an artifact of worst-case loading geometry, and it vanishes under orthogonal or sign-symmetric random loading orientations, yielding the optimal K^{−1/2} rate without iterative refinement. For genuinely heterogeneous views, the paper proposes HeteroJIVE, a weighted version whose weights down-weight noisy or interference-heavy views; it derives an oracle-optimal weighting scheme, a data-driven plug-in, and error bounds that separate statistical from structural heterogeneity. A reader should care because this suggests simple two-stage spectral methods are minimax-optimal over a much wider regime than previously understood, and because the weighting scheme is easy to implement.","feed_headline":"AJIVE's error barrier disappears under random loadings","feed_subtitle":"HeteroJIVE reweights noisy views to recover shared subspaces at the O(K^-1/2) rate.","key_machinery":"The estimator is the top-r eigenspace of P = Σ w_k Ũ_k Ũ_k^T, a weighted sum of per-view projection matrices. The analysis expands the spectral projector difference via a full all-orders perturbation formula, using the misalignment gap θ(w) = 1 − ||Σ w_k U_k U_k^T|| and the per-view loading angle δ_k = ||(V_k^T V_k)^{−1/2} V_k^T W_k (W_k^T W_k)^{−1/2}|| as the two geometric controls. Under sign symmetry of V_k conditional on W_k, the off-diagonal blocks of the inverse Gram matrices Γ_k^{−r} are odd functions, so their expectations vanish; this cancellation is what removes the previously identified non-diminishing second-order bias. The weighting map iteratively reweights views inversely to a","core_discovery":"The paper's central claim is that the estimation error of the two-stage spectral AJIVE estimator is governed by the geometry of the view loadings. In the equal-weight case, the previously reported non-diminishing second-order bias in the low-SNR regime is not unavoidable: when loadings V_k and W_k are nearly orthogonal, the bias term is small, and when loadings are sign-symmetric random, the bias terms have conditional mean zero and cancel across views, leaving a bound of order ε sqrt(r̄ log n / K) plus a smaller term. Thus the joint subspace can be recovered at the optimal K^{−1/2} rate without iterative refinement, under the condition that the individual subspaces are not too aligned. For","pith_inferences":["One could build a practical diagnostic from the first-stage SVDs: estimate δ_k and the effective θ to predict whether adding another view will reduce error or merely accumulate shared bias.","The fixed-point reweighting scheme resembles iteratively reweighted least squares; the stationarity result suggests a practical stopping rule based on the norm of the projected gradient.","The sign-symmetry cancellation suggests that deliberately randomizing loading signs across views—or collecting views with anti-correlated orientations—would accelerate bias cancellation; this is a testable data-augmentation extension.","The results imply minimax lower bounds for two-stage spectral AJIVE should be revisited under generic geometric assumptions rather than worst-case configurations."],"forward_implications":["Under sign-symmetric random loadings, equal-weight AJIVE achieves the O(K^{−1/2}) rate without iterative refinement, so alternating-projection refinement is theoretically redundant in that regime.","The non-diminishing error barrier in low-SNR AJIVE is confined to degenerate shared-and-aligned loading geometry; adding more views still helps when loadings are random or orthogonal.","For general weights, the optimal weighting behaves like w_k ∝ ε_k^{−2} when individual components are absent, and the data-driven plug-in approximates this oracle weight in non-asymptotic settings.","On a four-view breast cancer multi-omics dataset, the reweighted estimator improved SWISS score and clustering misclassification relative to equal-weight AJIVE."],"fun_headline_variants":["Random view geometry kills AJIVE's non-diminishing error","HeteroJIVE: optimal view reweighting hits K^-1/2 rate","Error barrier vanishes under sign-symmetric views","Geometry-aware AJIVE: error barrier is not universal"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The individual subspaces must be sufficiently misaligned on average: the weighted average of per-view individual projectors must be bounded away from the identity, i.e., θ(w) must dominate the noise level, otherwise the joint subspace is not identifiable and the error bounds diverge as 1/θ.","fun_headline_variants_meta":{"raw":{"variants":["Random view geometry kills AJIVE's non-diminishing error","HeteroJIVE: optimal view reweighting hits K^-1/2 rate","Error barrier vanishes under sign-symmetric views","Geometry-aware AJIVE: error barrier is not universal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001249,"raw_usage":{"total_tokens":4975,"prompt_tokens":780,"completion_tokens":4195,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":4136}},"tokens_in":524,"tokens_out":4195,"duration_ms":27928,"temperature":1.0,"reasoning_tokens":4136,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T18:54:00.443938+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate JIVE data with sign-symmetric random loadings, constant SNR across views, and θ bounded below by a constant; if the joint subspace error does not decay at the K^{−1/2} rate as K grows, the central rate claim fails. Equivalently, compute the empirical mean of the off-diagonal block of the inverse Gram matrix Γ_k^{−1}; a nonzero mean would directly contradict the bias-cancellation mechanism behind the theorem.","supporting_citations":[],"review_version":1}