{"id":"2b3820b5-2d35-4109-ba60-c738bb664c48","arxiv_id":"2506.02664","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"For a spiked matrix-tensor model with a shared latent vector, sequential matrix-then-tensor recovery reaches the optimal Bayesian weak-recovery thresholds, while joint risk minimization makes even the easy matrix part harder to recover.","lead":"This paper studies a mathematical model in which one shared hidden signal is observed through two noisy channels, a matrix and a tensor, and shows that recovering the matrix first makes the harder tensor part tractable while learning both at once hurts performance. It derives exact phase-transition thresholds and gives a simple spectral method that reaches the best Bayesian thresholds.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The ERM-degradation claim rests on equating ML-AMP fixed points with gradient-descent behavior; Prop. 4.1 proves only fixed-point equality, not dynamics or basins.","rationale":"The reader's weakest-assumption analysis identifies exactly the load-bearing gap: the paper derives the joint-degradation phase diagram for ML-AMP, not for gradient descent, and Prop. 4.1 only matches fixed points. Since the conclusion that joint empirical risk minimization fails is a headline contribution and is stated in the abstract without qualification, this gap is the most consequential weakness in the central argument. I checked whether a more severe concern exists, such as the rigor of state evolution for non-symmetric tensor AMP, but the paper's own framing and the broad AMP literature make the ML-AMP-to-GD bridge the clearest and most direct point where the argument is weakest. The numerical evidence in Figs. 8-9 supports SE and ML-AMP, so the fix is to add a direct analysis of spherical gradient descent or to explicitly restrict the claim to ML-AMP. The conditional verdict already reflects this requirement, so I recommend no change to the reader's verdict.","tokens_in":25131,"tokens_out":19602,"duration_ms":218551,"concrete_test":"Run spherical gradient descent exactly as in Eq. 4.1 on the ML loss L_{rho=1} for the finite-size model with N1=1000, alpha2=1.5, alpha3=0.8, alpha4=1, using both random and spectral initializations, sweeping (Delta_m, Delta_t) across the ML-AMP thresholds sqrt(alpha_tilde_2) and delta_tilde_c shown in Fig. 9; measure ensemble-averaged MSE for u, v, x, y. If the observed gradient-descent recovery boundary differs from the ML-AMP/SE boundary beyond finite-size effects, Proposition 4.1's fixed-point equivalence is insufficient for the ERM conclusion and the abstract's 'empirical risk minimization fails' statement must be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4's headline conclusion that empirical risk minimization for joint learning fails is not established. Proposition 4.1 proves only that fixed points of ML-AMP coincide with stationary points of spherical gradient descent on the loss L_rho of Eq. 4.1. The ML-AMP iteration in App. A.2 is a nonlocal, Onsager-corrected message-passing recursion with denoiser eta_ML(b) = sqrt(N) b / ||b||; it is not a gradient-descent trajectory, and equality of fixed points does not imply equality of attraction basins or finite-time behavior. The thresholds stated in Theorem 4.1 and the SE of Eq. 2.4 characterize ML-AMP's asymptotic overlaps from specified initializations, not those of gradient descent. In a rough non-convex landscape, gradient descent can miss stationary points that ML-AMP reaches, or be trapped in regions from which ML-AMP escapes. Therefore the abstract's claim that joint ERM 'significantly degrades performance' is unsupported by the provided analysis. The sequential spectral result of Theorem 5.1 and the Bayes-AMP thresholds in Theorem 3.1 do not depend on this particular bridge, so they remain intact.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a non-symmetric spiked matrix-tensor model in which the matrix spike and the tensor spike share the same latent vector u*. It analyzes three algorithmic settings: Bayes-AMP, a maximum-likelihood AMP variant used as a proxy for empirical risk minimization (ERM) via gradient descent, and a sequential spectral method that first estimates the matrix and then the tensor. The main claims are closed-form weak-recovery thresholds: for Bayes-AMP, matrix recovery at Δ_m = sqrt(α_2) and joint matrix-tensor recovery at Δ_t = sqrt(α_3 α_4)(α_2 − Δ_m^2)/(α_2 + Δ_m); for joint ML-AMP, the matrix threshold is strictly worse, Δ_m = sqrt(α~_2) with α~_2 = α_2 (α_2/Δ_m)/(α_2/Δ_m + α_3 α_4/(ρ^2 Δ_t)); and for sequential spectral estimation, the Bayes-optimal thresholds are recovered. The theoretical state-evolution predictions are compared with finite-size simulations for the sequential spectral method, Bayes-AMP, and ML-AMP.","tokens_in":25344,"tokens_out":5394,"duration_ms":55162,"significance":"If the results hold, this is a valuable contribution to the theory of multi-modal high-dimensional inference. It provides an exactly solvable model in which structural correlation between a hard tensor channel and an easy matrix channel creates a staircase effect, making tensor recovery possible at O(1) noise levels and yielding a simple curriculum-learning strategy with optimal thresholds. The main positive features are the explicit, parameter-free threshold formulas; the clear separation of the Bayesian, ML-AMP, and sequential regimes; and the finite-size simulations in Figures 3, 8, and 9 that match the theoretical SE curves. The conceptual message that learning order matters is interesting and well illustrated. However, the ERM-degradation claim is currently established only for ML-AMP, not for actual gradient descent, and the Bayes-AMP fixed-point analysis relies on an imported free-energy ansatz; these points need to be addressed before the paper's headline claims can be accepted.","major_comments":[{"comment":"The proof of Theorem 3.1(iii) uses a free-energy potential Φ~(q) in Appendix E that the authors state is \"adapted\" from the non-symmetric tensor model [13] and the symmetric matrix-tensor model [24], rather than derived for the non-symmetric model considered here. Because this free energy is used to assert that an all-nonzero stable fixed point exists and is selected in R_t, the sharp phase diagram of Theorem 3.1 is conditional on an unproven ansatz. Please either prove the free-energy expression (or its needed properties) for this model, or replace the free-energy selection argument with a direct stability and basin analysis, and state explicitly in the main text which parts of Theorem 3.1 are rigorous.","section":"Theorem 3.1 and Appendix E"},{"comment":"The weak-recovery definition appears dimensionally inconsistent. With estimators normalized so that ||w_k|| = Θ(√N_k), the quantity ||(u ⊗ v)^T (w_1 ⊗ w_2)|| equals |<u,w_1>| |<v,w_2>|, which is Θ(N_1 N_2) for estimators positively correlated with the spikes, not Θ(1) as written. As stated, neither the weak-recovery nor the no-recovery conditions can be satisfied in the asymptotic regime considered. The definition should rescale the overlap (for example by dividing by N_1 N_2 or by using normalized unit-norm vectors) so that the stated Θ(1) conditions match the mean-field overlaps q_k used throughout the paper.","section":"Definition 2.3"}],"minor_comments":[{"comment":"There is a typo in the third contribution bullet: \"in constrast\" should be \"in contrast.\"","section":"Introduction"},{"comment":"The paragraph beginning \"This result naturally arises by replacing u_t with u_mat...\" is repeated almost verbatim two paragraphs later; one copy should be deleted.","section":"Appendix B"},{"comment":"The statement that the sequential method \"recovers the optimal Bayesian thresholds\" is correct, but the wording could be sharpened: the sequential spectral method achieves the same threshold values as Bayes-AMP, but its overlaps (q_s in Eqs. 5.1–5.2) are not identical to the Bayes-AMP overlaps, so the claim concerns thresholds, not per-coordinate estimation performance.","section":"Section 5"},{"comment":"The ML-SE equations in Appendix A.2 are said to follow from [28], but the non-separable denoiser η_ML requires the augmented divergence definition introduced in Definition 2.2; it would be helpful to state exactly how the standard SE result is extended to this non-separable case, since this is not a direct application of the usual theorem.","section":"Appendix A.2"}],"recommendation":"major_revision","confidential_remarks":"The main substantive risk is the ML-AMP-to-gradient-descent bridge in Section 4; the Bayes-AMP and sequential spectral results are independent of that bridge and appear to be in good shape. I would not recommend rejection, but the abstract and conclusion overstate the ERM finding unless the dynamical equivalence is supplied or the claims are scaled back. I also recommend asking the authors to make the status of the free-energy ansatz in Appendix E explicit during revision, since it is load-bearing for the Bayes-AMP phase diagram."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper gives a clean analysis of a new non-symmetric spiked matrix-tensor model where one vector (u*) is shared between the matrix and tensor channels. The Bayes-AMP state evolution and the sequential spectral thresholds are worked out in closed form, with finite-size simulations backing the curves. The staircase effect—that the tensor becomes recoverable only after the matrix is recovered—is a nice and non-obvious finding, and the explicit formulas for the recovery regions appear correct. The paper is honest about importing the free-energy ansatz from prior symmetric matrix-tensor work and about the standard nature of the AMP machinery.\n\nThe soft spot is exactly where the stress-test lands. Section 4 claims that empirical risk minimization (ERM) for joint learning fails and that the matrix threshold degrades. What is actually analyzed is ML-AMP, and Proposition 4.1 proves only that fixed points of ML-AMP coincide with stationary points of spherical gradient descent. That does not establish that gradient descent trajectories follow the ML-AMP state evolution, nor that basins of attraction match. For rough non-convex losses like the tensor likelihood, fixed-point equality is a long way from dynamic equivalence. The paper's own text acknowledges that directly analyzing GD is hard and uses ML-AMP as a proxy, but the abstract and conclusion present the ERM failure as a proven fact. That is an overstatement. The Bayes-AMP thresholds and the sequential spectral theorem do not depend on this bridge, so the core of the paper stands. The ERM section needs to be reframed: either as a claim about ML-AMP specifically, or with a much more careful argument (or separate simulation) connecting ML-AMP dynamics to actual GD in this model.\n\nMinor points: the paper has a few duplicated paragraphs in the appendix, and the proof of Theorem 3.1 is partly a sketch (\"there exists at least one stable fixed point\" via a compactness/free-energy argument). That is probably fine for a paper of this type, but a referee will want the existence argument expanded.\n\nWho gets value from this: anyone working on multi-modal inference, spiked tensor models, or AMP theory. The sequential curriculum result is a concrete, testable prescription. I would not cite it immediately unless I worked directly on this model, but I would read it if it came out.\n\nRecommendation: send to a serious referee. The Bayesian and spectral results deserve publication, and the ERM claim needs to be either fixed or softened before acceptance.","headline":"Solid new model and clean thresholds for Bayes-AMP and sequential spectral methods, but the ERM-degradation claim overstates what ML-AMP fixed-point analysis can prove.","tokens_in":25864,"tokens_out":2747,"would_cite":false,"duration_ms":25706,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","62H25","82B44"],"pacs":[],"model":"deepseek-v4-flash","headline":"A shared latent signal lets an efficient algorithm recover matrix and tensor at thresholds that joint empirical risk minimization provably worsens, while a matrix-first curriculum restores the optimal thresholds.","keywords":["spiked matrix-tensor model","Bayes-AMP","state evolution","weak recovery threshold","phase transition","curriculum learning","multi-modal learning","tensor PCA"],"falsifier":"Fix $\\rho=1$ and choose parameters with $\\sqrt{\\tilde{\\alpha}_2} < \\Delta_m < \\sqrt{\\alpha_2}$, where joint ML predicts no recovery but the Bayes and matrix-first analyses predict recovery. Run spherical gradient descent on the joint ML loss from a matrix-SVD-informed initialization: if it systematically recovers the matrix, the ML-AMP proxy under-predicts empirical risk minimization, and if it fails from an uninformative initialization where the ML state evolution says the matrix fixed point is stable, the proxy over-predicts.","tokens_in":24936,"feed_emoji":"🧩","tokens_out":7586,"duration_ms":70407,"temperature":0.7,"pith_summary":"The paper introduces a non-symmetric spiked matrix-tensor model in which a noisy matrix and a noisy third-order tensor are correlated because they share one latent vector $u^\\star$: the matrix carries $Y_m = \\sqrt{\\Delta_m}Z + u^\\star(v^\\star)^\\top/\\sqrt{N_1}$ and the tensor carries $Y_t = \\sqrt{\\Delta_t}Z + u^\\star \\otimes x^\\star \\otimes y^\\star/N_1$. It aims to establish that this correlation changes what is computationally easy: a Bayes-AMP algorithm, whose updates are posterior means, recovers the matrix and then the tensor at $O(1)$ noise thresholds, whereas an isolated tensor spike is efficiently recoverable only at noise of order $O(N^{-1/2})$. For the more practical empirical-risk-minimization setting, studied through an ML-AMP proxy for gradient descent, the paper claims that joint fitting is self-defeating: the tensor term effectively shrinks the matrix aspect ratio from $\\alpha_2$ to $\\tilde{\\alpha}_2 \\leq \\alpha_2$, so the matrix threshold moves from $\\sqrt{\\alpha_2}$ to the strictly worse $\\sqrt{\\tilde{\\alpha}_2}$. Finally, it claims that a sequential curriculum---matrix PCA first, then contracted tensor PCA---restores the optimal Bayesian thresholds and can be implemented by two spectral decompositions. If true, the results give a precise, solvable illustration of why the order in which modalities are learned can matter as much as the information they contain.","feed_headline":"Learn the easy modality first: joint training hurts recovery","feed_subtitle":"Joint ML optimization worsens recovery thresholds; sequential PCA attains the Bayesian limits.","key_machinery":"The central object is state evolution for two variants of Approximate Message Passing. For the Bayes denoiser $\\eta(a,b)=b/(1+a)$, the empirical overlaps follow the low-dimensional recurrence $q^{t+1} = f^{\\mathrm{Bayes}}(g_{\\rho=1}(q^t))$ with $g_1 = \\alpha_2 q_2/\\Delta_m + \\alpha_3\\alpha_4 q_3q_4/\\Delta_t$; for the maximum-likelihood denoiser, which projects to the sphere, the analogous ML state evolution carries the effective ratio $\\tilde{\\alpha}_2$. Stability analysis of the zero fixed point, the matrix-only fixed point, and the matrix-tensor fixed point of these recurrences produces all thresholds. The sequential result is carried by the BBP phase transition: after PCA on the matrix, the contracted tensor is again a rank-one spiked matrix, so the second spectral step has closed-form overlaps and inherits the original matrix threshold.","core_discovery":"The paper studies weak recovery in the non-symmetric spiked matrix-tensor model with a shared spike, and its central claim is a complete phase diagram for three algorithms. Bayes-AMP recovers the matrix exactly when $\\Delta_m < \\sqrt{\\alpha_2}$ and recovers the tensor as well as soon as $\\Delta_t < \\delta_c^{\\mathrm{Bayes}} = \\sqrt{\\alpha_3 \\alpha_4}\\,\\frac{\\alpha_2 - \\Delta_m^2}{\\alpha_2 + \\Delta_m}$, both $O(1)$ noise thresholds, whereas an isolated tensor of this form is efficiently recoverable only at noise of order $O(N^{-1/2})$. The correlation thus induces a staircase phenomenon: learning $u^\\star$ from the matrix jump-starts learning $x^\\star$ and $y^\\star$ from the tensor, and the tensor cannot be recovered independently of the matrix. Joint ML-AMP, the paper's proxy for empirical risk minimization by gradient descent, replaces $\\alpha_2$ by the strictly smaller effective aspect ratio $\\tilde{\\alpha}_2 = \\alpha_2 \\frac{\\alpha_2/\\Delta_m}{\\alpha_2/\\Delta_m + \\alpha_3\\alpha_4/(\\rho^2 \\Delta_t)}$, so the matrix threshold moves from $\\sqrt{\\alpha_2}$ to $\\sqrt{\\tilde{\\alpha}_2}$ and the tensor-recovery region shrinks: adding the tensor modality makes the otherwise easy matrix task harder, and no finite weight $\\rho$ repairs the loss. Estimating the matrix first by PCA and then applying PCA to the contracted tensor $Y_t \\cdot \\hat{u}/\\sqrt{N_4}$ restores the Bayesian thresholds and produces closed-form overlaps for all four signal components.","pith_inferences":["Editor's extension: if ML-AMP faithfully mirrors gradient descent, the model predicts that in practical multi-modal training a shared representation trained jointly with a high-loss tensor branch can be degraded relative to training it with the easier modality alone, making curricula or staged fine-tuning the safer design.","Editor's extension: the staircase mechanism is stated for one shared vector and one tensor channel; the same state-evolution analysis would likely extend to multiple shared factors or higher-order tensor channels, producing a sequence of thresholds in which each learned factor lowers the noise needed for the next.","Editor's extension: because the pure tensor problem has a large statistical-to-computational gap, the $O(1)$ tensor threshold here predicts a strong finite-size benefit; a direct finite-$N$ comparison of Bayes-AMP and joint gradient descent across that gap would sharpen the practical reading.","Editor's extension: a falsifiable design rule follows: warm-starting tensor learning with a matrix-PCA estimate should beat joint training even when both use identical information, and performance should degrade monotonically as the tensor weight in the loss grows."],"forward_implications":["A pure tensor spike that is computationally hard at moderate noise becomes efficiently recoverable whenever it is statistically correlated with a matrix spike through one shared factor.","Joint empirical risk minimization over heterogeneous losses can be strictly worse than Bayesian optimal inference: the tensor term degrades the matrix threshold, and no finite loss weight fully repairs it.","Learning order is a computational resource: a matrix-first sequential curriculum attains the optimal Bayesian thresholds while simultaneous optimization does not.","The sequential scheme is implementable with off-the-shelf PCA, so the threshold gain is available to a simple spectral pipeline.","The phase diagrams contain hard phases and spinodal transitions where an informative initialization, such as a matrix SVD warm start, recovers signals that an uninformative initialization cannot reach."],"supporting_citations":[{"why":"Supplies the BBP phase-transition lemma used to derive the spectral thresholds for both steps of the sequential PCA algorithm.","marker":"[20]"},{"why":"Provides the spiked tensor model whose statistical-to-computational gap defines the baseline that the correlated matrix-tensor model improves upon.","marker":"[8]"},{"why":"Establishes the constrained low-rank matrix estimation phase transitions and the AMP derivation pattern on which Bayes-AMP and ML-AMP are built.","marker":"[21]"},{"why":"Introduces the staircase property that the paper identifies as the mechanism behind matrix-assisted tensor recovery.","marker":"[15]"},{"why":"Supplies the state evolution theorem used to justify the overlap recurrences in Proposition 2.1.","marker":"[28]"},{"why":"Treats the non-symmetric spiked tensor model whose layered structure and mutual information the free-energy analysis adapts.","marker":"[13]"}],"fun_headline_variants":["Easy-first learning beats joint training in multi-modal","Matrix-first curriculum hits optimal recovery thresholds","Sequential PCA beats joint ML for tensor-matrix signals","Learn matrix first, then tensor: joint effort fails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that empirical risk minimization fails depends on the unproven assumption that ML-AMP trajectories, basins of attraction, and state evolution faithfully represent spherical gradient descent, since only the fixed points of the two dynamics are shown to coincide.","fun_headline_variants_meta":{"raw":{"variants":["Easy-first learning beats joint training in multi-modal","Matrix-first curriculum hits optimal recovery thresholds","Sequential PCA beats joint ML for tensor-matrix signals","Learn matrix first, then tensor: joint effort fails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1495,"prompt_tokens":1067,"completion_tokens":428,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":368}},"tokens_in":683,"tokens_out":428,"duration_ms":4657,"temperature":1.0,"reasoning_tokens":368,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:20:30.001553+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fix $\\rho=1$ and choose parameters with $\\sqrt{\\tilde{\\alpha}_2} < \\Delta_m < \\sqrt{\\alpha_2}$, where joint ML predicts no recovery but the Bayes and matrix-first analyses predict recovery. Run spherical gradient descent on the joint ML loss from a matrix-SVD-informed initialization: if it systematically recovers the matrix, the ML-AMP proxy under-predicts empirical risk minimization, and if it fails from an uninformative initialization where the ML state evolution says the matrix fixed point is stable, the proxy over-predicts.","supporting_citations":[{"cited_title":"Phase transition of the largest eigenvalue for non-null complex sample covariance matrices","cited_arxiv_id":null,"evidence_quote":"Supplies the BBP phase-transition lemma used to derive the spectral thresholds for both steps of the sequential PCA algorithm."},{"cited_title":"Statistical and compu- tational phase transitions in spiked tensor estimation","cited_arxiv_id":null,"evidence_quote":"Provides the spiked tensor model whose statistical-to-computational gap defines the baseline that the correlated matrix-tensor model improves upon."},{"cited_title":"Constrained low-rank matrix estimation: phase transitions, approximate message passing and applications","cited_arxiv_id":null,"evidence_quote":"Establishes the constrained low-rank matrix estimation phase transitions and the AMP derivation pattern on which Bayes-AMP and ML-AMP are built."},{"cited_title":"The staircase property: How hierarchical structure can guide deep learning","cited_arxiv_id":null,"evidence_quote":"Introduces the staircase property that the paper identifies as the mechanism behind matrix-assisted tensor recovery."},{"cited_title":"State evolution for general approximate message passing algorithms, with applications to spatial coupling","cited_arxiv_id":null,"evidence_quote":"Supplies the state evolution theorem used to justify the overlap recurrences in Proposition 2.1."},{"cited_title":"The layered structure of tensor estimation and its mutual information","cited_arxiv_id":null,"evidence_quote":"Treats the non-symmetric spiked tensor model whose layered structure and mutual information the free-energy analysis adapts."}],"review_version":1}