{"id":"81407726-5015-42ca-b6e1-a5921cbf4bf3","arxiv_id":"2608.01282","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Prediction transfers across sources with different modality sets by first reweighting each source to the target outcome distribution, then aligning all modalities to a target CCA anchor via ridge maps.","lead":"The paper proposes a way to predict outcomes in a new population when training data come from several sources with different data modalities available and different outcome rates. A shared reference modality is used to estimate the target outcome mix, reweight the sources, and then align all other modalities to a target-defined representation before transferring predictions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's posterior-consistency claim rests on Assumption 3, which is not derived for the implemented density-ratio learner and, in particular, assumes the class-conditional density transport that Lemma 2 explicitly does not establish.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing concern: the alignment theorem is credible, but the consistency theorem depends on Assumption 3, which assumes both likelihood-ratio consistency and identifiability of the profile criterion. My read agrees and sharpens it: the transport of class-conditional densities from source ridge coordinates to target CCA coordinates is not a consequence of Lemma 2, and the paper explicitly isolates it as a separate nuisance condition. The supplement is transparent that for the flexible-tree implementation the required rate clauses remain assumed, and that the surrogate pathway is not covered by the rate theorem. These are addressable gaps rather than internal contradictions: the high-level assumption framework is legitimate, and the alignment result is independently supported by a detailed perturbation proof. However, because the central end-to-end consistency claim depends on an unverified premise, conditional acceptance is the appropriate verdict. The recommendation is therefore unchanged relative to the reader's verdict.","tokens_in":48282,"tokens_out":5612,"duration_ms":58295,"concrete_test":"Use the main simulation DGP at default settings to generate oracle population aligned coordinates from the known latent model, then compute the true class-conditional distributions of (u^(m), v^(m)_k) under the source ridge map G^(m) and of (u^(0), v^(0)_k) under the target CCA map B^(0). Estimate the total-variation distance between these distributions for each source domain and modality block. If this distance does not shrink to zero as source sample size grows, Assumption 3's transport clause fails and Theorem 1's exact-posterior conclusion does not apply; if the distance does vanish, the remaining unproved component is the rate kappa_n for the density-ratio learner, which should be verified separately by evaluating the implemented learner on oracle population coordinates across increasing source sizes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weak point is Assumption 3 (Section 4.2; Assumption 6 in Supplement S1.1), used in Theorem 1 to convert the alignment result into consistency of the target conditional outcome distribution. Lemma 2 establishes that fitted coordinates are recovered up to one common rotation, but the paper's own remark after Eq. (18) states that common orientation does not imply source-to-target likelihood-ratio transport and that transport is 'the separate nuisance condition in Assumption 3.' This is not a minor technicality: the source ridge maps G^(m) solve a penalized regression of A^(0)^T r on z, while the target CCA map B^(0) solves a constrained correlation problem; these maps coincide only under additional structural conditions. If they differ, the aligned source coordinates and target coordinates have different class-conditional laws, so a pooled source density-ratio estimator can be biased for the target ratio even with infinite source data. Assumption 3 simply assumes the average ratio error is o_p(1) at rate kappa_n, including this transport bias. The Supplement's rate ledger (item 4) is explicit that for the boosted-tree implementation 'both clauses remain assumed,' and the surrogate pathway is stated not to be covered by the rate theorem. The final paragraph of Section 4.3 narrows the claim to a 'working conditional distribution' when the composite ratio is not exact, but the abstract and Section 1.3 present consistency of the target conditional outcome distribution as the main theoretical result. Thus the central theorem is conditional on an unverified premise that is not derived from Assumptions 1-2.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a four-stage pipeline for multimodal domain adaptation under label shift and blockwise missingness: (1) unsupervised representation learning; (2) reference-based estimation of the target outcome distribution and reweighting of source observations; (3) target-anchored CCA/ridge alignment after reweighting; and (4) likelihood-ratio-based outcome transfer to estimate the target conditional outcome distribution. The main theoretical contributions are Lemma 2, establishing that the target CCA anchor, all source ridge maps, and all fitted aligned coordinates are recovered up to one common rotation at rate Delta_n, and Theorem 1, giving consistency and rates for the target conditional outcome distribution under Assumptions 1–3. The paper also develops a surrogate-label-assisted variant for sparse gold-standard source labels and reports simulations plus a renal cell carcinoma application.","tokens_in":48597,"tokens_out":5371,"duration_ms":53113,"significance":"The 'weight before alignment' idea is well motivated, and the one-common-rotation alignment result is a useful technical step for combining target CCA with source ridge regression under blockwise missingness. The paper is unusually transparent about the limits of its theory, explicitly stating in the supplement that the likelihood-ratio learner's consistency is assumed rather than derived for the boosted-tree implementation. If the outcome-transfer step were justified for a concrete ratio learner, the framework would be a solid contribution. As it stands, the central advertised claim—consistency of the target conditional outcome distribution—is conditional on an assumption that is not verified for the implemented method, which substantially weakens the contribution as stated.","major_comments":[{"comment":"Theorem 1's consistency claim for the target posterior is assumed, not derived, for the implementation actually used. Assumption 3 (Assumption 6 in the supplement) postulates both that the aligned likelihood-ratio estimator is consistent at rate kappa_n and that the oracle profile criterion identifies theta0. The supplement's own rate ledger states that for boosted trees 'both clauses remain assumed,' and the implementation indeed uses XGBoost-based density-ratio estimation. Therefore Theorem 1 does not establish consistency of the target conditional outcome distribution for the proposed pipeline; it reduces that claim to an unverified high-level condition. The authors should either supply a rate theorem for a concrete ratio-learning procedure under explicit conditions (e.g., a correctly specified parametric ratio model) or restate the contribution as a reduction theorem and confine the exact-posterior consistency claim to the 'working conditional distribution' mentioned only in the final paragraph of Section 4.3.","section":"§4.2–§4.3, Assumption 3; Supplement S1.1, rate ledger item 4"},{"comment":"The population ridge map G^(m) is defined as the minimizer of a weighted least-squares fit of A^(0)^T r onto z, whereas the target CCA map B^(0) solves a constrained correlation problem. These maps need not coincide, so the class-conditional law of the source aligned coordinates v^(m) need not equal that of the target aligned coordinates v^(0) even when the raw scores are transportable under Assumption 1. Lemma 2 only recovers each fitted coordinate up to a common rotation relative to its own population map, and the paper's own remark after Eq. (18) explicitly disclaims that common orientation implies likelihood-ratio transport. The 'intrinsic transfer rate' kappa_n in Assumption 3 therefore silently absorbs a source-to-target transport bias that the alignment result does not control. The paper should state explicit structural conditions under which the ridge-mapped source coordinates have the same outcome-conditional distribution as the target CCA coordinates, and ideally verify such conditions in the simulation design; otherwise the outcome-transfer step remains an unverified premise.","section":"§4.1, definition of G^(m); §4.2, Assumption 3; remark after Eq. (18)"}],"minor_comments":[{"comment":"The abstract and contribution section state consistency of the target conditional outcome distribution as a headline result without the qualification that appears only at the end of Section 4.3, where the authors note that the composite ratio (8) may define a working conditional distribution rather than the exact target posterior. The claims should be aligned with the theorem's actual scope.","section":"Abstract and §1.3"},{"comment":"The composite ratio is said to correspond to the joint conditional likelihood ratio when the aligned auxiliary blocks are conditionally independent given (U,Y). The paper does not state whether this conditional independence is part of the main theoretical assumptions or only an interpretation; clarifying this would help the reader assess the gap between the composite construction and the exact ratio in Theorem 1.","section":"§3.4, Eq. (8)"},{"comment":"The ablation terminology is inconsistent: the main text's 'target-anchor CCA/ridge match-up' is sometimes called the 'CCA switch' and sometimes 'target-anchor CCA,' while Table S8 lists four variants with slightly different names. Aligning the terminology across the main text, supplement, and tables would improve readability.","section":"Supplement S2.5 and Table S8"}],"recommendation":"major_revision","confidential_remarks":"The paper's core alignment lemma is solid and the transparency about the ratio-learning limitation is commendable, but the advertised posterior-consistency result is conditional on an assumption that the authors explicitly state remains unverified for the implemented boosted-tree density-ratio learner. This is a load-bearing gap that needs to be addressed either by providing a verifiable rate for a concrete learner or by appropriately narrowing the claims to a working conditional distribution. The revision should also make the distinction between the alignment contribution and the transport assumption visible in the abstract and introduction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWorth knowing: the alignment core is genuine and new; the headline consistency theorem is not fully proven as advertised. Lemma 2, deriving one common rotation for target CCA and all source ridge maps from standard perturbation theory, is the paper's real contribution and holds up. The \"weight before alignment\" idea—estimate the target outcome mixture from the reference modality, reweight source observations, then align—is a genuinely useful construction for EHR-style blockwise missingness. The literature survey is careful and positions the work correctly.\n\nThe soft spot is exactly where the stress test points. Theorem 1's consistency and rate for the target posterior are conditioned on Assumption 3, which assumes both consistency of the aligned likelihood-ratio learner and unique identification of theta0. More importantly, Assumption 3 embeds the density transport step: that the class-conditional laws of the source ridge coordinates coincide with those of the target CCA coordinates after alignment. Lemma 2 explicitly does not establish this; the paper says so after Eq (18), and the supplement's rate ledger admits that for the boosted-tree implementation \"both clauses remain assumed.\" So the abstract's consistency claim is conditional on an unverified premise. This is not a secret flaw—it is stated honestly—but it should be moved to the front: either present Theorem 1 as a conditional result with Assumption 3 flagged as a separate research problem, or prove transport under structural conditions on loadings and score distributions.\n\nOther concerns are minor. The EM tilt-back safeguard is active in about 95% of simulation replications, so the simulations mostly characterize the stabilized update rather than unrestricted EM; the paper discloses this clearly. Bootstrap standard errors without refitting are common in this literature but do overstate precision. The surrogate pathway is not covered by the rate theorem, and the real-data endpoint is mostly LATTE-derived; the supplement says plainly this is a mixed-endpoint transport analysis, not adjudicated validation.\n\nBottom line: the method is practical, the alignment theory is worthwhile, and the posterior-consistency theorem needs re-framing or strengthening. I would send this to a serious referee; conditional acceptance with targeted revisions is the right outcome. I'd also bring it to a reading group as an example of unusually honest assumption accounting.","headline":"Novel reference-anchored alignment with an honest but load-bearing assumption gap: the one-rotation result is solid, while the posterior-consistency headline rides on an unproven transport condition.","tokens_in":49185,"tokens_out":1986,"would_cite":true,"duration_ms":20124,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H20","62G05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A reference-anchored, reweight-before-align procedure is claimed to recover a common target-defined multimodal representation and consistent target conditional outcome estimates under label shift and blockwise missing modalities.","keywords":["domain adaptation","label shift","blockwise missing modalities","canonical correlation analysis","ridge regression","surrogate outcomes","multimodal data fusion","target calibration"],"falsifier":"Generate two synthetic replications that differ only in whether the class-conditional distribution of the reference score given the outcome is exactly transportable across source and target; the pipeline should show the claimed error rate in the transportable case and a systematic calibration gap in the non-transportable case, so observing equal errors would falsify the transport assumption.","tokens_in":48034,"feed_emoji":"🔄","tokens_out":9309,"duration_ms":76120,"temperature":0.7,"pith_summary":"The paper aims to show that multimodal domain adaptation remains possible when the unlabeled target population differs from labeled sources in its outcome distribution and each source observes only a subset of modalities. The proposed route is to first use a reference modality present in every domain to estimate the target outcome mixture, reweight source observations to that mixture, and only then align auxiliary modalities to a target-defined canonical correlation analysis anchor using ridge maps. Its central result is that this reweight-before-align construction recovers all fitted aligned coordinates up to one common rotation at rate $\\Delta_n$, and that under a separate likelihood-ratio condition the estimated target conditional outcome distribution is consistent at rate $n_0^{-1/2}+\\kappa_n+\\Delta_n$. If correct, the method yields calibrated target probabilities without requiring sources to share a common input space, and it degrades gracefully when gold-standard labels are sparse by borrowing surrogate outcomes.","feed_headline":"Reweight, then align: multimodal adaptation survives label shift","feed_subtitle":"Why it matters: calibrated predictions from partially observed multimodal sources despite prevalence shift.","key_machinery":"The central object is the target-calibrated, reference-anchored match-up. A reference modality observed in every domain is used to estimate the target outcome distribution via a profile likelihood; each source subject is reweighted by the estimated label-shift ratio $\\widehat{\\omega}_m(y)$, and the target defines a canonical correlation analysis (CCA, a method that finds linear combinations of two variable sets with maximal correlation) anchor between the reference scores and the concatenated auxiliary scores. Each source auxiliary block is then mapped into this target-anchored coordinate system by weighted ridge regression, so that all aligned coordinates share one common target-defined orientation up to a single orthogonal rotation. The proof's key move is to apply a subspace-perturbation theorem to a symmetric dilation of the CCA operator, which yields one common rotation $O$ for the target CCA pair, all source ridge maps, and every fitted coordinate block; ridge equivariance then propagates this same $O$ through the source normal equations.","core_discovery":"The central claim, established as Lemma 2 and Theorem 1, is that a target-defined representation can be learned from heterogeneous source blocks despite label shift and blockwise missingness. Under the label-shift and overlap assumption, the target outcome distribution is identified from the reference modality, and source observations are reweighted so that their marginal cross-modal associations match target associations; the target CCA anchor and every source ridge map are then shown to be recovered up to one common orthogonal rotation $O$ at rate $\\Delta_n$, with the same $O$ appearing in all domains and modalities. Under the additional Assumption 3 that the aligned likelihood-ratio estimator is consistent at population coordinates, the target outcome parameter and the target conditional outcome distribution are consistent, at rate $n_0^{-1/2}+\\kappa_n+\\Delta_n$. The paper is careful to distinguish representation recovery from outcome transport: common orientation alone does not carry source class-conditional densities to the target, so the ratio condition is a separate load-bearing requirement. The same pipeline with a surrogate bridge is claimed to approach the full-label oracle when source gold labels are sparse, supported by simulations and by an application to 12-month recurrence prediction in renal cell carcinoma.","pith_inferences":["The one-common-rotation result is a representation-level guarantee: it implies the aligned coordinates are a well-defined target-defined feature map, so the same match-up could serve downstream tasks other than the specific likelihood-ratio profile used here, such as clustering or classification with a different outcome model.","The reweight-before-align principle may transfer to other anchor-based alignment estimators; as long as the primitive error bounds in Assumption 2 hold, replacing CCA by another anchor estimator should preserve the common-rotation property.","A direct testable extension would compare this method with an align-first, reweight-second variant in the same simulation design; the paper's ablation isolates the CCA/ridge component but does not separately measure the ordering effect, which the simulations do not directly report."],"forward_implications":["When this pipeline is deployed, predicted risks in the target remain calibrated even if source prevalence differs sharply from target prevalence, because the outcome mixture is corrected before alignment rather than as a post-hoc recalibration step.","Auxiliary modalities can enter some sources and not others without being pooled into one shared input space; each observed block is projected into the target-anchored coordinate system by its own ridge map.","With sparse gold-standard labels, the surrogate-assisted version recovers most of the ranking performance of a full-label procedure and, with a small fraction of gold labels, matches its probability calibration at larger sample sizes (within 0.3% of oracle AUC at the largest simulated size).","Because the fitted coordinates are identified up to one common rotation, and because that rotation does not alter the conditional outcome distribution, any downstream learner applied to the aligned coordinates inherits the same consistency guarantee."],"supporting_citations":[{"why":"defines label shift and supplies the detection/correction framework that grounds the reweighting identity.","marker":"[Lipton et al., 2018]"},{"why":"provides the unified label-shift estimation rate and the profile-criterion machinery used to estimate the target outcome distribution.","marker":"[Garg et al., 2020]"},{"why":"supplies the Davis–Kahan variant used with a symmetric dilation to obtain one common rotation for the target CCA anchor.","marker":"[Yu et al., 2015]"},{"why":"provides the singular-subspace perturbation bound underlying the CCA recovery step.","marker":"[Wedin, 1972]"},{"why":"introduces canonical correlation analysis, the target anchor defining the shared aligned coordinate system.","marker":"[Hotelling, 1936]"},{"why":"justifies ridge regularization as the stabilizer that makes source alignment maps unique under collinearity and rank deficiency.","marker":"[Hastie, 2020]"},{"why":"supplies the argmax and M-estimation theorems used to turn ratio consistency into posterior consistency.","marker":"[Van der Vaart, 2000]"}],"fun_headline_variants":["Reweight then align: adaptation survives label shift","Target-defined CCA anchors reweighted multimodal sources","Unified rotation recovery for label-shift adaptation","Consistent estimation despite blockwise missing modalities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that, once the aligned coordinates are fixed, the outcome probabilities given those coordinates can be estimated consistently from reweighted source data at a known rate and the target outcome mixture is uniquely identifiable from that estimate; for the boosted-tree implementation in the paper this remains an assumption rather than a proven rate.","fun_headline_variants_meta":{"raw":{"variants":["Reweight then align: adaptation survives label shift","Target-defined CCA anchors reweighted multimodal sources","Unified rotation recovery for label-shift adaptation","Consistent estimation despite blockwise missing modalities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1602,"prompt_tokens":1005,"completion_tokens":597,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":538}},"tokens_in":621,"tokens_out":597,"duration_ms":5736,"temperature":1.0,"reasoning_tokens":538,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:09:24.219230+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate two synthetic replications that differ only in whether the class-conditional distribution of the reference score given the outcome is exactly transportable across source and target; the pipeline should show the claimed error rate in the transportable case and a systematic calibration gap in the non-transportable case, so observing equal errors would falsify the transport assumption.","supporting_citations":[],"review_version":2}