Pith. sign in

REVIEW 2 major objections 4 minor 73 references

Invariant latent factors in multi-environment data are identifiable from unlabeled covariates alone, and auxiliary labels then make the full latent signal transportable at near-oracle error.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:40 UTC pith:JN7LBIQD

load-bearing objection Solid multi-environment factor-model paper with a genuinely new auxiliary-label alignment step; the main catch is the central invariant-subspace assumption is unfalsifiable and the real-data analysis is thin. the 2 major comments →

arxiv 2607.18209 v1 pith:JN7LBIQD submitted 2026-07-20 math.ST cs.LGstat.MEstat.MLstat.TH

Unveiling Invariant and Transferable Latent Factors Across Heterogeneous Environments via ATLAS

classification math.ST cs.LGstat.MEstat.MLstat.TH MSC 62H2562J1262H12
keywords factor modelinvarianceheterogeneous environmentsmulti-environment datatransfer learninglatent factor regressionauxiliary labelsdiversified projection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish a clean separation: when data are collected from several environments whose covariate distributions differ, the latent structure splits into invariant factors (shared loadings across environments) and heterogeneous factors (environment-specific loadings), and this split is recoverable from unlabeled covariates alone under one geometric condition — the intersection of the per-environment loading spaces must be exactly the shared loading space. Given that split, ATLAS uses auxiliary labels — cheap, noisy proxies observed in a subset of environments — to find which of the heterogeneous factors predict the response in the same way everywhere, and to align them across environments. The result is a prediction rule that transfers: with auxiliary labels in the new environment, the method transports the full invariant signal; without them, it falls back to the invariant factors only, which the paper proves is the unique worst-case-optimal choice. The paper backs this with non-asymptotic guarantees: recovered factors carry dimension-free sub-Gaussian errors, and the transported predictor's error decomposes into one term from estimating the response coefficient, one from aligning factors with auxiliary labels, and one from recovering the latent factors themselves. If correct, this gives a principled recipe for building site-transportable predictors from abundant unlabeled data plus a few labels.

Core claim

On its own terms, the paper's central claim is that invariant and heterogeneous latent factors can be disentangled without any supervision, provided that the intersection of the column spaces of the per-environment loading matrices equals the shared loading space (Assumption 1(a)) and that invariant and heterogeneous factors are block-uncorrelated (Assumption 1(b)). Under those conditions, the invariant factors are identified up to a single common invertible transformation across all environments, and the heterogeneous factors up to environment-specific transformations; with a rank condition on the auxiliary-label coefficients, the response-relevant subset of each block is identified up to a

What carries the argument

The load-bearing object is the maximum invariant subspace: the intersection of the per-environment loading spaces, ∩_e col([B, A^(e)]) = col(B). ATLAS computes it empirically by averaging the top principal-subspace projectors of each environment's covariance matrix and taking the leading eigenspace of the average; this turns a set-theoretic intersection into a spectral step with a measurable eigen-gap (the parameter ϵ_A in Condition 4.2). The second piece is the diversified projection construction: each environment gets its own projection matrix that maps X to factor proxies while partialling out the heterogeneous factors from the invariant projection, so invariant and heterogeneous scores a

Load-bearing premise

The entire identification rests on Assumption 1(a): the intersection of loading spaces across environments is exactly the shared invariant space, meaning no environment-specific factor happens to load in the same direction at every observed site; if one does (a shared coding artifact, a common assay drift), it will be classified as invariant, the 'invariant' predictor will carry that spurious signal, and the failure is undetectable from the data.

What would settle it

Simulate two or more environments with a true invariant loading B and environment-specific loadings A^(e), then add one extra loading vector shared by all environments (violating Assumption 1(a)) and run ATLAS: the recovered invariant space will include the extra direction, and the transferred predictor's worst-case out-of-sample risk will visibly exceed the oracle built on the true B. A cheaper check: compute the average of the per-environment top-subspace projectors and inspect the eigenvalue just after the r_I-th; if it is not separated from 1 (ϵ_A near zero in Condition 4.2), the identific

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A fixed number of environments — as few as two, when heterogeneity is exhaustive — suffices to identify the invariant factors; the number of environments does not need to grow with the latent dimension.
  • Without auxiliary labels, restricting prediction to the invariant factors is the unique worst-case-optimal strategy under rotation uncertainty; heterogeneous factors cannot be transferred without additional supervision.
  • Auxiliary labels improve efficiency even when no heterogeneous factor is response-relevant, by reducing the variance of the estimated invariant signal.
  • The theory covers weak factors (loading strength need not scale with √d) and gives direction-wise, dimension-free sub-Gaussian error control, so downstream error does not accumulate over the factor dimension.
  • The framework extends to nonlinear mean functions: replacing the final regression step with nonparametric estimation substitutes a nonparametric rate for the parametric estimation error, leaving the factor-alignment machinery unchanged.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the method's guarantees inherit the assumption that no spurious 'shared' direction exists across sites; in practice, a coding or measurement convention common to all observed environments will masquerade as an invariant factor. Institutions applying ATLAS should define environments to straddle known convention breaks (coding systems, note templates, lab vendors) so artifacts a
  • Editorial inference: the eigen-gap of the averaged projector suggests a practical diagnostic — plot the spectrum of the averaged projection matrix; if the (r_I+1)-th eigenvalue sits close to 1, the environments are not providing the exhaustive heterogeneity the identification needs, and transfer claims should be downgraded.
  • Editorial inference: the paper's sketch for adapting to a new environment implies a streaming variant — a new site's loading subspace could be intersected with the stored invariant space using only its own covariance and a handful of auxiliary labels, letting a federation of institutions update the transferable model without re-running the entire pipeline.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper studies a multi-environment linear factor model in which each environment's covariates are driven by invariant factors with shared loadings and environment-specific heterogeneous factors. The authors show that, under a maximum-invariant-subspace condition plus block-uncorrelatedness, the invariant and heterogeneous factors are identifiable up to invertible transformations, and they prove instance-level necessity of the identification condition. They then propose ATLAS, a three-stage estimator: (i) an invariance-heterogeneity decomposition (IHD) that estimates shared and environment-specific loading subspaces and constructs diversified projections for the factors; (ii) a spectral step using auxiliary labels Z to select and align prediction-invariant heterogeneous factors; and (iii) a pooled GLM fit in labeled environments to estimate the invariant prediction rule and transfer it to new environments, with or without auxiliary labels. The theoretical core supplies non-asymptotic, dimension-free-in-direction sub-Gaussian error bounds for factor recovery, for aligning prediction-invariant factors, and for the transferred prediction error. Simulations and a temporal EHR application on rheumatoid arthritis are used to illustrate the method.

Significance. If the results hold as stated, the paper makes a useful contribution to multi-environment factor analysis and transfer learning. The identification theory is carefully developed: the maximum-invariant-subspace assumption is not merely imposed but shown to be instance-level necessary, and the use of auxiliary labels to go beyond invariant-factor-only prediction is a genuine extension of invariant causal prediction ideas to latent factor models. The non-asymptotic bounds are detailed and plausible, and the dimension-free sub-Gaussian control is a valuable technical contribution. The real-data demonstration, while limited, is appropriate for the motivating EHR setting. The main value is in providing a statistically rigorous method with explicit rates for a problem that is usually treated heuristically or under stronger distributional assumptions.

major comments (2)
  1. [Section 2.1 (Assumption 1(a)); Theorems 4.3 and 4.6] The entire identification and all downstream rate guarantees are conditional on the intersection condition ∩_e col([B,A^(e)]) = col(B) (or its quantitative version, Condition 4.2). The manuscript correctly acknowledges that this assumption is unfalsifiable from observational data and proves instance-level necessity in Theorem B.1. However, the practical claims in the title, abstract, and real-data section go beyond this conditional statement. If two or more environments share a heterogeneous loading direction a∉col(B) — a realistic scenario with shared coding conventions, laboratory normalizations, or documentation artifacts — then Condition 4.2 fails with ϵ_A=0, and IHD will include a in the estimated invariant subspace. The recovered 'invariant' factor may have environment-dependent association with Y in unseen environments, and the transfer guarantee in Theorem 4.6 no longer applies.
  2. [Abstract; Section 1.2; Theorems 4.3–4.6] The abstract and introduction describe the non-asymptotic bounds as 'sharp.' The only lower-bound argument in the paper is for the single-environment benchmark in Lemma 4.1, where the λ^{-1/2} noise floor is identified. For the multi-environment results — Theorem 4.3 for invariant/heterogeneous factor recovery, Theorem 4.5 for prediction-invariant factor alignment, and Theorem 4.6 for the transferred prediction error — no minimax lower bounds are provided. Terms such as sqrt(r^(e)/n_x) in δ_FI and the first-order terms in δ_Z are asserted as tight, but no matching lower-bound analysis is offered. The 'sharp' claim is therefore unsupported as stated. Please either provide lower bounds for these multi-environment problems or replace 'sharp' with 'near-oracle'/'non-asymptotic' and specify precisely which rates are known to be optimal.
minor comments (4)
  1. [Section 5.1, Figure 2(d) caption] In the caption and surrounding text, 'ATALS' appears to be a typo for 'ATLAS.'
  2. [Section 3.1, Algorithm 1 and Section 4] The theoretical results require choosing λ_ihd in an interval [C δ_WI, ϵ_A − C δ_WI] that depends on unknown population quantities, but the algorithm takes λ_ihd as an input without a data-driven selection rule. The same issue applies to λ_sel in Section 3.2. In the real-data experiment, λ_IHD is fixed at 0.01 and r^(e)=64 in Appendix E.2; a brief sensitivity analysis for these choices would strengthen the practical claims.
  3. [Theorem 4.6, Eq. following (4.11)] The display '∥Q^{-⊤} bβ − β*∥_2 / eC2' appears to have a typographical/subscript formatting issue; the intended statement is that the norm is bounded by eC2 times δ_y. Please correct the notation.
  4. [Section 6, 'No auxiliary labels Z'] The paragraph correctly assumes F_I ⊥⊥ F_H for the nonparametric no-Z extension. It may be worth stating explicitly that the derivation uses independence, not merely the block-uncorrelatedness used elsewhere, to justify E[g_H(F_SH)|F_I] = E[g_H(F_SH)].

Circularity Check

0 steps flagged

No significant circularity: identification and rates follow from explicit structural assumptions and self-contained spectral/GLM analysis.

full rationale

The paper's derivation chain is self-contained. The central identification result (Theorem 2.1 and Theorem B.1) is derived from the covariance structure of the multi-environment factor model under Assumption 1, which is stated explicitly as an untestable structural condition rather than something fitted from data. The IHD estimator's population justification—that the shared loading space is the unit eigenspace of the averaged per-environment projection matrices—is a direct algebraic consequence of Assumption 1(a). No fitted parameter is later relabeled as a prediction: factor recovery, auxiliary-label alignment via GLM/SVD, and the downstream transfer error bounds in Theorems 4.3, 4.5, and 4.6 are proved from stated conditions and standard regularity assumptions. Self-citations to Fan et al. (2024a) and Gu et al. (2025a) are used as background analogies for the invariance principle and are not load-bearing for the mathematical claims. Proposition 2.2 and Proposition 2.3 define worst-case uncertainty sets in terms of the model, which is a standard minimax formulation rather than a circular derivation. The acknowledged unfalsifiability of Assumption 1(a) is a genuine scope limitation, but it is an identifiability assumption, not a circular step.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The central claims rest on standard factor-model separating assumptions plus two strong domain assumptions: the maximum invariant subspace condition and block uncorrelatedness. The auxiliary-label rank conditions are also load-bearing. No new physical or probabilistic entities are introduced beyond latent factors and the ATLAS projection matrices.

free parameters (4)
  • r^(e) (per-environment factor count) = 64 in real data; known in theory/simulations
    Number of latent factors per environment; treated as known in the theory, set to 64 in the EHR application; not estimated by a data-driven rule with guarantees in this paper.
  • r_I (invariant factor count) = determined by λ_IHD=0.01 in real data
    Number of invariant factors; in Algorithm 1 selected by the threshold λ_IHD on eigenvalues of the averaged projection; in simulations assumed known, in the EHR application fixed implicitly by λ_IHD=0.01.
  • |S_I^*| and |S_H^*| (prediction-invariant subspace dimensions) = assumed known in simulations; unspecified in real data
    Dimensions of the response-invariant subsets; needed to choose the truncated SVD rank in (3.6)-(3.8). The real-data section does not describe how these are selected.
  • λ_IHD, λ_sel = λ_IHD=0.01; λ_sel not reported
    Hyperparameters for selecting r_I and the prediction-invariant ranks. Theory requires them in certain intervals; practice fixes λ_IHD at 0.01 without sensitivity analysis.
axioms (5)
  • domain assumption Assumption 1(a): ∩_e col([B, A^(e)]) = col(B) (maximum invariant subspace).
    Structural assumption that the only loading directions shared by all environments are the true invariant directions; unfalsifiable from X alone. Invoked in Theorem 2.1 and throughout.
  • domain assumption Assumption 1(b): E[F_I^(e) (F_H^(e))^T] = 0 (block uncorrelatedness).
    Separates invariant from heterogeneous factors; used in the partialling-out construction (3.3) and in the proof of Theorem 4.3.
  • domain assumption Condition B.1: rank(Ξ_I^*)=|S_I^*| and rank(Ξ_H^*)=|S_H^*| (exhaustive auxiliary labels).
    Requires q ≥ max(|S_I^*|,|S_H^*|) and that the auxiliary labels are informative for every prediction-relevant factor; otherwise F_{S_H}^* cannot be aligned.
  • standard math Condition 4.1: sub-Gaussian factors and errors, well-conditioned factor covariances.
    Distributional assumptions for the non-asymptotic bounds; standard in weak factor analysis.
  • domain assumption Condition 4.3/C.1: convex GLM loss, well-conditioned Hessians, sub-Gaussian noise for Z and Y.
    Regularity for the GLM steps; standard.

pith-pipeline@v1.3.0-alltime-deepseek · 47486 in / 17333 out tokens · 141772 ms · 2026-08-01T15:40:08.019014+00:00 · methodology

0 comments
read the original abstract

This paper considers a multi-environment factor model in which high-dimensional covariates are collected from heterogeneous environments, with auxiliary labels available in a subset of these environments. The joint distribution of the covariates may vary across environments, whereas the latent structure is decomposed into invariant factors with shared loadings and heterogeneous factors with environment-specific loadings. Such a model is motivated by transfer learning and latent factor regression, where one seeks stable low-dimensional representations for both interpretation and robust out-of-sample prediction of the response $Y$. Leveraging the invariance principle, we show that the invariant and heterogeneous factors are disentangled under a minimal structural condition. Based on this, we propose ATLAS, an Auxiliary-label and invariance-guided Transfer via Latent Alignment across heterogeneous environmentS. ATLAS is a unified procedure that leverages the invariance principle to separate aligned invariant and unaligned heterogeneous factors, and further exploits supervision from auxiliary labels to extract prediction-invariant and transferable factors from those unaligned heterogeneous factors. ATLAS yields near-oracle performance for downstream latent factor regression, enables transferable prediction in new environments through the full latent signal when auxiliary labels are available, and reduces to robust invariant-factor-only prediction otherwise. We establish sharp non-asymptotic error bounds for recovering invariant and heterogeneous factors, identifying all the response-invariant factors, and estimating the invariant signal in $Y$.

Figures

Figures reproduced from arXiv: 2607.18209 by Katherine Liao, Tianxi Cai, Yihong Gu.

Figure 1
Figure 1. Figure 1: A visualization of the solutions yielded by our proposed invariance-heterogeneity decomposition (IHD), PCA on all [PITH_FULL_IMAGE:figures/full_fig_p013_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Panel (a) (resp. (b)–(c)) depict the results on factor estimation of [PITH_FULL_IMAGE:figures/full_fig_p021_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Illustration of covariate/label construction. [PITH_FULL_IMAGE:figures/full_fig_p041_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

73 extracted references · 9 linked inside Pith

  1. [1]

    Agarwal, A., Negahban, S., & Wainwright, M. J. (2012). Noisy matrix decomposition via convex relaxation: Optimal rates in high dimensions. The Annals of Statistics , 40(2), 1171--1197

  2. [2]

    Ahn, S. C. & Horenstein, A. R. (2013). Eigenvalue ratio test for the number of factors. Econometrica , 81(3), 1203--1227

  3. [3]

    & Smolen, J

    Aletaha, D. & Smolen, J. S. (2018). Diagnosis and management of rheumatoid arthritis: a review. Jama , 320(13), 1360--1372

  4. [4]

    C., Hare, A

    Apathy, N. C., Hare, A. J., Fendrich, S., & Cross, D. A. (2022). Early changes in billing and notes after evaluation and management guideline change. Annals of internal medicine , 175(4), 499--504

  5. [5]

    Arjovsky, M., Bottou, L., Gulrajani, I., & Lopez-Paz, D. (2019). Invariant risk minimization. arXiv preprint arXiv:1907.02893

  6. [6]

    Z., Nicol, P

    Baharav, T. Z., Nicol, P. B., Irizarry, R. A., & Ma, R. (2025). Stacked svd or svd stacked? a random matrix theory perspective on data integration. arXiv preprint arXiv:2507.22170

  7. [7]

    Bai, J. (2003). Inferential theory for factor models of large dimensions. Econometrica , 71(1), 135--171

  8. [8]

    Bai, J. & Ng, S. (2002). Determining the number of factors in approximate factor models. Econometrica , 70(1), 191--221

  9. [9]

    Bai, J. & Ng, S. (2006). Confidence intervals for diffusion index forecasts and inference for factor-augmented regressions. Econometrica , 74(4), 1133--1150

  10. [10]

    Bai, J. & Ng, S. (2008). Forecasting economic time series using targeted predictors. Journal of Econometrics , 146(2), 304--317

  11. [11]

    Bai, J. & Ng, S. (2023). Approximate factor models with weaker loadings. Journal of Econometrics , 235(2), 1893--1916

  12. [12]

    Bair, E., Hastie, T., Paul, D., & Tibshirani, R. (2006). Prediction by supervised principal components. Journal of the American Statistical Association , 101(473), 119--137

  13. [13]

    B \"u hlmann, P. (2020). Invariance, causality and robustness. Statistical Science , 35(3), 404--426

  14. [14]

    K., Dunson, D

    Chandra, N. K., Dunson, D. B., & Xu, J. (2025). Inferring covariance structure from multiple data sources via subspace factor analysis. Journal of the American Statistical Association , 120(550), 1239--1253

  15. [15]

    & B \"u hlmann, P

    Chen, Y. & B \"u hlmann, P. (2021). Domain adaptation under structural causal models. Journal of Machine Learning Research , 22(261), 1--80

  16. [16]

    Chen, Y., Chi, Y., Fan, J., & Ma, C. (2021). Spectral methods for data science: A statistical perspective. Foundations and Trends in Machine Learning , 14(5), 566--806

  17. [17]

    Chen, Y., Rosenfeld, E., Sellke, M., Ma, T., & Risteski, A. (2022). Iterative feature matching: Toward provable domain generalization with logarithmic environments. Advances in neural information processing systems , 35, 1725--1736

  18. [18]

    C., Hanberg, J

    Cheng, D., Wang, X., McDermott, G. C., Hanberg, J. S., Love, Z., Zhong, K., Jeffway, M., Hou, J., Panickan, V., Sangar, R., et al. (2025). Inferring rheumatoid arthritis disease activity status from the electronic health records across health systems to enable real-world data studies. medRxiv , (pp.\ 2025--11)

  19. [19]

    & Yuan, M

    Choi, J. & Yuan, M. (2025). High dimensional factor analysis with weak factors. Journal of Econometrics , 252, 106086

  20. [20]

    De Vito, R., Bellio, R., Trippa, L., & Parmigiani, G. (2019). Multi-study factor analysis. Biometrics , 75(1), 337--346

  21. [21]

    De Vito, R., Bellio, R., Trippa, L., & Parmigiani, G. (2021). Bayesian multistudy factor analysis for high-throughput biological data. The annals of applied statistics , 15(4), 1723--1741

  22. [22]

    Fan, J., Fang, C., Gu, Y., & Zhang, T. (2024a). Environment invariant linear least squares. The Annals of Statistics , 52(5), 2268--2292

  23. [23]

    Fan, J. & Gu, Y. (2024). Factor augmented sparse throughput deep relu neural networks for high dimensional regression. Journal of the American Statistical Association , 119(548), 2680--2694

  24. [24]

    & Liao, Y

    Fan, J. & Liao, Y. (2022). Learning latent factors from diversified projections and its applications to over-estimated and weak factors. Journal of the American Statistical Association , 117(538), 909--924

  25. [25]

    Fan, J., Liao, Y., & Mincheva, M. (2013). Large covariance estimation by thresholding principal orthogonal complements. Journal of the Royal Statistical Society. Series B, Statistical methodology , 75(4)

  26. [26]

    Fan, J., Wang, D., Wang, K., & Zhu, Z. (2019). Distributed estimation of principal eigenspaces. Annals of statistics , 47(6), 3009

  27. [27]

    Fan, J., Xue, L., & Yao, J. (2017). Sufficient forecasting using factor models. Journal of econometrics , 201(2), 292--306

  28. [28]

    Fan, J., Yan, Y., & Zheng, Y. (2024b). When can weak latent factors be statistically inferred? arXiv preprint arXiv:2407.03616

  29. [29]

    Feng, Q., Jiang, M., Hannig, J., & Marron, J. (2018). Angle-based joint and individual variation explained. Journal of multivariate analysis , 166, 241--265

  30. [30]

    Flury, B. N. (1984). Common principal components in k groups. Journal of the American Statistical Association , 79(388), 892--898

  31. [31]

    Forni, M., Hallin, M., Lippi, M., & Reichlin, L. (2000). The generalized dynamic-factor model: Identification and estimation. Review of Economics and statistics , 82(4), 540--554

  32. [32]

    N., De Vito, R., Trippa, L., & Parmigiani, G

    Grabski, I. N., De Vito, R., Trippa, L., & Parmigiani, G. (2023). Bayesian combinatorial multistudy factor analysis. The annals of applied statistics , 17(3), 2212

  33. [33]

    Gu, Y., Fang, C., B \"u hlmann, P., & Fan, J. (2025a). Causality pursuit from heterogeneous environments via neural adversarial invariance learning . The Annals of Statistics , 53(5), 2230 -- 2257

  34. [34]

    Gu, Y., Fang, C., Xu, Y., Guo, Z., & Fan, J. (2025b). Fundamental computational limits in pursuing invariant causal prediction and invariance-guided regularization. arXiv preprint arXiv:2501.17354

  35. [35]

    & Li s ka, R

    Hallin, M. & Li s ka, R. (2007). Determining the number of factors in the general dynamic factor model. Journal of the American Statistical Association , 102(478), 603--617

  36. [36]

    Hallin, M., Paindaveine, D., & Verdebout, T. (2014). Efficient r-estimation of principal and common principal components. Journal of the American Statistical Association , 109(507), 1071--1083

  37. [37]

    He, Y., Liu, D., Sun, Y., & Wang, Y. (2025). Transpca for large-dimensional factor analysis with weak factors: Power enhancement via knowledge transfer. arXiv preprint arXiv:2503.08397

  38. [38]

    Huber, P. J. (1985). Projection pursuit. The Annals of Statistics , (pp.\ 435--475)

  39. [39]

    Jiang, P., Uematsu, Y., & Yamagata, T. (2023). Revisiting asymptotic theory for principal component estimators of approximate factor models. arXiv preprint arXiv:2311.00625

  40. [40]

    Khemakhem, I., Kingma, D., Monti, R., & Hyvarinen, A. (2020). Variational autoencoders and nonlinear ica: A unifying framework. In International conference on artificial intelligence and statistics (pp.\ 2207--2217).: PMLR

  41. [41]

    & Langer, S

    Kohler, M. & Langer, S. (2021). On the rate of convergence of fully connected deep neural network regression estimates. The Annals of Statistics , 49(4), 2231--2249

  42. [42]

    Kong, L., Xie, S., Yao, W., Zheng, Y., Chen, G., Stojanov, P., Akinwande, V., & Zhang, K. (2022). Partial identifiability for domain adaptation. Proceedings of Machine Learning Research , 162, 11455--11472

  43. [43]

    A., Strobl, E

    Lasko, T. A., Strobl, E. V., & Stead, W. W. (2024). Why do probabilistic clinical models fail to transport between sites. NPJ Digital Medicine , 7(1), 53

  44. [44]

    Li, Z., Cai, R., Chen, G., Sun, B., Hao, Z., & Zhang, K. (2023). Subspace identification for multi-source domain adaptation. Advances in Neural Information Processing Systems , 36, 34504--34518

  45. [45]

    F., Hoadley, K

    Lock, E. F., Hoadley, K. A., Marron, J. S., & Nobel, A. B. (2013). Joint and individual variation explained (jive) for integrated analysis of multiple data types. The annals of applied statistics , 7(1), 523

  46. [46]

    Ma, Z. & Ma, R. (2026). Optimal estimation of shared singular subspaces across multiple noisy matrices. IEEE Transactions on Information Theory

  47. [47]

    R., Cutrona, S

    Molloy-Paolillo, B., Mohr, D., Levy, D. R., Cutrona, S. L., Anderson, E., Rucci, J., Helfrich, C., Sayre, G., & Rinne, S. T. (2023). Assessing electronic health record (ehr) use during a major ehr transition: an innovative mixed methods approach. Journal of General Internal Medicine , 38(Suppl 4), 999

  48. [48]

    Moran, G. E. & Krishnan, A. (2026). Nonlinear multi-study factor analysis. arXiv preprint arXiv:2601.18128

  49. [49]

    F., Wendt, C

    Palzer, E. F., Wendt, C. H., Bowler, R. P., Hersh, C. P., Safo, S. E., & Lock, E. F. (2022). sjive: supervised joint and individual variation explained. Computational statistics & data analysis , 175, 107547

  50. [50]

    Peters, J., B \"u hlmann, P., & Meinshausen, N. (2016). Causal inference by using invariant prediction: identification and confidence intervals. Journal of the Royal Statistical Society. Series B (Statistical Methodology) , (pp.\ 947--1012)

  51. [51]

    u hlmann, P., & Sch \

    Pfister, N., Weichwald, S., B \"u hlmann, P., & Sch \"o lkopf, B. (2019). Robustifying independent component analysis by adjusting for group-wise stationary noise. Journal of Machine Learning Research , 20(147), 1--50

  52. [52]

    M., Williams, R

    Reps, J. M., Williams, R. D., Schuemie, M. J., Ryan, P. B., & Rijnbeek, P. R. (2022). Learning patient-level prediction models across multiple healthcare databases: evaluation of ensembles for increasing model transportability. BMC medical informatics and decision making , 22(1), 142

  53. [53]

    Rosenbaum, P. R. & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika , 70(1), 41--55

  54. [54]

    Rosenfeld, E., Ravikumar, P., & Risteski, A. (2021). The risks of invariant risk minimization. In International Conference on Learning Representations , volume 9

  55. [55]

    Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of educational Psychology , 66(5), 688

  56. [56]

    & Oberhauser, H

    Schell, A. & Oberhauser, H. (2023). Nonlinear independent component analysis for discrete-time and continuous-time signals. The Annals of Statistics , 51(2), 487--518

  57. [57]

    Schmidt-Hieber, J. (2020). Nonparametric regression using deep neural networks with relu activation function (with discussion). The Annals of Statistics , 48(4), 1875--1921

  58. [58]

    Seiter, B., Fries, A., von K \"u gelgen, J., & Peters, J. (2026). Anchor pca. arXiv preprint arXiv:2606.06233

  59. [59]

    & Al Kontar, R

    Shi, N. & Al Kontar, R. (2024). Personalized pca: Decoupling shared and unique features. Journal of machine learning research , 25(41), 1--82

  60. [60]

    S., Landew \'e , R., Breedveld, F

    Smolen, J. S., Landew \'e , R., Breedveld, F. C., et al. (2016). EULAR recommendations for the management of rheumatoid arthritis with synthetic and biological disease-modifying antirheumatic drugs: 2015 update. Annals of the Rheumatic Diseases , 75(3), 473--491

  61. [61]

    Stock, J. H. & Watson, M. W. (2002). Forecasting using principal components from a large number of predictors. Journal of the American statistical association , 97(460), 1167--1179

  62. [62]

    Stojanov, P., Li, Z., Gong, M., Cai, R., Carbonell, J., & Zhang, K. (2021). Domain adaptation with invariant representation learning: What transformations to learn? Advances in neural information processing systems , 34, 24791--24803

  63. [63]

    Torab-Miandoab, A., Samad-Soltani, T., Jodati, A., & Rezaei-Hachesu, P. (2023). Interoperability of heterogeneous health information systems: a systematic literature review. BMC medical informatics and decision making , 23(1), 18

  64. [64]

    Wang, P., Wang, H., Li, Q., Shen, D., & Liu, Y. (2024). Joint and individual component regression. Journal of Computational and Graphical Statistics , 33(3), 763--773

  65. [65]

    Wang, Z., Liu, M., Lei, J., Bach, F., & Guo, Z. (2025). Stablepca: Learning shared representations across multiple sources via minimax optimization. arXiv preprint arXiv:2505.00940

  66. [66]

    & Veitch, V

    Wang, Z. & Veitch, V. (2022). The causal structure of domain invariant supervised representation learning. arXiv preprint arXiv:2208.06987

  67. [67]

    Xia, Y. (2008). A multiple-index model and dimension reduction. Journal of the American Statistical Association , 103(484), 1631--1640

  68. [68]

    W., Cai, T., & Liu, M

    Xiong, X., Guo, Z., Zhu, H., Hong, C., Smoller, J. W., Cai, T., & Liu, M. (2026). Adversarial drift-aware predictive transfer: Toward durable clinical ai. arXiv preprint arXiv:2601.11860

  69. [69]

    M., Liu, M., Hong, C., Bonzel, C.-L., Panickan, V

    Xiong, X., Sweet, S. M., Liu, M., Hong, C., Bonzel, C.-L., Panickan, V. A., Zhou, D., Wang, L., Costa, L., Ho, Y.-L., et al. (2023). Knowledge-driven online multimodal automated phenotyping system. MedRxiv , (pp.\ 2023--09)

  70. [70]

    Yan, Y., Chen, Y., & Fan, J. (2024). Inference for heteroskedastic pca with missing data. The Annals of Statistics , 52(2), 729--756

  71. [71]

    Yang, Y. & Ma, C. (2025). Estimating shared subspace with ajive: the power and limitation of multiple data matrices. arXiv preprint arXiv:2501.09336

  72. [72]

    & Steele, R

    Young, Z. & Steele, R. (2022). Empirical evaluation of performance degradation of machine learning-based predictive models--a case study in healthcare information systems. International Journal of Information Management Data Insights , 2(1), 100070

  73. [73]

    T., Zhang, K., & Gordon, G

    Zhao, H., Des Combes, R. T., Zhang, K., & Gordon, G. (2019). On learning invariant representations for domain adaptation. In International conference on machine learning (pp.\ 7523--7532).: PMLR