REVIEW 3 major objections 5 minor 70 references
Robust and Adaptive Spectral Method for Representation Multi-Task Learning with Contamination
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A thresholded spectral method learns a shared multi-task representation even when an unknown, large fraction of tasks is contaminated, and its per-task estimates never fall below single-task accuracy.
desk verdict A useful adaptive spectral method for contaminated MTL with a promising h=0 theory, but the h>0 case has a proof gap and the implemented threshold isn't shown to satisfy the theoretical requirement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the perturbation matrix $N = \widehat{B}_{\mathrm{st}} - \widetilde{B}$, the difference between the single-task coefficient estimates and the ideal signal matrix, together with the threshold $\tau \asymp \|N\|_{\mathrm{op}}$ that cuts the singular values of $\widehat{B}_{\mathrm{st}}/\sqrt{T}$. The threshold performs two jobs at once: it filters out contamination and noise while selecting, without knowing the rank $r$, only directions whose signal clears the noise-and-heterogeneity floor. The minimum with the single-task rate in Theorem 9 is what converts thresholding into a no-negative-transfer guarantee, and a final biased-regularization step anchors per-task estimates to the learned subspace, falling back to the single-task rate when the multi-task bound is unfavorable.
What would settle it
Run RAS with its fixed threshold $\tau = 2.5\sqrt{(p+T)/(nT)}$ in a regime inside the paper's assumptions where inlier heterogeneity $h$ is nonzero and inlier signals are near the signal-to-noise threshold, and check whether the selected rank $\hat{k}$ misses inlier directions or admits outlier directions and whether the inlier error exceeds $\sqrt{(p+\log T)/n}$ with high probability. Exhibiting such an instance would show the implemented threshold violates the inequality that Proposition 4 and Theorem 9 require.
Extended reading notes
Core claim
The paper claims that unknown, arbitrarily structured task contamination does not have to break representation-based multi-task learning. RAS estimates the shared subspace from the leading singular vectors of the single-task coefficient matrix after an adaptive singular-value threshold chosen to dominate the perturbation matrix $N$ (estimation noise plus inlier heterogeneity plus outlier contamination). With the threshold set that way and a mild spectral gap, the estimated rank equals an effective signal rank and the subspace error is controlled by the inlier diversity $\sigma_{\min,\mathrm{in}}$, the heterogeneity $h$, and the contamination fraction $\varepsilon$. The central result, Theorem 9, bounds each inlier task's $\ell^2$ estimation error by the minimum of a multi-task rate depending on the estimated rank and the single-task rate $\sqrt{(p+\log T)/n}$, which is what guarantees no negative transfer; the only requirement on contamination is $\varepsilon < 1$, with no prior knowledge of $\varepsilon$ or $r$.
Load-bearing premise
The guarantees require the singular-value threshold to be at least as large as the worst-case perturbation that contamination, heterogeneity, and noise create in the coefficient matrix, and the threshold that provably works depends on the contamination fraction, the spread among inlier tasks, and their signal strengths — none of which are known; the fixed threshold used in practice is not shown to be large enough in those harder regimes.
Editorial extensions
If this is right
- Every inlier task keeps an estimation error of at most the single-task rate $\sqrt{(p+\log T)/n}$ with high probability, so representation sharing provably cannot hurt, no matter the contamination fraction (as long as it is below 1) and with no knowledge of the true rank.
- RAS needs neither the contamination proportion nor the intrinsic dimension as input, unlike the spectral method it benchmarks against; the threshold selects the rank automatically.
- Low-rank outlier structure is exploited: the estimated rank satisfies $\hat{k} \le r + r_{\mathrm{out}} - r_{\cap}$, improving the bound, while general-rank outliers degrade the bound smoothly instead of breaking the method.
- In transfer learning, the target-task error is the minimum of the single-task rate $\sqrt{p/n_0}$ and a representation-based rate, so reusing the learned representation is safe.
- Under standard scaling ($\zeta^{(t)}\asymp 1$, $\bar{\zeta}\asymp 1$, balanced outlier spectrum), RAS beats single-task learning when $r \ll |S| \wedge p$, the outlier dimension or mass is small relative to $p$, and inlier heterogeneity $h$ is small.
Reading between the lines
- The min-with-single-task structure makes RAS an automatic switch: a user can deploy it without deciding beforehand whether tasks are worth sharing, because in any regime where sharing fails the bound collapses to the single-task rate.
- Because the provably valid threshold depends on unknown quantities ($\varepsilon$, $h$, signal strengths) while the implemented threshold is a fixed heuristic, a natural testable extension is data-driven threshold calibration — for example sample-splitting or bootstrap on singular-value gaps — that would make the finite-sample guarantee fully parameter-free.
- The estimated rank $\hat{k}$ could double as a diagnostic of contamination structure: values near $r$ indicate low-rank outliers below the noise floor, while values near $r + r_{\mathrm{out}}$ reveal strong outlier signals that the threshold accepted.
- Since the argument only requires $\tau$ to dominate $\|N\|_{\mathrm{op}}$, and the paper notes this reasoning is model-agnostic, the same recipe plausibly extends to nonlinear losses whenever a concentration bound for the single-task estimates is available.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies representation-based multi-task learning when an unknown, possibly large fraction of tasks are contaminated, and inlier tasks may have heterogeneous representations. It proposes the Robust and Adaptive Spectral method (RAS): fit single-task regressions, form the coefficient matrix, perform an SVD with a data-driven singular-value threshold to select the shared subspace and its dimension, and then apply a biased regularization step that anchors per-task estimates to the learned subspace. The main theoretical results are non-asymptotic bounds on the subspace estimation error and per-task coefficient errors, a guarantee that the method performs no worse than single-task learning via an explicit minimum with the single-task rate, and an extension to transfer learning. The experiments cover low-rank and general-rank outliers, contamination proportions up to 80%, the effect of winsorization, and scaling with the number of tasks.
Significance. If the theoretical claims hold, this is a valuable contribution: RAS requires neither the contamination proportion nor the true representation dimension, handles arbitrary outlier tasks, and its no-negative-transfer guarantee is a practically important improvement over methods whose error bounds degrade with the contamination level. The paper provides detailed proofs in the appendix, and the experimental section is extensive, including comparisons against oracle methods that know the true rank and the contamination proportion. The conceptual idea of using an adaptive threshold calibrated to the perturbation level, and the explicit single-task fallback via the min in Theorem 9, are strengths. However, the central adaptivity claim is currently established only for an oracle threshold that depends on unknown quantities, and the proof of the subspace bound has a gap in the heterogeneous-representation case, so the advertised adaptivity in the h>0 regime is not yet fully supported.
major comments (3)
- [Section 2.2, Proposition 4, Theorem 9] The theoretical threshold τ is defined as τ ≍ sqrt((p+|S|)/(nT)) + h \barζ / sqrt(1-ε) · (σ_max(D*_S)/(√r σ_min(D*_S)) ∧ 1), which depends on the unknown contamination proportion ε, the heterogeneity h, and inlier signal strengths. The proof of Proposition 4 establishes the required domination ∥N/√T∥_op < 0.25τ only for this oracle choice. In Section 5 the implementation uses the fixed heuristic τ = 2.5√((p+T)/(nT)), which dominates the first term (since T ≥ |S|) but omits the h-dependent heterogeneity term. Consequently, when h > 0 and that term is non-negligible, the premise of Proposition 4, and therefore the rank-selection guarantee, the subspace bound of Proposition 7, and the adaptive first branch of Theorem 9, are not shown to hold for the actual algorithm. Since all experiments in Section 5 use h = 0, this gap is not probed empirically. The authors should either prove that the heuristic threshold satisfies the domination condition under explicit conditions, or provide a genuinely data-dependent threshold that adapts to the unknown heterogeneity and contamination level, with a matching theory.
- [Appendix A.3, Proposition 7] The proof of Proposition 7 claims the identity (bA^⊥)^T eB_S = ((bA^⊥)^T A) Θ*_S and then lower-bounds the left-hand side by ∥(bA^⊥)^T A∥_op · σ_r(Θ*_S) using Assumption 2. However, by the definitions in Section 2.1, eB_{:,t} = A A^T β(t)* = A A^T A^{(t)*} θ^{(t)*} for t ∈ S, so the matrix on the left is (bA^⊥)^T A · [A^T A^{(t)*} θ^{(t)*}]_{t∈S}, not (bA^⊥)^T A · [θ^{(t)*}]_{t∈S}, unless A^{(t)*} = A for every inlier task. The smallest singular value of the matrix with columns A^T A^{(t)*} θ^{(t)*} depends on the principal angles between A and the individual A^{(t)*} and can be substantially smaller than σ_r(Θ*_S). Thus the displayed subspace error bound does not follow for h > 0, and the ℓ₂ rate in Theorem 9 that relies on this bound is not established in the heterogeneous setting. Please either redefine Θ*_S to include the rotations A^T A^{(t)*} and derive the resulting bound with explicit dependence on h, or state Proposition 7 and the adaptive branch of Theorem 9 only for h = 0 with a separate treatment for the general case.
- [Theorem 9 and Lemma 6] The adaptive first rate in Theorem 9 uses k⋆ = min{r + r_out, r + Tε, p, T} as an upper bound on the estimated rank k̂, but r_out is not defined in the statement of Theorem 9. Moreover, the bound k̂ ≤ r + r_out − r∩ in Lemma 6 requires the outlier coefficient matrix (bBst)_{Sc} to decompose as a rank-r_out matrix plus a perturbation of operator norm at most (3/4)τ. This low-rank-plus-small-perturbation condition is not among the assumptions of Theorem 9, which only states that outlier tasks follow an arbitrary distribution. Without it, k̂ can be as large as min(p,T), and the claimed adaptivity to outlier structure is not proven in general; only the single-task fallback in the second branch of the minimum is unconditional. The authors should either explicitly incorporate the low-rank-outlier condition into Theorem 9, or present the adaptive rate as conditional on Lemma 6 and clarify what, if anything, is guaranteed for general-rank outliers.
minor comments (5)
- [Section 1.2, Section 5.4] There are several typos: "theoretical ganrantees" in Section 1.2, "kernal" in the related literature, and "varing" in Section 5.4; these should be corrected.
- [Figures 1-6] Several axis labels in the figures appear as garbled Unicode escape sequences (for example "/uni00000013/uni00000011/..."), making the labels unreadable; the figures should be regenerated with proper text rendering.
- [Section 2.2] The expression for the oracle threshold τ is given in prose rather than as a numbered equation; placing it in a displayed, numbered equation would make the cross-references in Section 3 and the appendix much easier to follow.
- [Section 2.1] The notation B*_S = {β(t)*} and B_S = {A A^T β(t)*} is ambiguous because B_S is later used as a matrix; using distinct symbols for the set of vectors and the matrix would improve clarity.
- [Proposition 4] The spectral gap assumption is stated as eλ_{r_eff+1} < 0.75τ while the effective signal rank is defined with the level 1.25τ; a remark explaining how these constants relate would help the reader understand the role of the gap condition.
Circularity Check
No circular reduction found; the oracle-threshold gap and Proposition 7 identity issue are correctness concerns, not circularity.
full rationale
The paper's theoretical claims are derived from explicit assumptions, and I could not exhibit any place where a predicted quantity equals a fitted input by construction. The threshold tau in Section 2.2 is an oracle quantity (tau approximately the operator norm of N, with a bound in terms of epsilon, h, zeta, and D*_S); it is not estimated from the data whose error is later reported as a prediction. Theorems 9 and Proposition 11 are conditional on that assumed threshold, so the practical concern that the implemented tau = 2.5 sqrt((p+T)/(nT)) (Section 5) may not dominate the h-dependent perturbation is a gap between theory and implementation, not a circular definition. Proposition 4's rank consistency is not forced by definition: reff is defined through singular values of eB (a population-like matrix) while k-hat is defined through bB_st, and the proposition supplies a perturbation argument (Weyl inequality plus concentration of N) linking them; the factor 1.25 tau versus tau gives it content. The no-negative-transfer fallback relies on Lemma 15 from Tian et al. (2023), which is co-authored by one of the present authors; however, that lemma is stated and proved under standard assumptions and does not assume the target bound, so under the review rules it counts as independent support rather than self-citation circularity. I also noted the Proposition 7 proof's identity (bA^perp)^T eB_S = (bA^perp)^T A Theta*_S requires reinterpreting Theta*_S as A^T B*_S rather than the theta(t)* defined via beta(t)* = A(t)* theta(t)*; this is a proof gap for h > 0, but it is an algebraic mismatch, not a circular reduction. Overall, the derivation chain is not circular; the advertised adaptivity is stronger than what the implemented tuning provably delivers, but that is a robustness and correctness concern that should be scored separately.
Assumptions & free parameters
free parameters (2)
- Threshold constant in tau =
2.5 in experiments; C' in theory (unspecified)
- Regularization parameter gamma =
0.5 sqrt(p+log T) in experiments; C' sqrt(p+log T) in theory
assumptions (7)
- domain assumption Features are sub-Gaussian with covariance bounded between c_min and c_max (Assumption 1)
- domain assumption Inlier task diversity: sigma_r(Theta*_S / sqrt(|S|)) >= sigma_min,in > 0 (Assumption 2)
- domain assumption n >= C(p+log T) (Assumption 3)
- domain assumption Spectral gap: lambda_tilde_{r_eff+1} < 0.75 tau in Proposition 4
- domain assumption sigma_min,in > 1.25 tau in Theorem 9
- domain assumption Target task assumptions (Assumptions 4-5) for transfer learning
- ad hoc to paper The lower-bound inequality ||X Theta||_op >= ||X||_op sigma_r(Theta) in the proof of Proposition 7
Cite this review
Pith. "Pith review of Robust and Adaptive Spectral Method for Representation Multi-Task Learning with Contamination." pith.science (2026). https://pith.science/paper/4ZXCXFJL
@misc{pith2026250906575,
author = {Pith},
title = {Pith review of: Robust and Adaptive Spectral Method for Representation Multi-Task Learning with Contamination},
year = {2026},
howpublished = {\url{https://pith.science/paper/4ZXCXFJL}},
note = {Machine review of arXiv:2509.06575}
}
read the original abstract
Representation-based multi-task learning (MTL) improves efficiency by learning a shared structure across tasks, but its practical application is often hindered by contamination, outliers, or adversarial tasks. Most existing methods and theories assume a clean or near-clean setting, failing when contamination is significant. This paper tackles representation MTL with an unknown and potentially large contamination proportion, while also allowing for heterogeneity among inlier tasks. We introduce a Robust and Adaptive Spectral method (RAS) that can distill the shared inlier representation effectively and efficiently, while requiring no prior knowledge of the contamination level or the true representation dimension. Theoretically, we provide non-asymptotic error bounds for both the learned representation and the per-task parameters. These bounds adapt to inlier task similarity and outlier structure, and guarantee that RAS performs at least as well as single-task learning, thus preventing negative transfer. We also extend our framework to transfer learning with corresponding theoretical guarantees for the target task. Extensive experiments confirm our theory, showcasing the robustness and adaptivity of RAS, and its superior performance in regimes with up to 80\% task contamination.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Ando, R. K. and Zhang, T. (2005). A framework for learning predictive structures from multiple tasks and unlabeled data. Journal of Machine Learning Research, 6 1817--1853
work page 2005
-
[2]
Argyriou, A. , Evgeniou, T. and Pontil, M. (2008). Convex multi-task feature learning. Machine Learning, 73 243--272
work page 2008
-
[3]
Barzilai, A. and Crammer, K. (2015). Convex multi-task learning by clustering. In Proceedings of the Eighteenth International Conference on Artificial Intelligence and Statistics, vol. 38 of Proceedings of Machine Learning Research. PMLR, San Diego, California, USA, 65--73
work page 2015
-
[4]
Baxter, J. (2000). A model of inductive bias learning. Journal of Artificial Intelligence Research, 12 149--198
work page 2000
-
[5]
Ben-David, S. , Blitzer, J. , Crammer, K. , Kulesza, A. , Pereira, F. and Vaughan, J. W. (2010). A theory of learning from different domains. Machine Learning, 79 151--175
work page 2010
-
[6]
Bengio, Y. , Courville, A. and Vincent, P. (2013). Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35 1798--1828
work page 2013
-
[7]
Blanchard, P. , Mhamdi, E. M. E. , Guerraoui, R. and Stainer, J. (2017). Machine learning with adversaries: Byzantine tolerant gradient descent. In Advances in Neural Information Processing Systems 30 (NeurIPS)
work page 2017
-
[8]
Cand \`e s, E. J. , Li, X. , Ma, Y. and Wright, J. (2011). Robust principal component analysis? Journal of the ACM, 58 11:1--11:37
work page 2011
Show all 70 references
-
[9]
Caruana, R. (1997). Multitask learning. Machine Learning, 28 41--75
1997
-
[10]
Catoni, O. (2012). Challenging the empirical mean and empirical variance: A deviation study. Annales de l'Institut Henri Poincar \'e , Probabilit \'e s et Statistiques , 48 1148--1185
2012
-
[11]
, Chi, Y
Chen, Y. , Chi, Y. , Fan, J. , Ma, C. et al. (2021). Spectral methods for data science: A statistical perspective. Foundations and Trends in Machine Learning , 14 566--806
2021
-
[12]
, Lei, Q
Chua, K. , Lei, Q. and Lee, J. D. (2021). How fine-tuning allows for effective meta-learning. Advances in Neural Information Processing Systems, 34 8871--8884
2021
-
[13]
, Kearns, M
Crammer, K. , Kearns, M. and Wortman, J. (2008). Learning from multiple sources. Journal of Machine Learning Research, 9
2008
-
[14]
Crawshaw, M. (2020). Multi-task learning with deep neural networks: A survey. arXiv preprint arXiv:2009.09796. ://arxiv.org/abs/2009.09796
2020 arXiv
-
[15]
Du, S. S. , Hu, W. , Kakade, S. M. , Lee, J. D. and Lei, Q. (2020). Few-shot learning via learning the representation, provably. arXiv preprint arXiv:2002.09434
2020 arXiv
-
[16]
and Wang, K
Duan, Y. and Wang, K. (2023). Adaptive and robust multi-task learning. The Annals of Statistics, 51 2015--2039
2023
-
[17]
, Feldman, V
Duchi, J. , Feldman, V. , Hu, L. and Talwar, K. (2022). Subspace recovery from heterogeneous data with non-isotropic noise. arXiv preprint arXiv:2210.13497
2022 arXiv
-
[18]
, Micchelli, C
Evgeniou, T. , Micchelli, C. A. and Pontil, M. (2005). Learning multiple tasks with kernel methods. Journal of Machine Learning Research, 6 615--637. ://jmlr.org/papers/v6/evgeniou05a.html
2005
-
[19]
and Pontil, M
Evgeniou, T. and Pontil, M. (2004). Regularized multi-task learning. In Proceedings of the 10th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 109--117
2004
-
[20]
Gong, P. , Ye, J. and Zhang, C. (2012). Robust multi-task feature learning. In Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining. 895--903
2012
-
[21]
, Han, Y
Gu, T. , Han, Y. and Duan, R. (2024). Robust angle-based transfer learning in high dimensions. Journal of the Royal Statistical Society Series B: Statistical Methodology qkae111
2024
-
[22]
Huber, P. J. (1964). Robust estimation of a location parameter. Annals of Mathematical Statistics, 35 73--101
1964
-
[23]
, Bach, F
Jacob, L. , Bach, F. and Vert, J.-P. (2008). Clustered multi-task learning: A convex formulation. In Advances in Neural Information Processing Systems. 745--752
2008
-
[24]
, Sanghavi, S
Jalali, A. , Sanghavi, S. , Ruan, C. and Ravikumar, P. (2010). A dirty model for multi-task learning. Advances in neural information processing systems, 23
2010
-
[25]
and Ye, J
Ji, S. and Ye, J. (2009). An accelerated gradient method for trace norm minimization. In Proceedings of the 26th International Conference on Machine Learning. ACM, 457--464. Montreal, Canada
2009
-
[26]
Johnson, W. E. , Li, C. and Rabinovic, A. (2007). Adjusting batch effects in microarray expression data using empirical bayes methods. Biostatistics, 8 118--127
2007
-
[27]
, Grauman, K
Kang, Z. , Grauman, K. and Sha, F. (2011). Learning with whom to share in multi-task feature learning. In Proceedings of the 28th International Conference on Machine Learning. Bellevue, WA, USA, 521--528
2011
-
[28]
Katzman, J. L. , Shaham, U. , Cloninger, A. , Bates, J. , Jiang, T. and Kluger, Y. (2018). Deepsurv: Personalized treatment recommender system using a cox proportional hazards deep neural network. BMC Medical Research Methodology, 18 24
2018
-
[29]
Kelley, D. R. , Snoek, J. and Rinn, J. L. (2016). Basset: Learning the regulatory code of the accessible genome with deep convolutional neural networks. Genome Research, 26 990--999
2016
-
[30]
, Gal, Y
Kendall, A. , Gal, Y. and Cipolla, R. (2018). Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7482--7491
2018
-
[31]
Kingma, D. P. and Ba, J. L. (2015). Adam: A method for stochastic gradient descent. In ICLR: international conference on learning representations. ICLR US., 1--15
2015
-
[32]
and Daum \'e III, H
Kumar, A. and Daum \'e III, H. (2012). Learning task grouping and overlap in multi-task learning. In Proceedings of the 29th International Conference on Machine Learning. 1723--1730
2012
-
[33]
and Orabona, F
Kuzborskij, I. and Orabona, F. (2013). Stability and hypothesis transfer learning. In International Conference on Machine Learning. PMLR, 942--950
2013
-
[34]
, Bengio, Y
LeCun, Y. , Bengio, Y. and Hinton, G. (2015). Deep learning. Nature, 521 436--444
2015
-
[35]
, Zame, W
Lee, C. , Zame, W. , Yoon, J. and van der Schaar, M. (2018). Deephit: A deep learning approach to survival analysis with competing risks. In Proceedings of the AAAI Conference on Artificial Intelligence. 2314--2321
2018
-
[36]
Leek, J. T. , Scharpf, R. B. , Bravo, H. C. , Simcha, D. , Langmead, B. , Johnson, W. E. , Geman, D. , Baggerly, K. and Irizarry, R. A. (2010). Tackling the widespread and critical impact of batch effects in high-throughput data. Nature Reviews Genetics, 11 733--739
2010
-
[37]
, Pontil, M
Lounici, K. , Pontil, M. , Tsybakov, A. B. and van de Geer, S. (2011). Oracle inequalities and optimal inference under group sparsity. Annals of Statistics, 39 2164--2204
2011
-
[38]
and Mendelson, S
Lugosi, G. and Mendelson, S. (2019). Mean estimation and regression under heavy-tailed distributions---a survey. Foundations of Computational Mathematics, 19 1145--1190
2019
-
[39]
, Mohri, M
Mansour, Y. , Mohri, M. and Rostamizadeh, A. (2009). Domain adaptation: Learning bounds and algorithms. arXiv preprint arXiv:0902.3430. ://arxiv.org/abs/0902.3430
2009 arXiv
-
[40]
, Pontil, M
Maurer, A. , Pontil, M. and Romera-Paredes, B. (2016). The benefit of multitask representation learning. Journal of Machine Learning Research, 17 1--32
2016
-
[41]
Minsker, S. (2015). Geometric median and robust estimation in banach spaces. Bernoulli, 21 2308--2335
2015
-
[42]
, Balduzzi, D
Muandet, K. , Balduzzi, D. and Sch "o lkopf, B. (2013). Domain generalization via invariant feature representation. In Proceedings of the 30th International Conference on Machine Learning, vol. 28 of Proceedings of Machine Learning Research. PMLR, Atlanta, GA, USA, 10--18
2013
-
[43]
, Niranjan, U
Netrapalli, P. , Niranjan, U. N. , Sanghavi, S. , Anandkumar, A. and Jain, P. (2014). Non-convex robust PCA . In Advances in Neural Information Processing Systems. 1107--1115
2014
-
[44]
Niu, X. , Su, L. , Xu, J. and Yang, P. (2024). Collaborative learning with shared linear representations: Statistical rates and optimal algorithms. arXiv preprint arXiv:2409.04919
2024
-
[45]
Pan, S. J. and Yang, Q. (2010). A survey on transfer learning. IEEE Transactions on Knowledge and Data Engineering, 22 1345--1359
2010
-
[46]
, Gross, S
Paszke, A. , Gross, S. , Massa, F. , Lerer, A. , Bradbury, J. , Chanan, G. , Killeen, T. , Lin, Z. , Gimelshein, N. , Antiga, L. et al. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32
2019
-
[47]
, Perotte, A
Ranganath, R. , Perotte, A. , Elhadad, N. and Blei, D. M. (2016). Deep survival analysis. In Proceedings of the 1st Machine Learning for Healthcare Conference, vol. 56 of Proceedings of Machine Learning Research. PMLR, Northeastern University, Boston, MA, USA, 101--114
2016
-
[48]
Rosenstein, M. T. , Marx, Z. , Kaelbling, L. P. and Dietterich, T. G. (2005). To transfer or not to transfer. In NIPS 2005 Workshop on Inductive Transfer: 10 Years Later
2005
-
[49]
Ruder, S. (2017). An overview of multi-task learning in deep neural networks. arXiv preprint arXiv:1706.05098. ://arxiv.org/abs/1706.05098
2017 arXiv
-
[50]
, Gupta, A
Saunshi, N. , Gupta, A. and Hu, W. (2021). A representation learning perspective on the importance of train-validation splitting in meta-learning. In Proceedings of the 38th International Conference on Machine Learning, vol. 139 of Proceedings of Machine Learning Research
2021
-
[51]
, Herbrich, R
Sch \"o lkopf, B. , Herbrich, R. and Smola, A. J. (2001). A generalized representer theorem. In International conference on computational learning theory. Springer, 416--426
2001
-
[52]
and Koltun, V
Sener, O. and Koltun, V. (2018). Multi-task learning as multi-objective optimization. In Advances in Neural Information Processing Systems. 525--536
2018
-
[53]
, Chiang, C.-K
Smith, V. , Chiang, C.-K. , Sanjabi, M. and Talwalkar, A. S. (2017). Federated multi-task learning. In Advances in Neural Information Processing Systems. 4424--4434. NIPS 2017, Long Beach, CA, USA
2017
-
[54]
, Zamir, A
Standley, T. , Zamir, A. R. , Chen, D. , Guibas, L. , Malik, J. and Savarese, S. (2020). Which tasks should be learned together in multi-task learning? In Proceedings of the 37th International Conference on Machine Learning, vol. 119 of Proceedings of Machine Learning Research...
2020
-
[55]
Thekumparampil, K. K. , Jain, P. , Netrapalli, P. and Oh, S. (2021). Statistically and computationally efficient linear meta-representation learning. Advances in Neural Information Processing Systems, 34 18487--18500
2021
-
[56]
Tian, Y. , Gu, Y. and Feng, Y. (2023). Learning from similar linear representations: Adaptivity, minimaxity, and robustness. arXiv preprint arXiv:2303.17765
2023 arXiv
-
[57]
, Jin, C
Tripuraneni, N. , Jin, C. and Jordan, M. (2021). Provable meta-learning of linear representations. In International Conference on Machine Learning. PMLR, 10434--10443
2021
-
[58]
, Jordan, M
Tripuraneni, N. , Jordan, M. and Jin, C. (2020). On the theory of transfer learning: The importance of task diversity. Advances in neural information processing systems, 33 7852--7862
2020
-
[59]
Vanschoren, J. (2018). Meta-learning: A survey. arXiv preprint arXiv:1810.03548
2018 arXiv
-
[60]
Vershynin, R. (2010). Introduction to the non-asymptotic analysis of random matrices. arXiv preprint arXiv:1011.3027
2010 arXiv
-
[61]
Wainwright, M. J. (2019). High-dimensional statistics: A non-asymptotic viewpoint, vol. 48. Cambridge university press
2019
-
[62]
Wedin, P.- . (1972). Perturbation bounds in connection with singular value decomposition. BIT Numerical Mathematics, 12 99--111
1972
-
[63]
, Khoshgoftaar, T
Weiss, K. , Khoshgoftaar, T. M. and Wang, D. (2016). A survey of transfer learning. Journal of Big data, 3 1--40
2016
-
[64]
, Caramanis, C
Xu, H. , Caramanis, C. and Sanghavi, S. (2012). Robust pca via outlier pursuit. IEEE Transactions on Information Theory, 58 3047--3064
2012
-
[65]
, Liao, X
Xue, Y. , Liao, X. , Carin, L. and Krishnapuram, B. (2007). Multi-task learning for classification with dirichlet process priors. Journal of Machine Learning Research, 8 35--63
2007
-
[66]
, Chen, Y
Yin, D. , Chen, Y. , Kannan, R. and Bartlett, P. (2018). Byzantine-robust distributed learning: Towards optimal statistical rates. In Proceedings of the 35th International Conference on Machine Learning, vol. 80 of Proceedings of Machine Learning Research. PMLR, 5650--5659
2018
-
[67]
and Yang, Q
Zhang, Y. and Yang, Q. (2017). A survey on multi-task learning. arXiv preprint arXiv:1707.08114. Submitted July 25, 2017, ://arxiv.org/abs/1707.08114
2017 arXiv
-
[68]
, Chen, J
Zhou, J. , Chen, J. and Ye, J. (2011). Clustered multi-task learning via alternating structure optimization. In Advances in Neural Information Processing Systems. 702--710
2011
-
[69]
and Troyanskaya, O
Zhou, J. and Troyanskaya, O. G. (2015). Predicting effects of noncoding variants with deep learning--based sequence model. Nature Methods, 12 931--934
2015
-
[70]
Zhuang, F. , Qi, Z. , Duan, K. , Xi, D. , Zhu, Y. , Zhu, H. , Xiong, H. and He, Q. (2020). A comprehensive survey on transfer learning. Proceedings of the IEEE, 109 43--76
2020
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.