REVIEW 2 major objections 1 minor 106 references
Subsampling for supervised learning in reproducing kernel Hilbert spaces
T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Studying asymptotics of the Horvitz-Thompson reweighted empirical risk minimizer in an RKHS reveals an optimal subsampling scheme based on the trace of the covariance operator that works via plug-in.
desk verdict The paper derives a trace-based optimal subsampling rule for Horvitz-Thompson reweighted kernel ERM from asymptotics and backs it with numerics, but the step showing the trace dominates the asymptotic variance needs explicit bounds to hold up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Horvitz-Thompson reweighted empirical risk minimizer in a reproducing kernel Hilbert space, whose asymptotic covariance supplies the optimality criterion via the trace of the covariance operator.
What would settle it
If the plug-in estimator of the trace-based subsampling probabilities fails to produce lower estimation error than uniform subsampling on large-scale data sets, the claimed optimality would be refuted.
Extended reading notes
Core claim
By studying the asymptotic properties of the Horvitz-Thompson reweighted empirical risk minimizer in an RKHS, the authors reveal an optimal subsampling scheme regarding the trace of the covariance operator and show that it can be used via plug-in.
Load-bearing premise
The asymptotic covariance of the estimator is dominated by the trace of the covariance operator in a way that directly yields an optimal subsampling distribution.
Editorial extensions
If this is right
- Subsampling probabilities are chosen proportional to the trace terms of the covariance operator for asymptotic optimality.
- The optimal scheme is realized by a plug-in estimator that replaces unknown quantities with empirical estimates.
- The procedure lowers computational cost for a fixed sample size while preserving the asymptotic efficiency of the full-data estimator.
- Numerical comparisons on synthetic and real-world data confirm practical gains over standard subsampling.
Reading between the lines
- The same trace-based criterion could be tested in other nonparametric settings that admit a covariance operator.
- Implementation would allow direct comparison of variance reduction against uniform or leverage-score sampling on benchmark data sets.
- The method supplies a concrete route to lower energy use in kernel-based training pipelines by reducing the number of kernel evaluations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript analyzes subsampling for supervised learning when the hypothesis class is a reproducing kernel Hilbert space. It considers the Horvitz-Thompson reweighted empirical risk minimizer, derives its asymptotic properties, and uses those properties to identify an optimal subsampling distribution that depends on the trace of the covariance operator; the distribution is then estimated by plug-in. A numerical study on synthetic and real data is included to illustrate practicality.
Significance. If the asymptotic expansion is valid and the trace term indeed governs the leading variance uniformly across sampling schemes, the work supplies a statistically grounded, computationally attractive subsampling rule for kernel methods. The plug-in implementation and the explicit link between asymptotics and sampling probabilities would be a useful contribution to large-scale nonparametric learning.
major comments (2)
- [Abstract] Abstract (and the paragraph on asymptotic properties): the optimality claim rests on the assertion that the trace of the covariance operator dominates the asymptotic covariance of the reweighted estimator. No explicit expansion, bound, or set of conditions (e.g., on the kernel, the sampling probabilities, or the regression function) is supplied in the abstract showing that this domination holds uniformly; without that step the optimality criterion does not necessarily follow from the asymptotics.
- [asymptotic analysis section] The transition from the asymptotic covariance expression to the trace-based objective appears to treat the trace as an external functional that can be minimized directly with respect to the inclusion probabilities. If the covariance expansion contains additional terms that also depend on the sampling scheme, the claimed optimality may be incomplete; the manuscript should state the precise domination argument and any uniformity conditions.
minor comments (1)
- Notation for the Horvitz-Thompson weights and the RKHS inner product should be introduced once and used consistently; several symbols appear without prior definition in the abstract.
Simulated Author's Rebuttal
We thank the referee for the thoughtful and constructive report. The comments highlight the need for greater explicitness regarding the domination argument in both the abstract and the asymptotic analysis. We address each point below and will revise the manuscript accordingly to strengthen the presentation.
read point-by-point responses
-
Referee: [Abstract] Abstract (and the paragraph on asymptotic properties): the optimality claim rests on the assertion that the trace of the covariance operator dominates the asymptotic covariance of the reweighted estimator. No explicit expansion, bound, or set of conditions (e.g., on the kernel, the sampling probabilities, or the regression function) is supplied in the abstract showing that this domination holds uniformly; without that step the optimality criterion does not necessarily follow from the asymptotics.
Authors: We agree that the abstract, being a concise summary, does not reproduce the full technical conditions. The asymptotic expansion and the uniform domination of the trace term are established in Theorem 2 under Assumptions 1--3 (bounded kernel, moment conditions on the regression function, and sampling probabilities bounded away from zero and one). These ensure the remainder is o(1) uniformly over admissible sampling schemes. To address the concern, we will revise the abstract to include a brief reference to these conditions and the resulting optimality criterion. revision: yes
-
Referee: [asymptotic analysis section] The transition from the asymptotic covariance expression to the trace-based objective appears to treat the trace as an external functional that can be minimized directly with respect to the inclusion probabilities. If the covariance expansion contains additional terms that also depend on the sampling scheme, the claimed optimality may be incomplete; the manuscript should state the precise domination argument and any uniformity conditions.
Authors: The asymptotic covariance derived in Theorem 1 takes the form trace(Σ_p) + R_n, where the remainder R_n is shown to be o_p(1) uniformly when the inclusion probabilities satisfy inf p_i ≥ c > 0 and the kernel is continuous and bounded. Because the leading term depends on the sampling distribution only through the trace functional, minimization with respect to the inclusion probabilities is valid. We will insert an explicit remark immediately after Theorem 1 stating the domination argument, the uniformity conditions, and why no other sampling-dependent terms enter at the leading order. revision: yes
Circularity Check
No circularity: optimality criterion derived from asymptotics as external functional
full rationale
The paper studies the asymptotic properties of the Horvitz-Thompson reweighted empirical risk minimizer in an RKHS and derives an optimal subsampling scheme based on the trace of the covariance operator, which is then implemented via plug-in estimation. This trace is an external functional of the operator, not defined in terms of the subsampling probabilities or fitted to the target quantity by construction. No self-definitional steps, fitted inputs renamed as predictions, or load-bearing self-citations appear in the provided abstract or reader's analysis. The derivation chain remains self-contained against external benchmarks, with the optimality following from the asymptotic expansion rather than tautological reduction. Minor self-citation (if any) is not load-bearing for the central claim.
Assumptions & free parameters
assumptions (2)
- domain assumption The hypothesis set lies in a reproducing kernel Hilbert space
- domain assumption The estimator is a minimizer of an empirical risk reweighted à la Horvitz-Thompson
Cite this review
Pith. "Pith review of Subsampling for supervised learning in reproducing kernel Hilbert spaces." pith.science (2026). https://pith.science/paper/4X5GBSU5
@misc{pith2026260621260,
author = {Pith},
title = {Pith review of: Subsampling for supervised learning in reproducing kernel Hilbert spaces},
year = {2026},
howpublished = {\url{https://pith.science/paper/4X5GBSU5}},
note = {Machine review of arXiv:2606.21260}
}
read the original abstract
In the era of big data, subsampling became a common practice in statistical learning. By selecting a subgroup of individuals based on which the learner is trained, subsampling aims at reducing the computational cost and time of the estimation step, and ideally leads to a decrease of its energy consumption and carbon footprint. This work focuses on a nonparametric setting, in which the hypotheses set lies in a reproducing kernel Hilbert space, and the estimator is a minimizer of an empirical risk reweighted \`a la Horvitz-Thompson. By studying the asymptotic properties of this estimator, we reveal an optimal subsampling scheme (regarding the trace of the covariance operator) and show that it can be used via plug-in. A numerical study on synthetic and real-world datasets shows the practicability and the benefit of the proposed approach.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
El Ahmad, Tamim and. Sketch. Proceedings of
-
[2]
Statistica Sinica , volume =
Optimal Subsampling Algorithms for Big Data Regressions , author =. Statistica Sinica , volume =
-
[3]
Alaoui, Ahmed and Mahoney, Michael , year = 2015, volume =. Fast. Advances in
2015
-
[4]
and Mendelson, Shahar , year = 2002, journal =
Bartlett, Peter L. and Mendelson, Shahar , year = 2002, journal =. Rademacher and
2002
-
[5]
Reproducing
Berlinet, Alain and. Reproducing
-
[6]
and Guyon, I.M
Boser, B.E. and Guyon, I.M. and Vapnik, V.N. , year = 1992, abstract =. A. Conference on
1992
-
[7]
Brault, Romain and Heinonen, Markus and. Random. Proceedings of
-
[8]
Automated
Campbell, Trevor and Broderick, Tamara , year = 2019, journal =. Automated
2019
Show all 106 references
-
[9]
Bayesian
Campbell, Trevor and Broderick, Tamara , year = 2018, pages =. Bayesian. Proceedings of the 35th
2018
-
[10]
Efficient
Chatalic, Antoine and Schreuder, Nicolas and Vito, Ernesto De and Rosasco, Lorenzo , year = 2025, journal =. Efficient
2025
-
[11]
Chatalic, Antoine and Carratino, Luigi and Vito, Ernesto De and Rosasco, Lorenzo , year = 2022, pages =. Mean. Proceedings of
2022
-
[12]
Chatalic, Antoine and Schreuder, Nicolas and Rosasco, Lorenzo and Rudi, Alessandro , year = 2022, pages =. Nyström. Proceedings of the 39th
2022
-
[13]
Journal of Statistical Planning and Inference , volume =
Information-Based Optimal Subdata Selection for Big Data Logistic Regression , author =. Journal of Statistical Planning and Inference , volume =
-
[14]
and Bertail, P
Clémençon, S. and Bertail, P. and Papa, G. , year = 2016, abstract =. Learning from. Proceedings of
2016
-
[15]
and Bertail, P
Clémençon, S. and Bertail, P. and Chautru, E. and Papa, G. , year = 2019, journal =. Optimal
2019
-
[16]
and Bellet, A
Clémençon, S. and Bellet, A. and Jelassi, O. and Papa, G. , year = 2015, journal =. Scalability of
2015
-
[17]
and Bertail, P
Clemencon, S. and Bertail, P. and Chautru, E. , year = 2014, abstract =. Scaling up. Proceedings of the
2014
-
[18]
Drineas, P. and. Fast. Journal of Machine Learning Research , volume =
-
[19]
Numerische Mathematik , volume =
Faster Least Squares Approximation , author =. Numerische Mathematik , volume =
-
[20]
, year = 2005, journal =
Drineas, Petros and Mahoney, Michael W. , year = 2005, journal =. On the
2005
-
[21]
Fercoq, Olivier and Bianchi, Pascal , year = 2019, journal =. A
2019
-
[22]
Local Case-Control Sampling:
Fithian, William and Hastie, Trevor , year = 2014, journal =. Local Case-Control Sampling:
2014
-
[23]
and Blanchard, G
Gribonval, R. and Blanchard, G. and Keriven, N. and Traonmilin, Y. , year = 2017, journal =. Compressive
2017
-
[24]
Journal of Multivariate Analysis , volume =
Asymptotic Normality of Support Vector Machine Variants and Other Regularized Kernel Methods , author =. Journal of Multivariate Analysis , volume =
-
[25]
and Heyde, C
Hall, P. and Heyde, C. C. , year = 1980, publisher =. Martingale
1980
-
[26]
and Wellner, J.A
Han, Q. and Wellner, J.A. , year = 2021, journal =. Complex Sampling Designs:
2021
-
[27]
Leverage Classifier:
Han, Yixin and Yu, Jun and Zhang, Nan and Meng, Cheng and Ma, Ping and Zhong, Wenxuan and Zou, Changliang , year = 2025, journal =. Leverage Classifier:
2025
-
[28]
and Tibshirani, R
Hastie, T. and Tibshirani, R. and Friedman, J. , year = 2013, publisher =. The
2013
-
[29]
Theoretical
Hsing, Tallen and Eubank, Randall , year = 2015, publisher =. Theoretical
2015
-
[30]
Hu, Guanyu and Wang, Haiying , year = 2021, journal =. Most
2021
-
[31]
and Campbell, T
Huggins, J. and Campbell, T. and Broderick, T. , year = 2016, abstract =. Coresets for. Neural
2016
-
[32]
Optimal Subsampling Designs , author =
-
[33]
Johnson, Tyler and Guestrin, Carlos , year = 2015, pages =. Blitz:. Proceedings of the 32nd
2015
-
[34]
and Liberty, E
Karnin, Z. and Liberty, E. , year = 2019, pages =. Discrepancy,. Proceedings of the
2019
-
[35]
Gaussian
Kpotufe, Samory and Sriperumbudur, Bharath , year = 2020, pages =. Gaussian. Proceedings of the
2020
-
[36]
and Wang, Y
Lacotte, J. and Wang, Y. and Pilanci, M. , year = 2021, abstract =. Adaptive. Proceedings of the 38th
2021
-
[37]
Adaptive and
Lacotte, Jonathan and Pilanci, Mert , year = 2022, journal =. Adaptive and
2022
-
[38]
and Wang, HaiYing , year = 2024, journal =
Lee, JooChul and Schifano, Elizabeth D. and Wang, HaiYing , year = 2024, journal =. Fast
2024
-
[39]
Towards a
Li, Zhu and Ton, Jean-Francois and Oglic, Dino and Sejdinovic, Dino , year = 2021, journal =. Towards a
2021
-
[40]
and Mahoney, M.W
Ma, P. and Mahoney, M.W. and Yu, B. , year = 2015, journal =. A
2015
-
[41]
, year = 2011, journal =
Mahoney, M.W. , year = 2011, journal =. Randomized
2011
-
[42]
and Rao, A.B
Mai, T. and Rao, A.B. and Musco, C. , year = 2021, abstract =. Coresets for. Neural
2021
-
[43]
Meanti, Giacomo and Carratino, Luigi and Rosasco, Lorenzo and Rudi, Alessandro , year = 2020, volume =. Kernel. Advances in
2020
-
[44]
and Schwiegelshohn, C
Munteanu, A. and Schwiegelshohn, C. and Sohler, C. and Woodruff, D.P. , year = 2018, abstract =. On Coresets for Logistic Regression , booktitle =
2018
-
[45]
SIAM Journal on Optimization , volume =
Newton Sketch: A near Linear-Time Optimization Algorithm with Linear-Quadratic Convergence , author =. SIAM Journal on Optimization , volume =
-
[46]
, year = 2015, journal =
Pilanci, Mert and Wainwright, Martin J. , year = 2015, journal =. Randomized
2015
-
[47]
Subsampling , author =
-
[48]
Rahimi, Ali and Recht, Benjamin , year = 2007, volume =. Random. Advances in
2007
-
[49]
and Mahoney, M.W
Raskutti, G. and Mahoney, M.W. , year = 2016, journal =. A
2016
-
[50]
Generalization
Rudi, Alessandro and Rosasco, Lorenzo , year = 2017, volume =. Generalization. Advances in
2017
-
[51]
Rudi, Alessandro and Camoriano, Raffaello and Rosasco, Lorenzo , year = 2015, volume =. Less Is. Advances in
2015
-
[52]
and Pruhs, K
Samadian, A. and Pruhs, K. and Moseley, B. and Im, S. and Curtin, R. , year = 2020, abstract =. Unconditional. Proceedings of the
2020
-
[53]
Estimation in
Sancetta, Alessio , year = 2021, journal =. Estimation in
2021
-
[54]
Mathematical Programming , volume =
Pegasos: Primal Estimated Sub-Gradient Solver for. Mathematical Programming , volume =
-
[55]
Journal of Machine Learning Research , volume =
Stochastic. Journal of Machine Learning Research , volume =
-
[56]
Shao, Jun and Tu, Dongsheng , year = 1995, series =. The
1995
-
[57]
Journal of Machine Learning Research , volume =
Optimal Subsampling for High-Dimensional Partially Linear Models via Machine Learning Methods , author =. Journal of Machine Learning Research , volume =
-
[58]
, year = 2022, journal =
Sterge, Nicholas and Sriperumbudur, Bharath K. , year = 2022, journal =. Statistical
2022
-
[59]
and Joachims, T
Swaminathan, A. and Joachims, T. , year = 2015, file =. Counterfactual. Proceedings of
2015
-
[60]
Ting, Daniel and Brochu, Eric , year = 2018, volume =. Optimal. Advances in
2018
-
[61]
and Kwok, James T
Tsang, Ivor W. and Kwok, James T. and Cheung, Pak-Ming , year = 2005, journal =. Core
2005
-
[62]
A Comparative Study on Sampling with Replacement vs
Wang, HaiYing and Zou, Jiahui , year = 2021, pages =. A Comparative Study on Sampling with Replacement vs. Proceedings of
2021
-
[63]
Information-
Wang, HaiYing and Yang, Min and Stufken, John , year = 2019, journal =. Information-
2019
-
[64]
Journal of Machine Learning Research , volume =
Maximum Sampled Conditional Likelihood for Informative Subsampling , author =. Journal of Machine Learning Research , volume =
-
[65]
Wang, HaiYing , year = 2019, journal =. More
2019
-
[66]
Wang, HaiYing and Zhu, Rong and Ma, Ping , year = 2018, journal =. Optimal
2018
-
[67]
Biometrika , volume =
Optimal Subsampling for Quantile Regression in Big Data , author =. Biometrika , volume =
-
[68]
Sampling
Wang, Jing and Zou, Jiahui and Wang, HaiYing , year = 2022, journal =. Sampling
2022
-
[69]
Scale-Invariant
Wang, Jing and Wang, HaiYing and Zhang, Hao , year = 2024, journal =. Scale-Invariant
2024
-
[70]
and Gittens, A
Wang, S. and Gittens, A. and Mahoney, M.W. , year = 2018, journal =. Sketched
2018
-
[71]
Using the
Williams, Christopher and Seeger, Matthias , year = 2000, volume =. Using the. Advances in
2000
-
[72]
, year = 2014, journal =
Woodruff, David P. , year = 2014, journal =. Sketching as a
2014
-
[73]
Statistics & Probability Letters , volume =
Some Results on the Convergence of Conditional Distributions , author =. Statistics & Probability Letters , volume =
-
[74]
Yang, Tianbao and Li, Yu-feng and Mahdavi, Mehrdad and Jin, Rong and Zhou, Zhi-Hua , year = 2012, volume =. Nyström. Advances in
2012
-
[75]
and Meng, X
Yang, J. and Meng, X. and Mahoney, M.W. , year = 2014, journal =. Quantile
2014
-
[76]
, year = 2017, journal =
Yang, Yun and Pilanci, Mert and Wainwright, Martin J. , year = 2017, journal =. Randomized Sketches for Kernels:
2017
-
[77]
Journal of Statistical Planning and Inference , volume =
Model Constraints Independent Optimal Subsampling Probabilities for Softmax Regression , author =. Journal of Statistical Planning and Inference , volume =
-
[78]
Statistical Papers , volume =
Optimal Subsampling for Softmax Regression , author =. Statistical Papers , volume =
-
[79]
Yao, Yaqiong and Wang, HaiYing , year = 2021, journal =. A
2021
-
[80]
Statistical Papers , volume =
Information-Based Optimal Subdata Selection for Non-Linear Models , author =. Statistical Papers , volume =
-
[81]
Yu, Jun and Wang, HaiYing and Ai, Mingyao and Zhang, Huiming , year = 2022, journal =. Optimal
2022
-
[82]
Approximating
Zhang, Haixiang and Zuo, Lulu and Wang, HaiYing and Sun, Liuquan , year = 2024, journal =. Approximating
2024
-
[83]
Computational Statistics , volume =
Optimal Subsample Selection for Massive Logistic Regression with Distributed Data , author =. Computational Statistics , volume =
-
[84]
Data Sparse Nonparametric Regression with \( \)-Insensitive Losses , booktitle =
Sangnier, Maxime and Fercoq, Olivier and. Data Sparse Nonparametric Regression with \( \)-Insensitive Losses , booktitle =
-
[85]
Probability and
Billingsley, Patrick , year = 1995, series =. Probability and
1995
-
[86]
Optimization
Bottou, L. Optimization. SIAM Review , volume =
-
[87]
Efficient
Braverman, Vladimir and Feldman, Dan and Lang, Harry and Statman, Adiel and Zhou, Samson , year = 2021, pages =. Efficient. Proceedings of
2021
-
[88]
Chen, Xiaohong and White, Halbert , year = 1998, journal =. Central. 3532640 , eprinttype =
1998
-
[89]
Accumulations of
Chen, Yifan and Yang, Yun , year = 2021, pages =. Accumulations of. Proceedings of
2021
-
[90]
Christmann, Andreas and Steinwart, Ingo , year = 2008, series =. Support
2008
-
[91]
Computational Statistics & Data Analysis , volume =
Consistency of Support Vector Machines Using Additive Kernels for Additive Models , author =. Computational Statistics & Data Analysis , volume =
-
[92]
Learning from
Clemencon, Stephan and Bertail, Patrice and Papa, Guillaume , year = 2016, pages =. Learning from. Proceedings of
2016
-
[93]
El Ahmad, Tamim and Laforgue, Pierre and. Fast. Transactions on Machine Learning Research , abstract =
-
[94]
Feller, William , year = 1991, publisher =. An
1991
-
[95]
Asymptotic
Hable, Robert , year = 2012, number =. Asymptotic. arXiv , keywords =:1203.4354 , publisher =
2012 arXiv
-
[96]
, year = 2021, journal =
Han, Qiyang and Wellner, Jon A. , year = 2021, journal =. Complex Sampling Designs:
2021
-
[97]
Leverage Classifier:
Han, Yixin and Yu, Jun and Zhang, Nan and Meng, Cheng and Ma, Ping and Zhong, Wenxuan and Zou, Changliang , year = 2023, number =. Leverage Classifier:. arXiv , keywords =:2308.12444 , primaryclass =
2023
-
[98]
and Hurwitz, William N
Hansen, Morris H. and Hurwitz, William N. , year = 1943, journal =. On the. 2235923 , eprinttype =
1943
-
[99]
Theoretical Foundations of Functional Data Analysis, with an Introduction to Linear Operators , author =
-
[100]
Coresets for
Huggins, Jonathan and Campbell, Trevor and Broderick, Tamara , year = 2016, volume =. Coresets for. Advances in
2016
-
[101]
arXiv , keywords =:2304.03019 , primaryclass =
Optimal Subsampling Designs , author =. arXiv , keywords =:2304.03019 , primaryclass =
-
[102]
Jakubowski, Adam , editor =. On. Mathematical
-
[103]
Probability Theory and Related Fields , volume =
Asymptotic Normality for Two-Stage Sampling from a Finite Population , author =. Probability Theory and Related Fields , volume =
-
[104]
and Kwok, James T
Tsang, Ivor W. and Kwok, James T. and Cheung, Pak-Ming , year = 2005, journal =. Core Vector Machines:
2005
-
[105]
Van Der Vaart, A. W. , year = 1998, edition =. Asymptotic
1998
-
[106]
Van Der Vaart, A. W. and Wellner, Jon A. , year = 2023, series =. Weak
2023
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.