REVIEW 3 minor
Optimal Poisson subsampling for quantile regression with large-scale longitudinal data
T0 review · 0 major / 3 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Optimal Poisson subsampling reduces computational burden for quantile regression on large longitudinal data while retaining asymptotic properties.
desk verdict This paper works out optimal Poisson subsampling probabilities for quantile regression on large longitudinal data and shows it beats uniform subsampling in simulations and a real example. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Optimal Poisson subsampling probabilities chosen to minimize the asymptotic variance of the weighted quantile GEE estimator.
What would settle it
A simulation study on a dataset of known size in which the mean squared error or coverage probability of the optimal-subsampling estimator is materially worse than that of the full-data estimator or of uniform subsampling.
Extended reading notes
Core claim
By deriving inclusion probabilities that optimize the asymptotic variance of the weighted quantile generalized estimating equation estimator, the optimal Poisson subsampling procedure produces consistent and asymptotically normal quantile regression estimates from a small fraction of the original longitudinal observations, with the same regularity conditions that validate the full-data estimator.
Load-bearing premise
The data satisfy the regularity conditions under which the asymptotic properties of the weighted quantile generalized estimating equations are valid.
Editorial extensions
If this is right
- Quantile regression becomes feasible on longitudinal datasets whose size would otherwise exceed available memory or time limits.
- Penalized versions of the procedure allow simultaneous variable selection and estimation without full-data computation.
- The same subsampling weights can be reused across multiple quantile levels once they are computed.
- Asymptotic normality supplies valid standard errors and confidence intervals after subsampling.
Reading between the lines
- The probability-calculation step itself could be parallelized or approximated by pilot subsamples to further reduce overhead on extremely large data.
- The optimality criterion might be adapted to other estimating-equation losses such as those arising in mean regression or generalized linear models for longitudinal data.
- If the initial parameter estimates used to form the sampling probabilities are themselves obtained from a crude pilot sample, the overall procedure remains consistent provided the pilot fraction grows appropriately.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an optimal Poisson subsampling algorithm for quantile regression on large-scale longitudinal data to reduce computational burden. It derives asymptotic properties of estimators obtained from weighted quantile generalized estimating equations under regularity conditions, presents an efficient algorithm for parameter estimation, extends the framework to penalized weighted smooth quantile GEE with corresponding theory, and demonstrates superior performance relative to uniform Poisson subsampling through simulations and a real-data application.
Significance. If the asymptotic normality results and the claimed efficiency of the optimal subsampling hold, the work provides a practical and theoretically grounded approach to scalable quantile regression for longitudinal data. The combination of Poisson subsampling with weighted GEE, the extension to regularization, and the empirical comparisons against uniform subsampling constitute a coherent contribution to computational statistics for big data.
minor comments (3)
- [Abstract / Introduction] The abstract refers to 'some regularity conditions' for the asymptotic results; these should be stated explicitly (perhaps in a dedicated assumptions section or theorem statement) so readers can assess their plausibility without searching the full text.
- [Methods] Notation for the optimal sampling probabilities and the weighted estimating equations should be introduced with a clear table or displayed equation early in the methods section to improve readability.
- [Simulations] The simulation section would benefit from reporting the actual subsample sizes used and the resulting wall-clock times alongside the statistical performance metrics to substantiate the computational savings claim.
Simulated Author's Rebuttal
We thank the referee for their positive assessment of our work on optimal Poisson subsampling for quantile regression with large-scale longitudinal data and for recommending minor revision. No specific major comments were provided in the report.
Circularity Check
No significant circularity; derivation self-contained under external regularity conditions
full rationale
The paper proposes an optimal Poisson subsampling procedure for quantile regression on longitudinal data, derives asymptotic normality for the resulting weighted quantile GEE estimators, and provides an efficient algorithm plus a penalized extension. All claims are conditioned on standard regularity conditions external to the derivation itself. No equation reduces a claimed prediction to a fitted input by construction, no self-citation is invoked as a uniqueness theorem or load-bearing premise, and no ansatz is smuggled via prior work. The central performance claims rest on simulation and real-data comparison against uniform subsampling rather than on any self-referential identity. This is the normal, non-circular case for a methodological statistics paper.
Assumptions & free parameters
assumptions (1)
- domain assumption Some regularity conditions
Cite this review
Pith. "Pith review of Optimal Poisson subsampling for quantile regression with large-scale longitudinal data." pith.science (2026). https://pith.science/paper/4GOKJT5R
@misc{pith2026260623363,
author = {Pith},
title = {Pith review of: Optimal Poisson subsampling for quantile regression with large-scale longitudinal data},
year = {2026},
howpublished = {\url{https://pith.science/paper/4GOKJT5R}},
note = {Machine review of arXiv:2606.23363}
}
read the original abstract
To address the computational challenges arising from large-scale longitudinal data, an optimal Poisson subsampling algorithm is proposed for quantile regression. The proposed method can substantially alleviate computational burden. Under some regularity conditions, we derive the asymptotic properties of the estimators from weighted quantile generalized estimating equations. For practical implementation, an efficient algorithm is proposed for parameter estimation. Furthermore, asymptotic theory is established for penalized weighted smooth quantile generalized estimating equations, and regularized parameter estimation is performed within the optimal Poisson subsampling framework. Both numerical simulations and a real data application demonstrate that the proposed optimal Poisson subsampling algorithm outperforms the uniform Poisson subsampling algorithm, and the regularized estimation exhibits satisfactory performance as well.
Figures
Figures from the paper (3 more)
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.