{"id":"972e326d-6fa1-4b0c-b7be-f0897b3ca6c6","arxiv_id":"2606.23363","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"An optimal Poisson subsampling algorithm is developed for quantile regression with large-scale longitudinal data, deriving asymptotic properties and outperforming uniform subsampling in simulations and real data.","lead":"The paper proposes an optimal Poisson subsampling algorithm for quantile regression on large longitudinal datasets to reduce computational demands. A smart generalist might read it to learn practical ways to analyze massive time-series data without full computation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's identification of the unspecified regularity conditions as the weakest point is accurate given the information available; the present review finds no additional, more specific technical vulnerability that would alter the UNVERDICTED status.","tokens_in":1588,"tokens_out":261,"duration_ms":14545,"concrete_test":"Recompute the asymptotic variance expression for the weighted quantile GEE estimator (main theorem) after substituting the optimal subsampling probabilities; check whether the resulting variance matches the claimed efficiency gain relative to uniform subsampling and whether any extra variability from estimating the probabilities appears.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that an optimal Poisson subsampling procedure yields computationally cheaper quantile regression estimators for large longitudinal data while preserving asymptotic normality (under regularity conditions) and outperforming uniform subsampling. No internal inconsistency, hidden assumption, or unsupported step is visible from the abstract or the stated claims. The use of weighted quantile GEE together with Poisson subsampling is a standard construction; the paper asserts that an efficient algorithm exists for the optimal probabilities and that penalized versions are also covered. Without a concrete flaw in any displayed equation or condition, the argument is not shown to be load-bearingly fragile.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes an optimal Poisson subsampling algorithm for quantile regression on large-scale longitudinal data to reduce computational burden. It derives asymptotic properties of estimators obtained from weighted quantile generalized estimating equations under regularity conditions, presents an efficient algorithm for parameter estimation, extends the framework to penalized weighted smooth quantile GEE with corresponding theory, and demonstrates superior performance relative to uniform Poisson subsampling through simulations and a real-data application.","tokens_in":1679,"tokens_out":346,"duration_ms":14215,"significance":"If the asymptotic normality results and the claimed efficiency of the optimal subsampling hold, the work provides a practical and theoretically grounded approach to scalable quantile regression for longitudinal data. The combination of Poisson subsampling with weighted GEE, the extension to regularization, and the empirical comparisons against uniform subsampling constitute a coherent contribution to computational statistics for big data.","major_comments":[],"minor_comments":[{"comment":"The abstract refers to 'some regularity conditions' for the asymptotic results; these should be stated explicitly (perhaps in a dedicated assumptions section or theorem statement) so readers can assess their plausibility without searching the full text.","section":"Abstract / Introduction"},{"comment":"Notation for the optimal sampling probabilities and the weighted estimating equations should be introduced with a clear table or displayed equation early in the methods section to improve readability.","section":"Methods"},{"comment":"The simulation section would benefit from reporting the actual subsample sizes used and the resulting wall-clock times alongside the statistical performance metrics to substantiate the computational savings claim.","section":"Simulations"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive assessment of our work on optimal Poisson subsampling for quantile regression with large-scale longitudinal data and for recommending minor revision. No specific major comments were provided in the report.","responses":[],"tokens_in":1091,"tokens_out":59,"duration_ms":7934,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know about this paper is that it proposes an optimal Poisson subsampling procedure for quantile regression applied to large longitudinal datasets. This is meant to reduce the computational load compared to using the full data, while still getting estimators with good asymptotic behavior.\n\nWhat stands out as new is the focus on finding the best sampling probabilities under the Poisson subsampling scheme specifically for the weighted quantile generalized estimating equations that account for the longitudinal dependence. They provide an efficient algorithm for parameter estimation and extend the theory to include penalized versions for regularized estimation. The numerical simulations and real data example are used to illustrate that this optimal approach performs better than the simpler uniform Poisson subsampling.\n\nThe paper does a reasonable job of addressing a practical problem in handling big data for this type of analysis. The claims about alleviating computational burden and the outperformance are backed by the experiments described.\n\nOn the soft spots, the asymptotic properties are derived under some regularity conditions, which is standard but means the results hold only when those are met; the abstract doesn't spell out what they are, so checking the paper for how reasonable they are would be good. There's also the question of how much time is spent calculating the optimal probabilities versus the time saved in the regression fit itself. If that calculation is costly, the net benefit could be smaller than claimed, though the paper likely includes some discussion of this.\n\nThis work is aimed at computational statisticians and data analysts who routinely deal with massive longitudinal data and need efficient ways to run quantile regressions. Someone looking for methods to scale statistical procedures to big data would find the specific algorithm and comparisons useful.\n\nI would recommend sending it for peer review. It has a clear methodological contribution with theory and practical validation, so referees can assess the details and suggest improvements where needed.","headline":"This paper works out optimal Poisson subsampling probabilities for quantile regression on large longitudinal data and shows it beats uniform subsampling in simulations and a real example.","tokens_in":2137,"tokens_out":436,"would_cite":false,"duration_ms":24354,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Optimal Poisson subsampling reduces computational burden for quantile regression on large longitudinal data while retaining asymptotic properties.","keywords":["quantile regression","Poisson subsampling","longitudinal data","generalized estimating equations","optimal subsampling","penalized estimation","computational efficiency","asymptotic properties"],"falsifier":"A simulation study on a dataset of known size in which the mean squared error or coverage probability of the optimal-subsampling estimator is materially worse than that of the full-data estimator or of uniform subsampling.","tokens_in":2488,"feed_emoji":"📊","tokens_out":629,"duration_ms":19420,"temperature":0.7,"pith_summary":"The paper develops an optimal Poisson subsampling algorithm for quantile regression applied to large-scale longitudinal data. Sampling probabilities are chosen to minimize the asymptotic variance of the estimator obtained from weighted quantile generalized estimating equations, rather than using equal probabilities. Asymptotic consistency and normality of the resulting estimators are derived under regularity conditions, and an efficient algorithm is supplied for both unpenalized and penalized estimation. Numerical simulations and a real-data example show that the optimal procedure requires far less computation than full-data analysis and yields lower variance than uniform Poisson subsampling. The same framework supports regularized estimation through penalized weighted smooth quantile generalized estimating equations.","feed_headline":"Optimal subsampling speeds quantile regression on large longitudinal data","feed_subtitle":"Probabilities are chosen to minimize estimator variance, cutting computation while beating uniform sampling in simulations and real data.","key_machinery":"Optimal Poisson subsampling probabilities chosen to minimize the asymptotic variance of the weighted quantile GEE estimator.","core_discovery":"By deriving inclusion probabilities that optimize the asymptotic variance of the weighted quantile generalized estimating equation estimator, the optimal Poisson subsampling procedure produces consistent and asymptotically normal quantile regression estimates from a small fraction of the original longitudinal observations, with the same regularity conditions that validate the full-data estimator.","pith_inferences":["The probability-calculation step itself could be parallelized or approximated by pilot subsamples to further reduce overhead on extremely large data.","The optimality criterion might be adapted to other estimating-equation losses such as those arising in mean regression or generalized linear models for longitudinal data.","If the initial parameter estimates used to form the sampling probabilities are themselves obtained from a crude pilot sample, the overall procedure remains consistent provided the pilot fraction grows appropriately."],"forward_implications":["Quantile regression becomes feasible on longitudinal datasets whose size would otherwise exceed available memory or time limits.","Penalized versions of the procedure allow simultaneous variable selection and estimation without full-data computation.","The same subsampling weights can be reused across multiple quantile levels once they are computed.","Asymptotic normality supplies valid standard errors and confidence intervals after subsampling."],"fun_headline_variants":["Optimal Poisson subsampling optimizes quantile regression variance","Consistent estimates from optimal Poisson subsampling of longitudinal data","Optimal inclusion probabilities minimize estimator variance in quantile regression","Asymptotically normal quantile regression via optimal Poisson subsampling"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The data satisfy the regularity conditions under which the asymptotic properties of the weighted quantile generalized estimating equations are valid.","fun_headline_variants_meta":{"raw":{"variants":["Optimal Poisson subsampling optimizes quantile regression variance","Consistent estimates from optimal Poisson subsampling of longitudinal data","Optimal inclusion probabilities minimize estimator variance in quantile regression","Asymptotically normal quantile regression via optimal Poisson subsampling"]},"model":"grok-4.3","cost_usd":0.006935,"raw_usage":{"total_tokens":3147,"prompt_tokens":531,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":69349500,"prompt_tokens_details":{"text_tokens":531,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2557,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":531,"tokens_out":59,"duration_ms":17924,"temperature":1.0,"reasoning_tokens":2557,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T05:55:29.960174+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A simulation study on a dataset of known size in which the mean squared error or coverage probability of the optimal-subsampling estimator is materially worse than that of the full-data estimator or of uniform subsampling.","supporting_citations":[],"review_version":1}