{"id":"15ecf8c8-019d-4463-8ff9-9701e84456b4","arxiv_id":"2412.15049","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A parametric linear regression model for quantile functions, applied to CT lung-density data, estimates treatment response with explicit confidence intervals.","lead":"Researchers built a new statistical method that models how entire lung-density distributions change after asthma treatment, and applied it to CT scans of 44 patients. The method gives doctors a way to measure, with confidence intervals, whether a patient's lung condition improved.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The d=1 Gaussian quantile assumption is unvalidated; all inference and clinical conclusions apply to projected Gaussian summaries, not the actual CT quantile functions.","rationale":"The reader's weakest assumption correctly identifies the unvalidated Gaussian restriction (d=1) as the main vulnerability. I agree: this assumption is structural, because the entire inference framework—ML estimators, confidence intervals, residual densities, and the clinical interpretation—operates on the Gaussian projection parameters μ and σ, not on the full empirical quantile functions. If the Gaussian fits are poor, then β2 and the derived threshold do not measure changes in the true HU distributions. A secondary issue is that μ_qx,i and σ_qx,i are treated as deterministic covariates despite being estimated, but with millions of voxels per scan the sampling error is likely negligible, so I do not treat this as the primary concern. The concern does not change the verdict: the mathematical development appears internally coherent, and the paper explicitly frames d=1 as a proof-of-concept, so CONDITIONAL is the appropriate outcome. The proposed check—applying the paper's own BIC procedure—would settle whether the d=1 assumption is empirically defensible.","tokens_in":34410,"tokens_out":8146,"duration_ms":72401,"concrete_test":"Run the BIC-based degree-selection procedure of Section 2.2 on all 88 empirical quantile functions (44 pre-treatment and 44 post-treatment) with D=3 or D=4, and report the distribution of the selected degrees d* for each patient. If a substantial fraction of patients (e.g., more than 10%) yields d* > 1, then the d=1 Gaussian model is not an adequate representation of the data, and the regression estimates, confidence intervals, and the −746.4 HU threshold in Section 4 should be re-interpreted as conditional on the Gaussian projection rather than as properties of the actual quantile functions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central statistical claim—a parametric linear regression for quantile function data with explicit inference—is built entirely on the assumption that each observed quantile function lies in the Gaussian family Q^2_1 (Section 3.1, model (3.3) with d=1). In Section 2.2 the paper proposes a BIC procedure to choose the polynomial degree d, but Section 4 simply sets d=1 as a proof-of-concept without reporting the BIC-selected degrees or any measure of how well the Gaussian projections approximate the empirical quantile functions (e.g., the L2 projection error in Proposition 6.1). Since the response Q_Y,i and covariates q_x,i are replaced by their Gaussian projections (μ_qx,i, σ_qx,i) and since the clinical claims—β2=0.638 as reduced variability, the −746.4 HU responder threshold, and the 75% response rate—are expressed in terms of these Gaussian parameters, a poor Gaussian fit would sever the link between the regression outputs and the actual CT voxel distributions. Lung HU histograms are typically skewed and often multi-modal; the adequacy of the Gaussian approximation is therefore not obvious and is never demonstrated. This is load-bearing because the method's novelty is precisely its claimed ability to do inference on quantile-function data, not merely on a coarse two-parameter summary.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a fully parametric linear regression model for quantile function data, QY,i = β0 + β1 μ_qx,i + β2 q^c_x,i + E_i, with errors following a quantile normal-exponential distribution, and derives explicit maximum likelihood estimators, unbiased versions, pivot distributions, confidence intervals, and an approximate confidence region for a new mean response. The methodology is applied to paired pre- and post-treatment CT scans of 44 asthma patients, after projecting each patient's empirical Hounsfield-unit quantile function onto Gaussian quantile functions (d = 1). The main reported findings are β2 = 0.638 (95% CI 0.614 to 0.647), interpreted as reduced post-treatment variability, and a pre-treatment mean threshold of −746.4 HU above which 75% of severe patients show a positive response.","tokens_in":34603,"tokens_out":7965,"duration_ms":70450,"significance":"If the model assumptions hold, the paper makes a useful theoretical contribution: it provides a rare quantile-function regression model with explicit finite-sample distributions for all estimators and closed-form confidence intervals, and the derivations in Sections 6 and 7 are detailed and appear internally consistent. The paper also ships reproducible R code and data. The countervailing issue is that all inference is performed on Gaussian (d = 1) projections of the quantile functions, and the application does not validate this projection or the exponential error assumption; there is also no simulation evidence for the coverage of the approximate confidence regions. Thus the practical significance for the CT application is not yet established, although the theoretical framework is promising and worth developing further.","major_comments":[{"comment":"The paper's central inference is performed on Gaussian (d = 1) projections, but the adequacy of this projection is never demonstrated. Section 2.2 introduces a BIC procedure to select the polynomial degree d, yet Section 3.2 and Section 4 simply set d = 1 'as a proof-of-concept' and report no BIC-selected degrees, no L2 projection errors, and no visual comparison between the empirical quantile functions and the fitted Gaussian quantile functions. Since q_x,i and Q_Y,i in model (3.3) are replaced by (μ_qx,i, σ_qx,i) and (μ_qy,i, σ_qy,i), and since the clinical statements in Section 4.1 (β2 = 0.638, the −746.4 HU threshold, the 75% response rate) are expressed in those Gaussian parameters, poor Gaussian fit would sever the link between the regression outputs and the actual CT HU distributions. Please report the projection diagnostics (e.g., the projection norm available from Proposition 6.1), the BIC-selected degrees for each patient, and a sensitivity analysis showing how the estimates and threshold change when d = 2 is used or when the empirical quantile functions enter the analysis directly.","section":"Section 2.2 and Section 4"},{"comment":"The covariates (μ_qx,i, σ_qx,i) are assumed non-random, but they are estimated from segmented images via a threshold-based segmentation and then orthogonal projection onto Q^2_1. With n = 44, the reported 95% confidence interval for β2 is extremely narrow (0.614 to 0.647), and this width reflects only the randomness in Q_Y,i under the model, not uncertainty in the preprocessing. The manuscript should either justify that the segmentation and projection variability is negligible (e.g., by repeating the analysis with different segmentation thresholds and with d = 1 versus d = 2) or incorporate this variability into the inference. As written, the confidence statements in Corollary 3.5 and Section 4.1 condition on the exact Gaussian projections being the true covariates.","section":"Section 3.2, model (3.3)"},{"comment":"No simulation study is included to verify finite-sample behavior. The confidence intervals in Corollary 3.5 rest on exact sampling distributions under model (3.3), but the confidence region in Proposition 3.6 is only approximate because it plugs estimators into the density, and the residual-density p-values in Section 4 are also based on plug-in densities. With n = 44 and a non-normal pivot for β2, the actual coverage of these procedures is unknown. A simulation study under model (3.3) reporting coverage of the confidence intervals and of CR_{1−α}, together with the type I error of the test H0: β2 ≥ 1, would directly address this gap.","section":"Section 3.3 (Proposition 3.6) and Section 4"},{"comment":"The manuscript reports that the residual standard deviations fit the exponential distribution 'to a lesser extent', with mass concentrated in the center at the expense of the left tail, and proceeds because the exponential assumption enables explicit formulas. This is the distributional assumption on which all inference about β2 and β rests, so a qualitative histogram is not sufficient. Please add a formal goodness-of-fit assessment for the exponential error component (e.g., tests of exponentiality of the residual scale components or a likelihood-ratio comparison with a Gamma or Weibull error family) and show that the conclusion β2 < 1 and the reported confidence interval are robust to this choice. The p-value of 0.0508 from the Hellinger correlation test addresses only the independence of μ_E and σ_E, not the marginal exponential assumption.","section":"Section 4, residual diagnostics"}],"minor_comments":[{"comment":"The threshold −746.4 HU and the responder rates (36% and 75%) are derived from the estimated coefficients β0 and β1, but no standard error or confidence interval is provided for these derived quantities; a delta-method or bootstrap interval would help calibrate the clinical claims.","section":"Section 4.1"},{"comment":"In the console output, the statistic for β2hat is printed as −0.007, which is inconsistent with the estimate 0.638 and standard error 0.0088; please correct the formatting or explain what is being reported.","section":"Section 4, R output"},{"comment":"The statement that none of the existing Wasserstein regression techniques 'provide the capability for inference' is broad; for instance, Chen, Lin and Müller (2023) and related work provide asymptotic inference for Wasserstein regression. Please qualify the statement to 'no closed-form finite-sample inference' or cite the relevant inference results.","section":"Introduction, Section 1"},{"comment":"The interactive 3D widgets are referenced by external URLs; for archival reproducibility, please embed static versions as supplementary material or provide them in the GitHub repository.","section":"Figures 3.3 and 3.4"},{"comment":"There are several typographical errors, including 'practionners' in Section 5 and 'Morevover' in Section 4.1; a careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses an important applied problem and provides a substantial theoretical apparatus, but the lack of validation of the Gaussian projection is the main obstacle. In my view, the paper can be made publishable if the authors add projection diagnostics, a sensitivity analysis, and a simulation study; these are substantial but feasible revisions. I would not recommend acceptance in the current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: the paper delivers what it promises—a parametric linear regression model for quantile functions with explicit MLEs, confidence intervals, and residual densities. The derivations in Section 7 are detailed and appear internally correct. The QNE error model is new, and the qlm() implementation plus the CT application are real additions to the Wasserstein-space regression literature.\n\nThe soft spot is exactly where the stress-test points: the paper proposes a BIC procedure for choosing the polynomial degree d, but then simply sets d=1 and never reports projection errors or BIC values. With d=1, all inference is on Gaussian quantile projections, i.e., just the mean and SD of each patient's HU distribution. If those distributions are skewed or multi-modal—as lung HU histograms often are—the clinical interpretation of beta2 and the -746.4 threshold applies to the projected Gaussian summary, not the actual quantile functions. The paper does acknowledge this as a proof-of-concept and says d>1 is future work, but for a paper whose novelty is \"inference on quantile functions,\" the adequacy of the Gaussian approximation is load-bearing and needs at least a diagnostic.\n\nThe lack of a simulation study is a second soft spot, though not fatal. The theoretical distributions are exact under the model, but the approximate confidence region in Proposition 3.6 uses plug-in densities, and finite-sample coverage is untested. A small simulation would substantially strengthen the paper.\n\nI'm less worried about two things the reader flagged. Treating mu_qx and sigma_qx as deterministic covariates is standard in this setting, and with 3–15 million voxels per patient the estimation error is negligible. And the -746.4 threshold plus 75% response rate is post hoc in-sample exploration; the paper should label it as such, but it's not a methodological flaw.\n\nBottom line: this is a real contribution written by people who know the literature. It belongs in the review process, not the reject pile. I'd send it to a referee and ask for a Gaussian-approximation diagnostic and a small simulation before publication.","headline":"A genuine methodological contribution with closed-form inference for quantile-function regression, but the d=1 Gaussian restriction is unvalidated and the clinical claims rest on it.","tokens_in":35199,"tokens_out":3547,"would_cite":true,"duration_ms":35811,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J05","62F10","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A fully parametric linear regression model for quantile function data yields closed-form estimators, confidence intervals, and residual densities for paired CT lung scans.","keywords":["quantile function regression","Hounsfield units","asthma treatment response","air trapping","Wasserstein space","maximum likelihood","computed tomography","functional data analysis"],"falsifier":"Re-fit the analysis with the polynomial degree $d$ chosen by BIC (allowing $d > 1$) on the same 44 patients and compare the empirical quantile functions to their Gaussian projections; if the average $L^2$ approximation error is large relative to the noise variance, or if the $\\beta_2$ confidence interval shifts materially or crosses 1, the Gaussian restriction decides the central claim.","tokens_in":34164,"feed_emoji":"🫁","tokens_out":9378,"duration_ms":66579,"temperature":0.7,"pith_summary":"This paper claims that treatment response in asthma can be read from paired CT scans through a new parametric regression model for quantile functions. The model writes each post-treatment lung-density distribution as a linear combination of the pre-treatment mean and centered quantile curve plus a quantile normal-exponential error, yielding explicit maximum-likelihood estimators, confidence intervals, and residual densities. On 44 patients treated with Benralizumab, the estimated slope $\\beta_2 = 0.638$ (95% CI 0.614 to 0.647) indicates that the spread of the lung-density distribution shrinks after treatment, which the authors interpret as reduced air trapping. The authors further derive a baseline severity threshold of $-746.4$ HU: among patients with more severe initial disease ($\\mu_{q_x} \\leq -746.4$ HU), 75% show positive response.","feed_headline":"Asthma drug shrinks lung density spread, CT quantile model finds","feed_subtitle":"Parametric quantile regression yields explicit confidence intervals; beta2=0.638 means less air trapping.","key_machinery":"The load-bearing object is the space $Q_1^2$ of Gaussian quantile functions, where every quantile curve has the form $q(p) = a_0 + a_1 \\Phi^{-1}(p)$. This restriction reduces each patient's entire lung voxel distribution to two scalars (mean and standard deviation), and the regression model becomes a bivariate linear model on these scalars. The errors $E_i$ are built from independent normal and exponential coefficients, giving a quantile normal-exponential distribution and making the likelihood tractable: the maximum-likelihood estimators of the intercept and slope on the means are the usual Gaussian linear-regression estimators, while the slope on the centered quantile curve is the minimum of the patient-wise ratio of post- to pre-treatment standard deviations. Explicit pivots for $\\beta_2$ and the exponential scale $\\beta$ follow from the memoryless property of the exponential distribution, which is what makes closed-form confidence intervals possible.","core_discovery":"The central discovery is a fully parametric linear regression model for quantile function data that supports closed-form statistical inference, analogous to ordinary least squares. Under the restriction to Gaussian quantile functions, each pre-treatment quantile function is represented by its mean and standard deviation, and the post-treatment quantile function is modeled as $Q_{Y,i} = \\beta_0 + \\beta_1 \\mu_{q_{x,i}} + \\beta_2 q_{x,i}^c + E_i$, with errors following a quantile normal-exponential distribution. The paper derives explicit maximum-likelihood estimators for all five parameters, their exact sampling distributions, pivots, and confidence intervals, as well as an approximate density for residual quantile functions and a confidence region for the mean response of a new patient. In the lung application, the estimate $\\beta_2 = 0.638$ with 95% CI $[0.614, 0.647]$ is the key result: since $\\beta_2 < 1$, the post-treatment quantile curves are shrunk toward their mean, implying lower variability of Hounsfield units and, per the clinical hypothesis, reduced air trapping.","pith_inferences":["If the Gaussian quantile projection is a poor approximation to the actual CT voxel distributions, the reported confidence intervals may understate uncertainty, since the pre-treatment $(\\mu, \\sigma)$ values are estimated rather than known covariates; a simulation study that resamples the voxel-level data would quantify this.","The method could be adapted to test treatment effects at the lobe level by applying the same regression separately to segmented lobes, which the authors suggest and which would preserve spatial information.","Because the error distribution conflates measurement error with biological noise, the $\\beta_2$ estimate likely reflects attenuation bias from noisy $\\sigma_{q_x}$; a measurement-error variant of the model is a natural next step."],"forward_implications":["If the model is correct, $\\beta_2 < 1$ provides a quantitative, scan-derived biomarker for air-trapping improvement after one year of Benralizumab treatment.","Clinicians can apply the method to any paired quantile-function dataset and obtain p-values and confidence intervals with the same ease as classical linear regression.","The threshold $-746.4$ HU derived from $\\beta_0$ and $\\beta_1$ gives a baseline severity level above which a majority of patients (75% in this cohort) show measurable improvement.","The residual quantile functions provide a per-patient diagnostic tool, flagging outliers such as patients #9 and #11 for chart review.","The method sets up an extension to polynomial degrees $d > 1$, which would allow non-Gaussian quantile shapes at the cost of losing explicit formulas."],"supporting_citations":[{"why":"Provides the original least-squares quantile-function regression model that this paper makes parametric and inferential.","marker":"Verde and Irpino (2010)"},{"why":"Extends the Wasserstein least-squares approach for histogram data that the new model builds on.","marker":"Irpino and Verde (2015)"},{"why":"Supplies the Wasserstein-space statistical background that motivates treating quantile functions as data objects.","marker":"Panaretos and Zemel (2019)"},{"why":"A Wasserstein regression approach with no explicit inference, serving as the benchmark the paper contrasts.","marker":"Chen, Lin and Müller (2023)"},{"why":"The classical quantile regression monograph used to distinguish the model from scalar quantile regression.","marker":"Koenker (2005)"},{"why":"The Hellinger correlation test used to check the independence assumption between the mean and scale residuals.","marker":"Geenens and Lafaye de Micheaux (2022)"},{"why":"Supplies the threshold-based lung segmentation algorithm that isolates the lung voxels.","marker":"Heuberger, Geissbühler and Müller (2005)"},{"why":"Provides the semidefinite programming approach used to compute the polynomial projections.","marker":"Siem, de Klerk and Den Hertog (2008)"}],"fun_headline_variants":["Parametric quantile regression for CT asthma response","Quantile model: beta2=0.638, less air trapping","Closed-form inference in quantile regression on CT data","Asthma CT quantile analysis: shrunk density spread","New quantile regression gives explicit CIs in lung CT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole inferential chain assumes each patient's pre- and post-treatment lung-density distributions are adequately described by Gaussian quantile functions, and that the fitted pre-treatment means and standard deviations are exact covariates with no estimation error; if either fails, the confidence intervals and the $-746.4$ HU threshold no longer have their stated interpretation.","fun_headline_variants_meta":{"raw":{"variants":["Parametric quantile regression for CT asthma response","Quantile model: beta2=0.638, less air trapping","Closed-form inference in quantile regression on CT data","Asthma CT quantile analysis: shrunk density spread","New quantile regression gives explicit CIs in lung CT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000731,"raw_usage":{"total_tokens":3271,"prompt_tokens":947,"completion_tokens":2324,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":2243}},"tokens_in":563,"tokens_out":2324,"duration_ms":15570,"temperature":1.0,"reasoning_tokens":2243,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:39:59.544096+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-fit the analysis with the polynomial degree $d$ chosen by BIC (allowing $d > 1$) on the same 44 patients and compare the empirical quantile functions to their Gaussian projections; if the average $L^2$ approximation error is large relative to the noise variance, or if the $\\beta_2$ confidence interval shifts materially or crosses 1, the Gaussian restriction decides the central claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the original least-squares quantile-function regression model that this paper makes parametric and inferential."},{"cited_title":"u hler , A. A. M \\","cited_arxiv_id":null,"evidence_quote":"Supplies the threshold-based lung segmentation algorithm that isolates the lung voxels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the semidefinite programming approach used to compute the polynomial projections."}],"review_version":1}