REVIEW 3 major objections 5 minor 30 references
Practically significant differences between conditional distribution functions
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper establishes a consistent, asymptotically pivotal test for whether the $L^2$ distance between two conditional distribution functions exceeds a prespecified threshold $\Delta$, using self-normalization so no variance estimation…
desk verdict Useful pivotal test for relevant differences in conditional CDFs; main proof gap is the uniform sequential Bahadur expansion, and the local power formula needs correction. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Self-normalization is the engine. For each sample $\ell=1,2$, the parameter $\beta^\ell(y)$ of the distribution regression model $F^\ell_{Y|X}(y|x)=\Lambda(x^\top\beta^\ell(y))$ is re-estimated on the first $\lfloor n_\ell t\rfloor$ observations, giving sequential estimates $\hat F^\ell_{Y|X}(t,y|x)$ and the difference process $\hat\Delta(t,y|x)=\hat F^1_{Y|X}(t,y|x)-\hat F^2_{Y|X}(t,y|x)$. The numerator is $\hat T_n=\int_I\hat\Delta^2(1,y|x)dy$, an estimator of the squared $L^2$ distance, and the self-normalizer is $\hat V_n=(\int_\epsilon^1(\int_I\hat\Delta^2(t,y|x)dy-\int_I\hat\Delta^2(1,y|x)dy)^2dt)^{1/2}$. The key identity is that the limiting Gaussian process for $\sqrt n(\hat\Delta-\Delta)$ has covariance $(t_1\wedge t_2)/(t_1t_2)\cdot H(y_1,y_2)$, so the leading fluctuation $\sqrt n D_n(t)$ converges to $\tau B(t)/t$. Consequently the ratio $(\hat T_n-\int_I\Delta^2(y|x)dy)/\hat V_n$ converges to $B(1)/(\int_\epsilon^1(B(t)/t-B(1))^2dt)^{1/2}$, a pivotal distribution because the unknown scale $\tau$ cancels.
What would settle it
Simulate two independent samples in which the true link function differs between the groups (for example, a logistic link in one and a skewed link such as the Gumbel distribution in the other) while both are estimated with a normal link, set $\Delta$ equal to the true $L^2$ distance, and record the empirical rejection rate of rule (3.8) at the boundary. If the rejection rate does not stay close to $\alpha$ as the sample sizes grow, Assumption 3.1 is violated and the level claim fails. A second check is to compare the empirical distribution of $(\hat T_n-\int_I\Delta^2(y|x)dy)/\hat V_n$ with simulated quantiles of $W$ under strongly dependent errors; systematic discrepancies show the pivot does not hold outside the i.i.d. setting.
Extended reading notes
Core claim
The paper's central claim is Theorem 3.1: under Assumption 3.1, rejecting $H_0: \|F^1_{Y|X}(\cdot|x)-F^2_{Y|X}(\cdot|x)\|_2\le\Delta$ whenever $\hat T_n>\Delta^2+q_{1-\alpha}\hat V_n$ gives a test whose rejection probability tends to $0$, $\alpha$, or $1$ according as the true squared $L^2$ distance is below, equal to, or above $\Delta^2$, provided $\tau^2>0$ at the boundary. The proof establishes a self-normalized limit law: $\sqrt n(\hat T_n-\int_I\Delta^2(y|x)dy)/\hat V_n$ converges in distribution to $W=B(1)/(\int_\epsilon^1(B(t)/t-B(1))^2dt)^{1/2}$, which depends only on Brownian motion, so the critical value $q_{1-\alpha}$ can be simulated once and reused. The same argument yields nontrivial power against local alternatives of order $1/\sqrt n$, an asymptotically pivotal one-sided confidence interval for the $L^2$ distance, and a data-driven threshold $\hat\Delta_\alpha$ that summarizes the evidence against the null.
Load-bearing premise
The whole argument rests on the assumption that the distribution regression model with the chosen link function is correctly specified in both samples, and that the sequential estimators converge weakly as a stochastic process; if that convergence fails, the pivotal limit distribution and the claimed error control are not established.
Editorial extensions
If this is right
- A researcher who can specify a meaningful threshold $\Delta$ obtains a test with controlled error probability for the question of practical significance, rather than only for whether a nonzero difference exists.
- When $\Delta$ is hard to fix in advance, the quantity $\hat\Delta_\alpha$ from (3.23) gives the smallest threshold at which the null is accepted, serving as a data-driven measure of evidence for similarity.
- The test has nontrivial power against local alternatives converging to the boundary at the parametric $1/\sqrt n$ rate, so it does not lose sensitivity just at the relevance margin.
- The same pivot yields a one-sided asymptotic confidence interval for the squared $L^2$ distance, providing an effect-size statement alongside the hypothesis test.
- The self-normalization principle extends to testing relevant endogeneity, where one conditional distribution is estimated through a control function, preserving the same pivotal structure.
Reading between the lines
- This construction is not tied to the particular link function $\Lambda$: any estimator with a uniform Bahadur expansion and functional weak convergence should carry the same self-normalization argument, so the method could be transferred to quantile regression, duration models, or other semiparametric estimators.
- The data-driven threshold $\hat\Delta_\alpha$ could be re-read as an equivalence bound and reported as 'the two distributions are indistinguishable up to this distance at level $\alpha$', a statement closer to what policy discussions require than a $p$-value.
- The paper's level guarantee rests on one link function being correct in both samples; a stress test the paper does not run is misspecification in which the two samples have different link functions or non-i.i.d. dependence, where the pivot's behavior is not covered by Theorem 3.1 and remains to be checked.
- Connecting to the equivalence-testing tradition, the threshold $\Delta$ could be treated as a design parameter set by cost or welfare considerations, making 'practically significant' explicit before the data are analyzed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a self-normalized test for the null hypothesis that the L2 distance between two conditional distribution functions in a semiparametric distribution regression model is at most a given threshold Delta. The test statistic is the estimated squared L2 difference of the conditional distributions, normalized by a sequential-process version of the same quantity; the limiting distribution of the resulting ratio is a functional of a standard Brownian motion, so critical values can be simulated once and for all. The paper states and proves (Theorem 3.1) that the decision rule has asymptotic level alpha and is consistent, claims parametric-rate local power (Theorem 3.2), reports a small Monte Carlo study (including link misspecification and a jump in the response distribution), and applies the method to SOEP income data. The central theoretical device is the uniform weak convergence in (3.10) of the estimated conditional-distribution difference process, established in the appendix via a uniform Bahadur expansion and a functional central limit theorem for the sequential Z-estimator process.
Significance. If the technical gaps identified below are filled, this paper would make a useful contribution: it addresses a question that is often more relevant than exact equality of conditional distributions, and it does so with a method that avoids bootstrap or variance estimation. The self-normalization idea is elegant and the limit distribution is genuinely pivotal under the stated assumptions. The paper also ships a concrete simulation study and a real-data application, which increases its practical value. The claim that the method is robust to link misspecification is only partially supported (one logistic-versus-normal configuration), but this is an empirical robustness check rather than a central claim. Overall the contribution is significant for empirical economics if the proofs are made rigorous.
major comments (3)
- [Appendix, Theorem A.1 and Theorem 3.1, Eq. (3.10)] The proof of the main theorem is conditional on the uniform weak convergence in (3.10), but the appendix does not actually prove this convergence. Theorem A.1 states a uniform Bahadur expansion with sup_{(t,y)} |D_n(t,y)| -> 0, yet the proof is only a reference to "similar arguments as in the proof of Theorem 5.2 in Chernozhukov et al. (2013)", which does not contain the sequential time dimension t. The sup over both t and y requires stochastic equicontinuity of the normalized score process in both arguments; this is not established and does not follow from the pointwise consistency assumed in Assumption 3.1(2). Likewise, Theorem A.2 invokes Sheehy and Wellner (1992) but does not verify the relevant Donsker conditions for the class of functions indexed by y in the sequential setting. Since the covariance factorization (A.5), the self-normalizing pivot (3.17)-(3.18), and the level claim in Theorem 3.1 all rest on (3.10), this is a load-bearing gap. The authors should either supply a complete proof of (3.10) under Assumption 3.1 or state (3.10) as an explicit high-level assumption and verify it in a separate appendix with sufficient detail.
- [Section 3, Eq. (3.16)] The error bound in (3.16) is stated as o_P(1/sqrt(n)) for the approximation of \hat V_n^2 by the integral of [D_n(t)-D_n(1)]^2 dt. This bound is too weak to justify the joint convergence in (3.17): after multiplication by n, a remainder of order o_P(1/sqrt(n)) becomes o_P(sqrt(n)), which does not vanish and would invalidate the continuous mapping step. The dropped terms in the derivation of (3.16) are actually O_P(1/n^{3/2}) = o_P(1/n) under the uniform weak convergence in (3.10), so the stated result is true with the stronger bound. The proof as written, however, does not state or prove the stronger bound, and the displayed o_P(1/sqrt(n)) is not sufficient. Please correct (3.16) and explicitly justify the uniform O_P(1/sqrt(n)) bound on the estimated difference process that yields the o_P(1/n) remainder.
- [Section 3, Theorem 3.2] The local power result is asserted with the proof deferred to "an inspection of the proof of Theorem A.4", but under local alternatives of the form (3.21) the parameters beta^\ell(y) may depend on n, so the uniform weak convergence in (3.10) must be re-established for a sequence of drifting data-generating processes. This is not an immediate corollary of the proof of Theorem 3.1, which is written for fixed beta^\ell(y). The theorem should either be given a rigorous proof with the required drifting-parameter conditions made explicit, or its statement should be weakened to a high-level condition on the weak convergence of the local rescaled process. As it stands, the claim of nontrivial power at parametric local alternatives is not fully supported.
minor comments (5)
- [Theorem 3.1] In the statement of Theorem 3.1, the limiting rejection probability is written as depending on \int_I \hat\Delta^2(y|x)dy, but it should be \int_I \Delta^2(y|x)dy (the true squared L2 distance); the proof uses the correct quantity.
- [Appendix, Theorem A.1] The definition of R_i^\ell(\beta^\ell(y)) in Theorem A.1 appears to contain a typesetting error: both terms are shown with denominator \Lambda(x^\top\beta^\ell(y)), and since 1 - 1\{Y_i^\ell > y\} = 1\{Y_i^\ell \le y\}, the two displayed terms would cancel. The second denominator should be 1-\Lambda(...).
- [Remark 3.2, Eq. (3.25)] The displayed probability in (3.25) has a dimension mismatch: the term \sqrt{n}(\Delta - M^2) should read \sqrt{n}(\Delta^2 - M^2) (or, alternatively, the notation for the threshold should be aligned with the squared L2 norm used in the hypotheses).
- [Section 5] The application reports p-values but does not explain how the quantiles of the limiting distribution W were simulated for the p-value computation; the text gives only the 95% quantile 1.7546 for epsilon = 0.1. Please add a sentence describing the simulation of W or cite a table with the simulated quantiles.
- [Section 4, Scenario 2b] The text says that the jump specification "violates Assumption 1.3"; this should be Assumption 3.1(3) (the numbering appears to be inconsistent).
Circularity Check
No significant circularity: the test's asymptotic pivot is derived from external empirical-process results, and the rejection boundary is simulated from a Brownian-motion functional rather than fitted from the data.
full rationale
The paper's central derivation is self-contained in the sense required for a non-circular finding. The test statistic T_n and the self-normalizer V_n are constructed directly from the sequential distribution-regression estimators; the critical value q_{1-alpha} is simulated from the Brownian-motion functional W in (3.9), not estimated from the sample that is being tested. Theorem 3.1 is proved by showing weak convergence of the sequential process (3.10), which is then established in the appendix (Theorems A.1-A.4) from Assumption 3.1 plus external empirical-process results: Chernozhukov et al. (2013, Theorem 5.2), Sheehy and Wellner (1992, Theorem 1.1), and van der Vaart and Wellner (2023, functional delta method). None of these citations is authored by the present paper's authors in a load-bearing way. The only self-citations (Dette and Schumann 2024; Kutta and Dette 2024; Wied 2024) are contextual or describe extensions, and the paper does not invoke a uniqueness theorem from the authors' prior work to forbid alternatives. The skeptic's concern that Assumption 3.1(2) states only pointwise consistency while the proof needs uniform-in-(t,y) weak convergence is a proof-completeness or correctness risk, not circularity, because the uniform result is not assumed as the conclusion; it is supposed to be derived in Theorems A.1-A.4. That derivation may be incomplete, but an incomplete proof is not an input-output identity. The simulation study and the SOEP application use the same decision rule for testing, not a fitted parameter disguised as a prediction. Thus no step in the claimed derivation chain reduces by construction or by self-citation to its own inputs.
Assumptions & free parameters
free parameters (1)
- epsilon (trimming constant in the self-normalizing statistic) =
0.1 (default); 0.05 and 0.2 in robustness checks
assumptions (6)
- domain assumption The conditional distribution is correctly specified as F_Y|X(y|x) = Lambda(x^T beta(y)) for a known link function Lambda in both samples.
- standard math Uniform weak convergence of the sequential Z-estimator process, equation (3.10), inherited from Chernozhukov et al. (2013) Theorem 5.2 and the Sheehy-Wellner equivalence.
- domain assumption The response interval I is compact or finite, and in the continuous case the conditional density exists, is uniformly bounded, and is uniformly continuous.
- domain assumption Regressors have finite second moments and the information matrix eigenvalue condition holds uniformly in y.
- ad hoc to paper The trimming constant epsilon > 0 is fixed and bounded away from zero; the quantile q_{1-alpha} is simulated from a standard Brownian motion functional.
- domain assumption The two samples are independent and each is i.i.d., with n1/(n1+n2) converging to c in (0,1).
Cite this review
Pith. "Pith review of Practically significant differences between conditional distribution functions." pith.science (2026). https://pith.science/paper/UUYO36YO
@misc{pith2026250606545,
author = {Pith},
title = {Pith review of: Practically significant differences between conditional distribution functions},
year = {2026},
howpublished = {\url{https://pith.science/paper/UUYO36YO}},
note = {Machine review of arXiv:2506.06545}
}
abstract
In the framework of semiparametric distribution regression, we consider the problem of comparing the conditional distribution functions corresponding to two samples. In contrast to testing for exact equality, we are interested in the (null) hypothesis that the $L^2$ distance between the conditional distribution functions does not exceed a certain threshold in absolute value. The consideration of these hypotheses is motivated by the observation that in applications, it is rare, and perhaps impossible, that a null hypothesis of exact equality is satisfied and that the real question of interest is to detect a practically significant deviation between the two conditional distribution functions. The consideration of a composite null hypothesis makes the testing problem challenging, and in this paper we develop a pivotal test for such hypotheses. Our approach is based on self-normalization and therefore requires neither the estimation of (complicated) variances nor bootstrap approximations. We derive the asymptotic limit distribution of the (appropriately normalized) test statistic and show consistency under local alternatives. A simulation study and an application to German SOEP data reveal the usefulness of the method.
Figures
Reference graph
Works this paper leans on
-
[1]
(1997): A Conditional Kolmogorov Test, Econometrica, 65(5), 1097--1128
Andrews, D. (1997): A Conditional Kolmogorov Test, Econometrica, 65(5), 1097--1128
work page 1997
-
[2]
Berger, J. O. and M. Delampady (1987): Testing Precise Hypotheses , Statistical Science, 2, 317 -- 335
work page 1987
-
[3]
Bradley, R. C. (2007): Introduction to Strong Mixing Conditions. Vol 1-3, Heber City, Utah: Kendrick
work page 2007
-
[4]
Brise \ n o-Sanchez, G., M. Hohberg, A. Groll, and T. Kneib (2020): Flexible Instrumental Variable Distributional Regression, Journal of the Royal Statistical Society, Series A, 183(4), 1553--1574
work page 2020
-
[5]
Chernozhukov, V., I. Fernández-Val, and S. Luo (2025+): Distribution Regression with Sample Selection and UK Wage Decomposition, Journal of Political Economy, forthcoming
work page 2025
-
[6]
Chernozhukov, V., I. Fernández-Val, and B. Melly (2013): Inference on Counterfactual Distributions, Econometrica, 81(6), 2205--2268
work page 2013
-
[7]
Chernozhukov, V., I. Fernández-Val, W. Newey, S. Stouli, and F. Vella (2020): Semiparametric Estimation of Structural Functions in Nonseparable Triangular Models, Quantitative Economics, 11(2), 503--533
work page 2020
-
[8]
Delgado, M., A. García-Suaza, and P. Sant'Anna (2022): Distribution Regression in Duration Analysis: An Application to Unemployment Spells, Econometrics Journal, 25(3), 675--698
work page 2022
Show all 30 references
-
[9]
Destatis (2025): Verbraucherpreisindex und Inflationsrate, https://www.destatis.de/DE/Themen/Wirtschaft/Preise/Verbraucherpreisindex/ \_inhalt.html (07.04.2025)
2025
-
[10]
Dette, H. and M. Schumann (2024): Testing for Equivalence of Pre-Trends in Difference-in-Differences Estimation, Journal of Business and Economic Statistics, 42(4), 1289--1301
2024
-
[11]
Foresi, S. and F. Peracchi (1995): The Conditional Distribution of Excess Returns: An Empirical Analysis, Journal of the American Statistical Association, 90, 451--466
1995
-
[12]
(1988): The Role of Wages in the Inflation Process, American Econonomic Review, 78(2), 276--283
Gordon, R. (1988): The Role of Wages in the Inflation Process, American Econonomic Review, 78(2), 276--283
1988
-
[13]
H \"o rmann, S. and P. Kokoszka (2010): Weakly dependent functional data , The Annals of Statistics, 38, 1845 -- 1884
2010
-
[14]
Hu, X. and J. Lei (2024): A Two-Sample Conditional Distribution Test Using Conformal Prediction and Weighted Rank Sum, Journal of the American Statistical Association, 119(546), 1136--1154
2024
-
[15]
Jordà, O. and F. Nechio (2023): Inflation and Wage Growth since the Pandemic, European Econonomic Review, 156, 104474
2023
-
[16]
Silbersdorff, and B
Kneib, T., A. Silbersdorff, and B. Säfken (2023): Rage Against the Mean - A Review of Distributional Regression Approaches, Econometrics and Statistics, 26, 99--123
2023
-
[17]
Kutta, T. and H. Dette (2024): Validating approximate slope homogeneity in large panels, Journal of Econometrics, 246, 105898
2024
-
[18]
Rothe, C. and D. Wied (2013): Misspecification Testing in a Class of Conditional Distributional Models, Journal of the American Statistical Association, 108, 314--324
2013
-
[19]
Sheehy, A. and J. Wellner (1992): Uniform Donsker Classes of Functions, Annals of Probability, 20(4), 1983--2020
1992
-
[20]
Sozialpolitik-aktuell (2025): Sozialpolitik-aktuell, https://www.sozialpolitik-aktuell.de/files/sozialpolitik-aktuell/\_Politikfelder/Einkommen-Armut/Datensammlung/PDF-Dateien/tabIII1.pdf (07.04.2025)
2025
-
[21]
Sozio-oekonomisches Panel (SOEP) (2022): Version 37, Daten der Jahre 1984-2020 (SOEP-Core v37, EU-Edition), DOI: 10.5684/soep.core.v37eu
2022 doi
-
[22]
Spady, R. and S. Stouli (2025): Gaussian Transforms Modeling and the Estimation of Distributional Regression Functions, Econometrica, forthcoming
2025
-
[23]
Troster, V. and D. Wied (2021): A Specification Test of Dynamic Conditional Distributions, Econometric Reviews, 40(2), 109--127
2021
-
[24]
Tukey, J. W. (1991): The Philosophy of Multiple Comparisons, Statistical Science, 6, 100--116
1991
-
[25]
van der Vaart, A. and J. Wellner (2023): Weak Convergence and Empirical Processes: With Applications to Statistics, Springer Series in Statistics, Cham, Switzerland: Springer, 2nd ed
2023
-
[26]
Oka, and D
Wang, Y., T. Oka, and D. Zhu (2023): Bivariate Distribution Regression with Application to Insurance Data, Insurance: Mathematics and Economics, 113, 215--232
2023
-
[27]
(2010): Testing Statistical Hypotheses of Equivalence and Noninferiority, CRC Press
Wellek, S. (2010): Testing Statistical Hypotheses of Equivalence and Noninferiority, CRC Press
2010
-
[28]
(2024): Semiparametric Distribution Regression with Instruments and Monotonicity, Labour Economics, 90, 102565
Wied, D. (2024): Semiparametric Distribution Regression with Instruments and Monotonicity, Labour Economics, 90, 102565
2024
-
[29]
Wu, W. B. (2005): Nonlinear system theory: Another look at dependence, Proceedings of the National Academy of Sciences, 102, 14150--14154
2005
-
[30]
Zheng, and Z
Zhou, W.-X., C. Zheng, and Z. Zhang (2017): Two-Sample Smooth Tests for the Equality of Distributions, Bernoulli, 23(2), 951--989
2017
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.