Pith. sign in

REVIEW 5 minor 26 references

Reserving a tiny slice of a goodness-of-fit test's Type I error budget for secondary diagnostics can sharply broaden its power without sacrificing the primary test's performance.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 00:25 UTC pith:5F3GKCM7

load-bearing objection Solid, honestly-scoped augmentation method: the calibration logic checks out and the experiment is careful, but prespecification of the budgets is doing real work and the new bottle is an old class of rules.

arxiv 2607.15015 v2 pith:5F3GKCM7 submitted 2026-07-16 stat.ME

Priority-preserving augmentation of goodness-of-fit tests by conditional calibration

classification stat.ME MSC 62G1062F40
keywords goodness-of-fit testingtest augmentationconditional calibrationType I error allocationpower decompositionKolmogorov-Smirnov testsecondary statisticsexact finite-sample test
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that an omnibus goodness-of-fit test can be kept as the primary decision rule and still gain sensitivity to specific departures, by reserving a small, prespecified fraction of its nominal Type I error probability for secondary diagnostics such as sample variance and skewness. The mechanism is an ordered chain in which each secondary acceptance region is calibrated under the null conditional on having passed all earlier stages, so the overall null rejection probability remains exactly alpha while the ordering supplies a clean accounting of what each added statistic contributes. In a normal-null experiment with n=10, handing 0.75% of the 0.05 budget to variance and skewness leaves power against a location shift essentially unchanged (0.2742 to 0.2733) yet lifts power against a scale alternative N(0,0.2^2) from 0.4693 to 0.9827. The paper supports the principle with strong consistency of the Monte Carlo calibration and an exact finite-m pooled-rank variant.

Core claim

The central claim is the augmentation principle: you can broaden the power of an established test by spending a tiny, unconditional slice of its rejection budget on feature-specific statistics, with each secondary statistic's acceptance boundaries calibrated conditional on acceptance by all earlier stages. The resulting test is a fixed rectangular acceptance rule; the ordering is only a mechanism for choosing boundaries and assigning first-rejection credit, not a sequential-sampling scheme. The paper's unconditional budget parameterization b_0 + ... + b_L = alpha makes the trade-off explicit, and the experiment shows that transferring 0.75% of the total 0.05 budget from Kolmogorov-Smirnov to

What carries the argument

The ordered conditional-calibration chain T0 -> T1 -> ... -> TL. Each acceptance region Br is chosen so that Q0(Tr not in Br | X in Ar-1) = gamma_r, with unconditional stage budgets b_r = (product_{s<r}(1-gamma_s)) gamma_r summing to alpha = 1 - product(1-gamma_r). This turns the design problem into a transparent choice of how much null rejection probability to move from primary to secondary diagnostics, and it turns the final test into a fixed rectangular acceptance region. Supporting machinery includes an order-invariance proposition under conditional null independence, a proof that weighted minimum-p tests are equivalent to hyperrectangle tests, a strong-consistency theorem for the Monte

Load-bearing premise

All secondary statistics, their tail rules, their order, and the budget fractions rho and omega must be fixed before looking at the data; if any of these are chosen after inspecting the sample, the claimed null rejection level is no longer guaranteed.

What would settle it

Simulate a practitioner who inspects a few samples from N(0,0.2^2) and then chooses rho and omega to maximize power, while still using the paper's calibration formulas on the same null bank. Under H0, measure the null rejection rate of this post-hoc-tuned chain: it should exceed the nominal 0.05, whereas the same chain with prespecified budgets stays at 0.05.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Any prespecified diagnostic — variance, skewness, tail indices, or problem-specific summaries — can be appended to any primary omnibus test without asymptotically inflating the null level beyond alpha.
  • The cost of augmentation is directly measurable in units of primary rejection probability (alpha - b0), and the benefit is measured stagewise by ordered first-rejection contributions under alternatives.
  • The weak equivalence of weighted minimum-p and hyperrectangle tests means the augmentation results transfer to simultaneous decision rules; the chain differs only in how coordinate thresholds are selected.
  • Under conditional null independence given primary acceptance, the order of secondary statistics does not change the final test, so attribution, not total power, is the order-sensitive output.
  • The pooled-rank variant gives exact finite-m null size for a fixed null law, with the trade-off that the boundaries become observation-specific rather than reusable.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The method suggests a general recipe for any omnibus test: preselect a small set of interpretable diagnostics and fix a tiny budget fraction before seeing data; the same conditional-calibration chain should work for multivariate or two-sample settings with permutation calibration.
  • The near order-invariance seen at small budgets implies that the choice of ordering mainly affects how credit is reported, not the decision; this could let practitioners choose order for interpretability without worrying about power changes.
  • A principled, data-independent rule for selecting rho and omega is still missing; the sensitivity grid is descriptive rather than an optimization, so the next step would be a utility- or worst-case-based budget selector.
  • The exact pooled-rank version may be especially useful when calibration reproducibility matters per observation, at the cost of recomputing boundaries for each new sample.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper proposes an augmentation principle for goodness-of-fit testing: retain a primary omnibus statistic but allocate a small, prespecified fraction of the total null rejection budget to secondary statistics that capture simple feature-specific departures. The implementation is an ordered chain T0→...→TL in which each stage's acceptance region is calibrated under the null conditional on acceptance by all earlier stages. Unconditional stage budgets b_r are converted into conditional levels γ_r via Eq. (15), giving the exact population size identity α=1−∏(1−γ_r). The final test is a fixed rectangular acceptance rule; ordering is a boundary-selection and attribution device. Theoretical support includes strong consistency of the reusable Monte Carlo calibration (Prop. 3.1), an exact randomized finite-m pooled-rank variant (Prop. 3.2), a sufficient condition for order invariance (Prop. 2.1), and an algebraic equivalence to weighted minimum-p/hyperrectangle rules (Prop. 2.2). The experiment augments KS with variance and skewness under a standard normal null with n=10, m=h=10^6, reporting that transferring 0.75% of α=0.05 changes power against N(0.5,1) from 0.2742 to 0.2733 while raising power against N(0,0.2^2) from 0.4693 to 0.9827. A sensitivity grid over ρ and ω, null-rejection checks, and stagewise decompositions are provided.

Significance. If the result holds, the paper supplies a simple, transferable tool: rather than replacing an omnibus test, one can add a few interpretable diagnostics with a tightly controlled error-budget trade-off. The theoretical core is clean: the budget-to-conditional-level mapping (Eqs. (11)–(15)) is algebraically sound, and the proofs of Props. 2.1, 2.2, 3.1, and 3.2 are credible and appropriately conditional on standard regularity assumptions. Strengths include complete, reproducible R code, a prespecified sensitivity grid rather than post-hoc tuning, explicit treatment of finite-m calibration uncertainty, and plain acknowledgment of the proof-of-principle scope. The prespecification requirement (Section 2.8) is a standard and explicit condition, not an internal inconsistency. The reported size checks and paired comparisons make the power claims credible. The paper is not overclaimed: it repeatedly states that the result is not a uniform dominance proof and that the allocation illustrated is not an optimal budget. The methodology should be of interest to practitioners of goodness-of-fit testing and to statistical methodologists working on multi-statistic tests under a single null hypothes

minor comments (5)
  1. [3.3 / abstract] The phrase 'exact randomized finite-m null size' is correct, but the randomized character should be emphasized more prominently at first use. A reader might otherwise infer exactness in the ordinary nonrandomized sense; Section 3.3 itself is clear, but the conclusion restates only 'exact finite-m decision' without the qualifier.
  2. [4.2 / Appendix E] The weighted minimum-p baseline uses self-ranks r/m for calibration and (rank+1)/(m+1) for evaluation. This is asymptotically negligible, and the null-size check in Table 6 directly covers it, but one sentence explaining the finite-sample convention and why it does not affect the comparisons would improve transparency.
  3. [2.6, after Prop. 2.1] The line 'Q(bRσ_m △ bRτ_m) → 0' is an abuse of notation: the probability Q is applied to the symmetric difference. Please write Q(bRσ_m △ bRτ_m)→0 or clarify that it is the Q-measure of the symmetric difference.
  4. [Table 1 caption] The caption 'Bold indicates the largest unrounded estimated power in a column; exact maxima are all shown in bold' is confusing because both clauses appear to say the same thing. It would be clearer to state simply that bold marks the maximum in each column, with exact unrounded values retained.
  5. [4.3.7 / 5] The paper repeatedly and correctly warns that ρ and ω must be prespecified. A short paragraph in the conclusion on possible principled, data-free selection heuristics—e.g., utility functions or worst-case power criteria over a prespecified alternative grid—would strengthen practical applicability. This is a scope suggestion, not a required fix.

Circularity Check

0 steps flagged

No significant circularity: the budgets are design inputs, the power figures are simulated outputs, and the supporting theorems are self-contained.

full rationale

The paper's derivation chain is self-contained. The stage budgets b_r are prespecified design parameters summing to alpha; the conditional levels gamma_r are derived from them by the identity in Eq. (15), and the calibration procedures are then shown to realize those budgets. This is a construction and a finite-sample calibration check, not a prediction of something already contained in the inputs. The headline power figures come from Monte Carlo simulation with m = h = 10^6 and t = 100 repetitions, with full code in Appendices D and E; no parameter is fitted to data to produce the reported power, and the sensitivity grid over 65 (rho, omega) allocations is reported rather than hidden. The central theorems are proved in the appendices: Proposition 3.1 is a Glivenko-Cantelli/quantile-consistency argument, and Proposition 3.2 is an exchangeability argument for the pooled-rank version. Neither imports an unverified result from prior work. The only self-citation, reference [6], appears in the contextual sentence 'This paper continues our earlier work on simultaneous acceptance regions...' and is not used as load-bearing evidence for any theorem, uniqueness claim, or fitted parameter. Proposition 2.1's uniqueness assumption is stated and used within the proof itself, not imported from the author's prior papers. The prespecification requirement in Section 2.8 is a validity condition for the test, not an input that is then re-asserted as an output. Overall, no step in the derivation reduces, by the paper's own equations or by self-citation, to its own inputs.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

No invented physical or probabilistic entities are introduced. The chain is an algorithmic construction, not a new entity. Two design parameters, rho and omega, are hand-chosen and explored with a sensitivity grid. The remaining assumptions are standard regularity conditions or explicitly stated modeling conditions.

free parameters (2)
  • secondary budget fraction rho = 0.0075
    Hand-chosen fraction of the total alpha (0.05) transferred from KS to the secondary statistics in the main experiment. The paper explicitly frames it as an interpretable illustration, with a sensitivity grid over 0 to 0.5 reported in Section 4.3.7.
  • variance split omega = 0.5
    Hand-chosen fraction of the secondary budget assigned to variance rather than skewness. The sensitivity grid over 0, 0.25, 0.5, 0.75, 1 is reported in Section 4.3.7 and Appendix C.
axioms (5)
  • standard math Classical Glivenko-Cantelli theorem for VC classes of axis-aligned rectangles in R^{L+1}.
    Used in the proof of Proposition 3.1 (Appendix A.3) to obtain uniform convergence of the calibration empirical measure over all relevant random region sets.
  • domain assumption At each stage, the population conditional CDF F_r is continuous and strictly increasing at the quantiles defining B_r, and Q0(A_{r-1}) > 0.
    Assumption 2 of Proposition 3.1; required for consistency of empirical conditional quantiles. It is plausible for continuous statistics such as KS, variance, and skewness, but not verified in the paper.
  • domain assumption All statistics, tail rules, orderings, and budgets are fixed independently of the observed sample.
    Stated in Section 2.8; the null calibration is valid only under this prespecification. The paper does not provide a data-independent rule for choosing rho and omega.
  • domain assumption Pooled rows are exchangeable and stagewise selection rules are permutation-equivariant with independent tie-breaking.
    Conditions of Proposition 3.2 (Section 3.3) for the exact finite-m null size of the pooled-rank chain.
  • domain assumption Conditional null independence of secondary statistics given primary acceptance.
    Assumption of Proposition 2.1 for exact order invariance; not used in the main experiment, where order invariance is only approximate.

pith-pipeline@v1.3.0-alltime-deepseek · 35246 in / 17538 out tokens · 200500 ms · 2026-08-02T00:25:02.767668+00:00 · methodology

0 comments
read the original abstract

An omnibus goodness-of-fit statistic can fail to exploit simple diagnostic evidence efficiently. For example, under a standard normal null, an exceptionally large observation is direct evidence of a scale or tail departure even when the primary omnibus statistic does not cross its critical value. This motivates a simple augmentation principle: retain an established primary test, but reserve a small part of its null rejection budget for secondary statistics that encode natural features such as variation, asymmetry, or tail behavior. We implement this principle by calibrating each secondary acceptance region under the null conditional on acceptance at all preceding stages. Once calibrated, the resulting test is a fixed rectangular acceptance rule; the ordering is a mechanism for choosing its boundaries and assigning ordered first-rejection contributions, not a sequential-sampling scheme. An unconditional stage-budget parameterization makes the central trade-off explicit: additional sensitivity is purchased by removing a prespecified, usually small, amount of rejection probability from the primary stage. We establish strong consistency of the quantile-based Monte Carlo calibration and give an observation-specific pooled-rank version with exact randomized finite-m null size. In an experiment under a standard normal null with n=10, we augment the Kolmogorov--Smirnov statistic with sample variance and sample skewness. Assigning only 0.75% of the total Type I error budget to the two secondary statistics changes power against N(0.5,1) from 0.2742 to 0.2733, while increasing power against N(0,0.2^2) from 0.4693 to 0.9827. This focused experiment is a proof of principle rather than an exhaustive comparison of normality tests: a small unconditional allocation to prespecified diagnostics can greatly broaden power while preserving almost all of the primary test's power.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 7 canonical work pages · 3 internal anchors

  1. [1]

    Brown, Andreas Buja, Wolfgang Rolke, and Robert A

    Sivan Aldor-Noiman, Lawrence D. Brown, Andreas Buja, Wolfgang Rolke, and Robert A. Stine. The power to see: A new graphical test of normality.The American Statistician, 67(4):249–260, 2013.doi: 10.1080/00031305.2013.847865

  2. [2]

    Test of significance based on wavelet thresholding and Neyman’s truncation.Journal of the American Statistical Association, 91(434):674–688, 1996.doi:10.2307/2291663

    Jianqing Fan. Test of significance based on wavelet thresholding and Neyman’s truncation.Journal of the American Statistical Association, 91(434):674–688, 1996.doi:10.2307/2291663

  3. [3]

    Fisher.Statistical Methods for Research Workers

    Ronald A. Fisher.Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh, 4 edition, 1932. URL:https://archive.org/details/in.ernet.dli.2015.205971/mode/2up

  4. [4]

    Adaptive goodness-of-fit tests in a density model.The Annals of Statistics, 34(2):680–720, 2006.arXiv:math/0607013,doi:10.1214/009053606000000119

    Magalie Fromont and B´ eatrice Laurent. Adaptive goodness-of-fit tests in a density model.The Annals of Statistics, 34(2):680–720, 2006.arXiv:math/0607013,doi:10.1214/009053606000000119

  5. [5]

    Goeman and Aldo Solari

    Jelle J. Goeman and Aldo Solari. The sequential rejection principle of familywise error control.The Annals of Statistics, 38(6):3782–3810, 2010.doi:10.1214/10-AOS829

  6. [6]

    Goodness of Fit Tests Based on Joint Densities of Multiple Sample Statistics

    Roman Guchenko. Goodness of fit tests based on joint densities of multiple sample statistics.arXiv preprint arXiv:2607.02285, 2026. URL:https://arxiv.org/abs/2607.02285,doi:10.48550/arXiv.2607.02285

  7. [7]

    Wilbert C. M. Kallenberg and Teresa Ledwina. Consistency and Monte Carlo simulation of a data driven version of smooth goodness-of-fit tests.The Annals of Statistics, 23(5):1594–1608, 1995.doi: 10.1214/aos/1176324315

  8. [8]

    King, Xibin Zhang, and Muhammad Akram

    Maxwell L. King, Xibin Zhang, and Muhammad Akram. Hypothesis testing based on a vector of statistics. Journal of Econometrics, 219(2):425–455, 2020.doi:10.1016/j.jeconom.2020.03.010

  9. [9]

    Data-driven version of Neyman’s smooth test of fit.Journal of the American Statistical Association, 89(427):1000–1005, 1994.doi:10.1080/01621459.1994.10476834

    Teresa Ledwina. Data-driven version of Neyman’s smooth test of fit.Journal of the American Statistical Association, 89(427):1000–1005, 1994.doi:10.1080/01621459.1994.10476834

  10. [10]

    Yaowu Liu and Jun Xie. Cauchy combination test: A powerful test with analyticp-value calculation under arbitrary dependency structures.Journal of the American Statistical Association, 115(529):393–402, 2020. doi:10.1080/01621459.2018.1554485

  11. [11]

    smooth test

    Jerzy Neyman. “smooth test” for goodness of fit.Skandinavisk Aktuarietidskrift, 20(3–4):149–199, 1937

  12. [12]

    Jerzy Neyman.A Selection of Early Statistical Papers of J. Neyman. University of California Press, 1 edition, 1967. URL:http://www.jstor.org/stable/jj.8501421

  13. [13]

    Supplemental studies for simultaneous goodness-of-fit testing.arXiv preprint arXiv:2007.04727, 2020

    Wolfgang Rolke. Supplemental studies for simultaneous goodness-of-fit testing.arXiv preprint arXiv:2007.04727, 2020. URL:https://arxiv.org/abs/2007.04727,doi:10.48550/arXiv.2007.04727

  14. [14]

    Simulation studies for goodness-of-fit and two-sample methods for univariate data.arXiv preprint arXiv:2411.05839, 2024

    Wolfgang Rolke. Simulation studies for goodness-of-fit and two-sample methods for univariate data.arXiv preprint arXiv:2411.05839, 2024. URL:https://arxiv.org/abs/2411.05839,doi:10.48550/arXiv. 2411.05839

  15. [15]

    Power Studies For Two-Sample and Goodness-of-Fit Methods For Multivariate Data

    Wolfgang Rolke. Power studies for two-sample and goodness-of-fit methods for multivariate data.arXiv preprint arXiv:2605.12089, 2026. URL:https://arxiv.org/abs/2605.12089,doi:10.48550/arXiv. 2605.12089

  16. [16]

    Romano and Michael Wolf

    Joseph P. Romano and Michael Wolf. Exact and approximate stepdown methods for multiple hypoth- esis testing.Journal of the American Statistical Association, 100(469):94–108, 2005.doi:10.1198/ 016214504000000539

  17. [17]

    Teemu S¨ ailynoja, Paul-Christian B¨ urkner, and Aki Vehtari. Graphical test for discrete uniformity and its applications in goodness-of-fit evaluation and multiple sample comparison.Statistics and Computing, 32(2):1–21, 2022.doi:10.1007/s11222-022-10090-6

  18. [18]

    KSD aggregated goodness-of-fit test

    Antonin Schrab, Benjamin Guedj, and Arthur Gretton. KSD aggregated goodness-of-fit test. InAdvances in Neural Information Processing Systems, volume 35, 2022. URL:https://arxiv.org/abs/2202.00824

  19. [19]

    MMD aggregated two-sample test.Journal of Machine Learning Research, 24(194):1–81, 2023

    Antonin Schrab, Ilmun Kim, M´ elisande Albert, B´ eatrice Laurent, Benjamin Guedj, and Arthur Gretton. MMD aggregated two-sample test.Journal of Machine Learning Research, 24(194):1–81, 2023. URL: https://arxiv.org/abs/2110.15073

  20. [20]

    Spokoiny

    Vladimir G. Spokoiny. Adaptive hypothesis testing using wavelets.The Annals of Statistics, 24(6):2477– 2498, 1996.doi:10.1214/aos/1032181163

  21. [21]

    Tamhane, Jiangtao Gou, Christopher Jennison, Cyrus R

    Ajit C. Tamhane, Jiangtao Gou, Christopher Jennison, Cyrus R. Mehta, and Teresa Curto. A gatekeeping procedure to test a primary and a secondary endpoint in a group sequential design with multiple interim looks.Biometrics, 74(1):40–48, 2018.doi:10.1111/biom.12732

  22. [22]

    Leonard H. C. Tippett.The Methods of Statistics. Williams and Norgate, London, 1931. URL:https: //archive.org/details/in.ernet.dli.2015.189563/mode/2up. 23

  23. [23]

    Admissible ways of merging p-values under arbitrary dependence

    Vladimir Vovk, Bin Wang, and Ruodu Wang. Admissible ways of mergingp-values under arbitrary dependence.arXiv preprint arXiv:2007.14208, 2020. URL:https://arxiv.org/abs/2007.14208,doi: 10.48550/arXiv.2007.14208

  24. [24]

    Combiningp-values via averaging.Biometrika, 107(4):791–808, 2020

    Vladimir Vovk and Ruodu Wang. Combiningp-values via averaging.Biometrika, 107(4):791–808, 2020. doi:10.1093/biomet/asaa027

  25. [25]

    Westfall and S

    Peter H. Westfall and S. Stanley Young.Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment. John Wiley & Sons, New York, 1993. A Proofs of the theoretical results A.1 Proof of Proposition 2.1 Proof of Proposition 2.1.Fix a secondary statisticT j and a setS⊆ {1, . . . , L} \ {j}. Conditional null inde- pendence gives Q0(Tj(X)/∈B ...

  26. [26]

    mean=", s)) 19), 20function(mean) function(n) rnorm(n, mean, 1) 21), 22lapply( 23setNames( 24normal.sd.vector, 29 25sapply(normal.sd.vector, function(s) paste0(

    { 7int m = samples.nrow(); // infer number of samples from samples matrix 8int n = samples.ncol(); // infer sample size from samples matrix 9 10NumericVector out(m); // allocate memory for the final result 11 12for (int i = 0; i < m; i++) { // loop over rows of samples matrix 13NumericVector sample = samples(i, _); // get specific sample 14sample.sort(); ...