Pith. sign in

REVIEW 5 minor 1 cited by

Calibrating aggregated test statistics on the permutation distribution itself is finite-sample valid and at least as powerful as Bonferroni-type worst-case calibration, with sequential and two-batch extensions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Calibrating aggregated test statistics on the permutation distribution itself is finite-sample valid and at least as powerful as Bonferroni-type worst-case calibration, with sequential and two-batch extensions.

T0 review reviewed 2026-08-01 challenge →

load-bearing objection A clean, correct paper on permutation-based aggregation with a genuine dominance theorem; the main caveat is the exact-exchangeability assumption, which the authors state plainly.

arxiv 2607.15823 v1 pith:ODIEGRT3 submitted 2026-07-17 stat.ME cs.LGmath.STstat.MLstat.TH

Aggregation of Statistical Evidence under Exchangeability

classification stat.ME cs.LGmath.STstat.MLstat.TH
keywords aggregationevidencedependencetransformedunderacrosscalibrationdatasets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Suppose you want to run a scientific test but the best way to summarize the evidence is unclear: one statistic catches one kind of signal, another catches a different kind. A natural tactic is to compute several statistics and combine them into a single p-value. The classic fix for the resulting multiple testing problem, Bonferroni correction, stays valid no matter how the statistics are correlated, but it pays for this by assuming the worst-case correlation.

This paper works with a construction that is more careful. It applies all K statistics to the original data and to B randomly shuffled copies. Within each copy, every statistic is converted to a permutation p-value, and the K p-values are merged with a function of the user's choice (minimum, mean, median, anything). Because every copy is processed in exactly the same way, the B+1 merged scores are exchangeable under the null hypothesis: the original data is just one of B+1 symmetric siblings. Ranking the original score among its shuffled siblings therefore gives an exact, finite-sample p-value (Theorem 1).

The central power result (Theorem 2) is a dominance argument: any threshold that is valid for every possible dependence pattern must in particular be valid for the specific pattern displayed by the permutation copies, so the data-calibrated threshold is always at least as large as the worst-case one. The paper then extends the idea in three ways: a sequential version that spends the significance budget across ordered stages and can stop early; a two-batch version that standardizes on one batch of shuffles and calibrates on another, allowing the merging rule itself to be learned; and a conformal-prediction application where, for absolute-residual scores with minimum merging, the merged prediction set is an expli

Core claim

Theorem 2: 'Suppose that c_α,K is a deterministic constant satisfying sup_ν P_ν(f(U_1,...,U_K) ≤ c_α,K) ≤ α. Then it holds almost surely that c_α,K < û^SB_α. Consequently, ... P(f_0 ≥ û^SB_α) ≤ P(f_0 > c_α,K).' Together with Theorem 1 (P(p_SB ≤ α) ≤ α for all α, B, any merge function f), the load-bearing assertion is that single-batch permutation-calibrated aggregation is finite-sample valid under the group-invariance null and at least as powerful as any deterministic calibration valid under arbitrary dependence (Bonferroni, O/M-family merging), while adapting to the realized dependence (Propositions 2-3). If correct, SB/TB aggregation can replace worst-case calibration in adaptive tests and conformal merging without losing level control.

Load-bearing premise

Exact joint exchangeability of the transformed rows under the null: the transformations g_1,...,g_B must be chosen independently of the data (i.i.d. uniform on G or without-replacement sampling) so that the row vectors (T^1_b,...,T^K_b) are exchangeable in b, and the standardization, merge, and tie-breaking maps must be row-wise equivariant. Location: Section 2.1 (Definition 1 and 'Unless stated otherwise, we assume that the transformations are chosen so that the transformed data g_0(X),...,g_B(X) are exchangeable'); Section 3.2 (row-wise application of f); Section 5.3 (learning the rule on the testing batch 'would generally destroy the exchangeability'); Section 8 ('The current guarantees assume that the transformed rows are exactly exchangeable'). If the null holds only approximately, or transformations are chosen with the data, every finite-sample guarantee in Theorems 1-3 and Proposition 5 is void.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper studies aggregation of multiple test statistics under the group-invariance (randomization) null. In single-batch (SB) aggregation, the same B+1 transformed datasets are used to compute permutation p-values for each of K statistics, merge them row-wise through an arbitrary function f, and calibrate the merged value by its rank among all rows. Theorem 1 establishes finite-sample type I error control for any f. Theorem 2 shows that the SB critical value almost surely exceeds any deterministic threshold that is valid under arbitrary dependence among super-uniform p-values, yielding power dominance over worst-case calibrated rules such as O- and M-family merging (Corollary 1). Propositions 2 and 3 show adaptivity to perfect rank alignment and asymptotic oracle calibration. Section 4 adds a sequential alpha-spending version (Proposition 5), and Section 5 adds a two-batch (TB) construction that allows reference-batch-dependent aggregation rules while preserving validity (Theorem 3). Applications to adaptive nonparametric testing and conformal prediction are given, including a closed-form interval intersection for minimum-merging TB in Proposition 6. Numerical experiments illustrate finite-sample power, level, and computational gains.

Significance. If the results hold, they provide a clean and useful unification: exact finite-sample validity for essentially arbitrary merging functions, a simple argument showing uniform dominance over deterministic worst-case calibrations, and explicit dependence adaptivity. The paper's strengths include elementary, checkable proofs, no fitted constants or circular calibration, explicit credit to the prior permutation-combination literature, and honest discussion of limitations in Section 8. The main scope caveat is the exact-exchangeability assumption on the transformed rows: without it, Theorems 1–3 and Proposition 5 have no finite-sample content. The paper states this condition in Section 2.1 and acknowledges the approximate-exchangeability gap in Section 8, and the primary applications (permutation and split-conformal methods) satisfy the condition. I therefore regard the assumption as a stated scope limitation rather than an internal inconsistency. The empirical work supports, rather than establishes, the theoretical claims, which is appropriate here.

minor comments (5)
  1. [Sections 2.1 and 8] The exact-exchangeability condition on the transformed rows is the single most important scope condition. Since Section 8 correctly notes that approximate exchangeability is open, I suggest adding one sentence in the abstract or introduction making explicit that all finite-sample guarantees require exact exchangeability, so that the phrase 'potentially complex dependence' is not over-read as relaxing this structural assumption.
  2. [Section 3.3, Theorem 2] The 'almost surely' qualifier in c_{α,K} < û^SB_α appears unnecessary: Lemma 1 is deterministic for any realized statistic matrix, so the inequality holds for every realization. Making this pointwise nature explicit would simplify the statement and clarify that no additional probabilistic regularity is needed.
  3. [Sections 3.5 and Proposition 1] The strict inequality in the SB minimum rejection rule is essential for the 1/(B+1) shift in Proposition 4. Since the threshold definition in Eq. (6) uses a supremum over {F(u) ≤ α}, the open/closed half-line distinction is easy to misread. A short remark after Proposition 1 explaining the strict-vs-non-strict convention would prevent errors.
  4. [Figures 1 and 2] The figure captions are compressed. In Figure 1, the left panel uses L1, L4, L∞ while the right panel labels L1, L2, L3, and the 'SB SB ties TB TB ties' entries do not map unambiguously to line types. In Figure 2, the decision-stage legend is unclear. Expanding the captions or adding explicit legends would materially improve readability.
  5. [Section 3.2, Eq. (4)] The permutation p-value definition uses the non-strict inequality 1(T^i_k ≥ T^b_k). This is standard, but the paper later emphasizes the importance of tie-breaking; connecting Eq. (4) explicitly to the tie-breaking discussion in Appendix B.6 in the main text would help readers understand why the conservative definition is used before tie-breaking is introduced.

Circularity Check

0 steps flagged

No significant circularity; the validity and dominance theorems are proved from first principles, and self-citations serve as comparison baselines rather than load-bearing support.

full rationale

The paper's central claims are derived within the manuscript. Theorem 1 and Theorem 3 rest on the standard exchangeability rank argument (row vectors (T^1_b,...,T^K_b) are exchangeable under the group-invariance null, so the rank statistic p_SB is super-uniform), stated and justified in Sections 3.2 and 5.2, with proofs in the supplement. Theorem 2 is not an identity: the SB threshold û^SB_α is defined directly from the realized merged values (Proposition 1), while the deterministic constant c_{α,K} is defined by a supremum over all joint laws with super-uniform margins; the link between them is Lemma 1, a deterministic count bound on empirical super-uniformity of permutation p-values, from which the feasibility of c for the empirical constraint and the strict inequality c < û follow. This is a genuine two-step argument, not a re-statement of an input. The sequential (Proposition 5) and TB results likewise reduce to row-removal counting and conditional exchangeability. No parameter is fitted to data and then reported as a prediction; thresholds are internally calibrated. Self-citations ([26], [61]–[63], [34], [38], [77], etc.) are used as historical attribution ([47], [49], [64], [75]) or as comparison baselines (MaxT with Monte Carlo calibration, rank-transformed subsampling), and are not the foundation of any load-bearing theorem; the alleged liberality of MaxT is proved in-appendix (Proposition S.8), not imported. No uniqueness theorem, no ansatz-smuggling citation, and no renaming of a known result is performed: Section 3.5 explicitly acknowledges that minimum merging ``recovers the Westfall--Young single-step method.'' The paper's own stated limitations (Section 8: exact exchangeability of transformed rows; Section 5.3: learning the aggregation rule on the calibration batch would destroy exchangeability) are honest scope conditions, not circular steps. Under the stated exact-exchangeability assumption, the derivation chain is self-contained, so the circularity score is minimal.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

Counts: 0 free parameters, 6 axioms (all standard domain assumptions or standard math, none ad hoc to the paper), 0 invented entities. The framework imports the randomization hypothesis, the row-wise equivariance discipline, and standard rank-exchangeability facts; everything else is derived. User-set inputs (B, α_j spending sequence, n_1/n_2 batch split, choice of f) are method parameters, not fitted values. The auxiliary tie-breaking variables (Appendix B.6) are standard randomization devices, not evidence-generating postulates.

axioms (6)
  • domain assumption Group-invariance hypothesis (Definition 1): under H0, g(X) has the same distribution as X for every g ∈ G.
    The validity theorems (Thm 1, Thm 3, Prop 5) hold only under this exact randomization hypothesis; Section 8 states the framework does not cover approximate randomization: 'The current guarantees assume that the transformed rows are exactly exchangeable.'
  • domain assumption The transformations g_1,...,g_B are drawn independently of X and jointly so that g_0(X),...,g_B(X) are exchangeable (i.i.d. uniform on G, or without-replacement sampling).
    Stated in Section 2.1: 'Unless stated otherwise, we assume that the transformations are chosen so that the transformed data g_0(X),...,g_B(X) are exchangeable under the group-invariance hypothesis.' If transformations are chosen using the data, row exchangeability and hence Theorem 1 fail.
  • domain assumption The standardization map (permutation p-values), the merging function f, and any tie-breaking augmentation are applied row-wise and are equivariant under row permutations.
    Section 3.2: 'Applying the merging function row-wise therefore yields aggregated values (f_0,...,f_B) that are also exchangeable.' Section 5.3: learning the aggregation rule from the testing batch 'would generally destroy the exchangeability underpinning the SB procedure.'
  • domain assumption TB/conformal: conditional exchangeability of testing-batch rows (or aggregation-batch scores plus test point) given the reference batch; the aggregation rule may depend only on reference-batch information.
    Section 5.2: testing rows are 'conditionally exchangeable given the reference batch'; Section 6.2 (split conformal setup). This legitimizes learned aggregation rules and the precomputed TB threshold.
  • standard math Rank-p-value super-uniformity of exchangeable tuples (Lemma 1 and its generalization Lemma S.8, extending Harrison 2012, Lemma A1).
    Proved in the paper; the load-bearing analytic fact behind Theorem 2 and all validity statements.
  • domain assumption Asymptotic regime condition for Proposition 3: (f_{1,n}, f_{2,n}) converge jointly to i.i.d. limits with a well-separated oracle quantile Q*_α.
    Section 3.4 imposes this to conclude SB adaptivity to the oracle quantile. Pairwise convergence to independent limits suffices for EDF convergence of bounded exchangeable indicators, so the condition is adequate, but it is imported rather than derived.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Aggregation of Statistical Evidence under Exchangeability." pith.science (2026). https://pith.science/paper/ODIEGRT3

@misc{pith2026260715823,
  author       = {Pith},
  title        = {Pith review of: Aggregation of Statistical Evidence under Exchangeability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ODIEGRT3}},
  note         = {Machine review of arXiv:2607.15823}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We study aggregation of statistical evidence under unknown and potentially complex dependence using group-invariance. Building on permutation-based constructions that treat transformed datasets as exchangeable units, we aggregate evidence across statistics for each transformed dataset and calibrate the resulting aggregates across transformations. We develop a finite-sample power and adaptivity theory for this framework, together with extensions to sequential and data-dependent aggregation that preserve validity. For single-batch aggregation, which uses one collection of transformed datasets for both standardization and calibration, we show that the critical values uniformly improve on deterministic calibrations valid under arbitrary dependence, including Bonferroni correction, while adapting to the unknown dependence structure. We also introduce a sequential alpha-spending version that permits early rejection when evidence is strong, and a two-batch extension that separates standardization from calibration to accommodate learned aggregation rules and reduce computation. Applications to adaptive nonparametric testing and conformal prediction illustrate how these results sharpen existing aggregation methods.

Figures

Figures reproduced from arXiv: 2607.15823 by Antonin Schrab, Arthur Gretton, Ilmun Kim, Rajen Shah.

Figure 1
Figure 1. Figure 1: Power estimation for two-sample d-dimensional mean shift detection. by ∆/2 and of the second sample by −∆/2, yielding a true mean difference of ∆ across the d shifted dimensions. The empirical power is averaged over 1000 independent repetitions at a nominal significance level of α = 0.05. For the first experiment varying signal sparsity in [PITH_FULL_IMAGE:figures/full_fig_p023_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Power estimation for kernel-based MMD two-sample nonparametric testing in a sequen￾tial setting. 7.2 Sequential two-sample nonparametric testing We evaluate the empirical performance of the Sequential SB minimum test (SeqSB, see Algorithm 2) using simulated data evaluated across K = 10 sequential stages with equal spending αj = α/K for j ∈ [K], designed to introduce a controlled correlation structure along… view at source ↗
Figure 3
Figure 3. Figure 3: Marginal coverage, efficiency (average prediction set size) and runtimes (in seconds) for merging multiple conformal prediction sets. Appendix B.1) such as the minimum p-merging function f(p1, . . . , pK) = K min(p1, . . . , pK), the mean p-merging function f(p1, . . . , pK) = 2 mean(p1, . . . , pK) and the median p-merging function f(p1, . . . , pK) = 2 median(p1, . . . , pK), with scaling specifically ch… view at source ↗
Figure 4
Figure 4. Figure 4: Schematic illustrations of (a) the SB aggregation procedure of Algorithm 1 and (b) the TB aggregation procedure of Algorithm 3, using statistics T 1 , . . . , T K and merging function f. 33 [PITH_FULL_IMAGE:figures/full_fig_p033_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Type I error estimation for the SB/TB, Data-Driven SB/TB, SeqSB and MaxT tests. experiments in the rest of this section to ensure fair comparisons across tests. We note that for K > 1 the MaxT test exhibits a similar behavior (failing to control the type I error at level α) but deviates from the theoretical level derived for K = 1. All other tests have type I error bounded above by α as desired. Due to the… view at source ↗
Figure 6
Figure 6. Figure 6: Power estimation for one-sample zero mean testing of an equicorrelated multivariate Gaussian. dures. First, the worst-case (WC) tests are highly conservative, achieving near-zero empirical power across almost all settings. In the sparse signal regime (left panel), among the three single-step methods, only the SB Min test demonstrates strong performance. Because the mean shift is iso￾lated to a single dimen… view at source ↗
Figure 7
Figure 7. Figure 7: Power estimation for kernel-based HSIC independence nonparametric testing. with Gaussian kernels, is computed over a multi-scale grid where the total number of bandwidths evaluated, K, corresponds to the product of the number of candidate bandwidths for X and Y (i.e., K = KX × KY ). Empirical power is averaged over 1000 independent repetitions, all tests are calibrated using B = 199 permutations, and the n… view at source ↗
Figure 8
Figure 8. Figure 8: Power estimation for aggregated kernel-based MMDAgg and HSICAgg nonparametric testing. C.4 Improved MMDAgg and HSICAgg optimal tests In the experimental setting of [PITH_FULL_IMAGE:figures/full_fig_p048_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Schematic illustration of the Data-driven SB aggregation procedure of Appendix D.2 (Algorithm 7), using statistics T 1 , . . . , T K and merging functions f 1 , . . . , fM. 50 [PITH_FULL_IMAGE:figures/full_fig_p050_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Sharp Minimax Rates for Smooth Two-Sample Testing under Central Differential Privacy

    math.ST 2026-07 conditional novelty 7.0

    Under central differential privacy, the sharp L1 separation radius for two-sample testing of Hölder-smooth densities is the maximum of the classical rate and three privacy barriers, and adapting to unknown smoothness ...

Reference graph

Works this paper leans on

79 extracted references · 5 linked inside Pith · cited by 1 Pith paper

  1. [1]

    (2015).Tests of independence by bootstrap and permutation: an asymptotic and non- asymptotic study

    Albert, M. (2015).Tests of independence by bootstrap and permutation: an asymptotic and non- asymptotic study. Application to neurosciences.PhD thesis, Université Nice Sophia Antipolis

  2. [2]

    Albert, M., Laurent, B., Marrel, A., and Meynaoui, A. (2022). Adaptive test of independence based on HSIC measures.The Annals of Statistics, 50(2):858–879

  3. [3]

    N., Barber, R

    Angelopoulos, A. N., Barber, R. F., and Bates, S. (2024). Theoretical Foundations of Conformal Prediction. arXiv preprint arXiv:2411.11824

  4. [4]

    Baraud, Y., Huet, S., and Laurent, B. (2003). Adaptive tests of linear hypotheses by model selection. The Annals of Statistics, 31(1):225–251

  5. [5]

    B., Kontoyiannis, I., and Samworth, R

    Berrett, T. B., Kontoyiannis, I., and Samworth, R. J. (2021). Optimal rates for independence testing via U-statistic permutation tests.The Annals of Statistics, 49(5):2457–2490

  6. [6]

    B., Wang, Y., Barber, R

    Berrett, T. B., Wang, Y., Barber, R. F., and Samworth, R. J. (2020). The conditional permu- tation test for independence while controlling for confounders.Journal of the Royal Statistical Society Series B: Statistical Methodology, 82(1):175–197. 27

  7. [7]

    Biggs, F., Schrab, A., and Gretton, A. (2023). MMD-FUSE: Learning and combining kernels for two-sample testing without data splitting.Advances in Neural Information Processing Systems, 36

  8. [8]

    Candes, E., Fan, Y., Janson, L., and Lv, J. (2018). Panning for gold: ‘Model-X’ knockoffs for high dimensional controlled variable selection.Journal of the Royal Statistical Society Series B: Statistical Methodology, 80(3):551–577

  9. [9]

    Caughey, D., Dafoe, A., and Seawright, J. (2017). Nonparametric combination (NPC): A frame- work for testing elaborate theories.The Journal of Politics, 79(2):688–701

  10. [10]

    Cha, S., Lee, S., Schrab, A., and Kim, I. (2026). More Permutations Do Not Always Increase Power: Non-monotonicity in Monte Carlo Permutation Tests.arXiv preprint arXiv:2605.03886

  11. [11]

    L., Schrab, A., Gretton, A., Sejdinovic, D., and Muandet, K

    Chau, S. L., Schrab, A., Gretton, A., Sejdinovic, D., and Muandet, K. (2025). Credal two- sample tests of epistemic uncertainty. InProceedings of The 28th International Conference on Artificial Intelligence and Statistics, volume 258 ofProceedings of Machine Learning Research, pages 127–135. PMLR

  12. [12]

    and Kim, I

    Choi, W. and Kim, I. (2023). Averaging p-values under exchangeability.Statistics & Probability Letters, 194:109748

  13. [13]

    and Romano, J

    Chung, E. and Romano, J. P. (2013). Exact and asymptotically robust permutation tests.The Annals of Statistics, 41(2):484–507

  14. [14]

    Cox, D. R. (1975). A note on data-splitting for the evaluation of significance levels.Biometrika, 62(2):441–444

  15. [15]

    Domingo-Enrich, C., Dwivedi, R., and Mackey, L. (2025). Cheap permutation testing.arXiv preprint arXiv:2502.07672

  16. [16]

    Fisher, R. A. (1925).Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh

  17. [17]

    Fisher, R. A. (1935).The Design of Experiments. Oliver and Boyd, Edinburgh

  18. [18]

    and Laurent, B

    Fromont, M. and Laurent, B. (2006). Adaptive goodness-of-fit tests in a density model.The Annals of Statistics, 34(2):680–720

  19. [19]

    Fromont, M., Laurent, B., and Reynaud-Bouret, P. (2013). The two-sample problem for pois- son processes: Adaptive tests with a nonasymptotic wild bootstrap approach. The Annals of Statistics, 41(3):1431–1461

  20. [20]

    and Ramdas, A

    Gasparin, M. and Ramdas, A. (2024). Merging uncertainty sets via majority vote. arXiv preprint arXiv:2401.09379

  21. [21]

    Gasparin, M., Wang, R., and Ramdas, A. (2025). Combining exchangeable p-values.Proceed- ings of the National Academy of Sciences, 122(11):e2410849122

  22. [22]

    (2005).Permutation, parametric and bootstrap tests of hypotheses

    Good, P. (2005).Permutation, parametric and bootstrap tests of hypotheses. Springer

  23. [23]

    Gretton, A. (2015). A simpler condition for consistency of a kernel independence test.arXiv preprint arXiv:1501.06103. 28

  24. [24]

    M., Rasch, M

    Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A. (2012). A kernel two-sample test.Journal of Machine Learning Research, 13(25):723–773

  25. [25]

    Gretton, A., Herbrich, R., Smola, A., Bousquet, O., and Schölkopf, B. (2005). Kernel methods for measuring independence.Journal of Machine Learning Research, 6:2075–2129

  26. [26]

    Guo, F. R. and Shah, R. D. (2025). Rank-transformed subsampling: inference for multiple data splitting and exchangeable p-values.Journal of the Royal Statistical Society Series B: Statistical Methodology, 87(1):256–286

  27. [27]

    Hagrass, O., Sriperumbudur, B., and Li, B. (2024). Spectral regularized kernel two-sample tests. The Annals of Statistics, 52(3):1076–1101

  28. [28]

    Conservativehypothesistestsandconfidenceintervalsusingimportance sampling

    Harrison, M.T.(2012). Conservativehypothesistestsandconfidenceintervalsusingimportance sampling. Biometrika, 99(1):57–69

  29. [29]

    I., and Dieuleveut, A

    Hegazy, M., Aolaritei, L., Jordan, M. I., and Dieuleveut, A. (2025). Valid selection among conformal sets. InAdvances in Neural Information Processing Systems, volume 38

  30. [30]

    and Goeman, J

    Hemerik, J. and Goeman, J. (2018). Exact testing with random permutations.Test, 27(4):811– 825

  31. [31]

    and Suslina, I

    Ingster, Y. and Suslina, I. A. (2012). Nonparametric goodness-of-fit testing under Gaussian models, volume 169. Springer Science & Business Media

  32. [32]

    D., Bühlmann, P., and Samworth, R

    Janková, J., Shah, R. D., Bühlmann, P., and Samworth, R. J. (2020). Goodness-of-fit testing in high dimensional generalized linear models.Journal of the Royal Statistical Society Series B: Statistical Methodology, 82(3):773–795

  33. [33]

    B., and Yu, Y

    Kent, A., Berrett, T. B., and Yu, Y. (2026). Locally Differentially Private Two-Sample Testing. Biometrika, page asag034

  34. [34]

    Kim, I., Balakrishnan, S., and Wasserman, L. (2022). Minimax optimality of permutation tests. The Annals of Statistics, 50(1):225–251

  35. [35]

    Kim, I., Neykov, M., Balakrishnan, S., and Wasserman, L. (2024). Conditional indepen- dence testing for discrete distributions: Beyondχ2- and G-tests.Electronic Journal of Statistics, 18(2):4767–4794

  36. [36]

    and Ramdas, A

    Kim, I. and Ramdas, A. (2024). Dimension-agnostic inference using cross U-statistics. Bernoulli, 30(1):683–711

  37. [37]

    Kim, I., Ramdas, A., Singh, A., and Wasserman, L. (2021). Classification accuracy as a proxy for two-sample testing.The Annals of Statistics, 49(1):411–434

  38. [38]

    and Schrab, A

    Kim, I. and Schrab, A. (2026). Differentially Private Permutation Tests.Journal of the Amer- ican Statistical Association, pages 1–13

  39. [39]

    and Romano, J

    Lehmann, E. and Romano, J. P. (2022). Testing Statistical Hypotheses. Springer Texts in Statistics. Springer, 4th edition. 29

  40. [40]

    J., and Wasserman, L

    Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J., and Wasserman, L. (2018). Distribution-free predictiveinferenceforregression. Journal of the American Statistical Association, 113(523):1094– 1111

  41. [41]

    Liu, F., Xu, W., Lu, J., Zhang, G., Gretton, A., and Sutherland, D. J. (2020). Learning deep kernels for non-parametric two-sample tests. InInternational Conference on Machine Learning, pages 6316–6326

  42. [42]

    R., Kim, I., Shah, R

    Lundborg, A. R., Kim, I., Shah, R. D., and Samworth, R. J. (2024). The projected covariance measure for assumption-lean variable significance testing.The Annals of Statistics, 52(6):2851– 2878

  43. [43]

    H., and Bühlmann, P

    Meinshausen, N., Maathuis, M. H., and Bühlmann, P. (2011). Asymptotic optimality of the westfall–young permutation procedure for multiple testing under dependence. The Annals of Statistics, 39(6):3369–3391

  44. [44]

    Meng, X.-L. (1994). Posterior predictivep-values. The Annals of Statistics, 22(3):1142–1160

  45. [45]

    Moran, P. A. (1973). Dividing a sample into two parts a statistical dilemma.Sankhy¯ a: The Indian Journal of Statistics, Series A, pages 329–333

  46. [46]

    Mun, J., Kwak, S., and Kim, I. (2025). Minimax optimal two-sample testing under local differential privacy.Journal of Machine Learning Research, 26(252):1–79

  47. [47]

    Paik, S., Celentano, M., Green, A., and Tibshirani, R. J. (2025). Integral Probability Metrics Meet Neural Networks: The Radon-Kolmogorov-Smirnov Test. Journal of Machine Learning Research, 26(86):1–57

  48. [48]

    Papadopoulos, H. (2008). Inductive conformal prediction: Theory and application to neural networks. INTECH Open Access Publisher Rijeka

  49. [49]

    and Salmaso, L

    Pesarin, F. and Salmaso, L. (2010).Permutation Tests for Complex Data: Theory, Applications and Software. Wiley Series in Probability and Statistics. John Wiley & Sons

  50. [50]

    Pitman, E. J. (1937). Significance tests which may be applied to samples from any populations. Supplement to the Journal of the Royal Statistical Society, 4(1):119–130

  51. [51]

    J., and Gretton, A

    Pogodin, R., Schrab, A., Li, Y., Sutherland, D. J., and Gretton, A. (2024). Practical Kernel Tests of Conditional Independence.arXiv preprint arXiv:2402.13196

  52. [52]

    F., Candès, E

    Ramdas, A., Barber, R. F., Candès, E. J., and Tibshirani, R. J. (2023). Permutation tests using arbitrary permutation distributions.Sankhya A, 85(2):1156–1177

  53. [53]

    and Wang, R

    Ramdas, A. and Wang, R. (2025). Hypothesis testing with E-values.Foundations and Trends® in Statistics, 1(1-2):1–390

  54. [54]

    Ribero, M., Schrab, A., and Gretton, A. (2026). Regularizedf-divergence kernel tests. InThe 29th International Conference on Artificial Intelligence and Statistics

  55. [55]

    Romano, J. P. and Wolf, M. (2005). Exact and approximate stepdown methods for multiple hypothesis testing. Journal of the American Statistical Association, 100(469):94–108. 30

  56. [56]

    Rüschendorf, L. (1982). Random Variables with Maximum Sums.Advances in Applied Proba- bility, 14(3):623–632

  57. [57]

    Rüger, B. (1978). Das maximale Signifikanzniveau des Tests: „LehneH0 ab, wenn k unter n gegebenen Tests zur Ablehnung führen.".Metrika, 25:171–178

  58. [58]

    Schrab, A. (2025a). A practical introduction to kernel discrepancies: MMD, HSIC & KSD. arXiv preprint arXiv:2503.04820

  59. [59]

    (2025b).Optimal Kernel Hypothesis Testing

    Schrab, A. (2025b).Optimal Kernel Hypothesis Testing. PhD thesis, UCL (University College London)

  60. [60]

    Schrab, A. (2025c). A unified view of optimal kernel hypothesis testing. arXiv preprint arXiv:2503.07084

  61. [61]

    Schrab, A., Guedj, B., and Gretton, A. (2022a). KSD Aggregated Goodness-of-fit Test. InAd- vances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022

  62. [62]

    Schrab, A., Kim, I., Albert, M., Laurent, B., Guedj, B., and Gretton, A. (2023). MMD aggregated two-sample test.Journal of Machine Learning Research, 24(194):1–81

  63. [63]

    Schrab, A., Kim, I., Guedj, B., and Gretton, A. (2022b). Efficient aggregated kernel tests using incomplete U-statistics. Advances in Neural Information Processing Systems, 35:18793–18807

  64. [64]

    Shah, R. D. and Bühlmann, P. (2018). Goodness-of-fit tests for high dimensional linear models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 80(1):113–135

  65. [65]

    Shekhar, S., Kim, I., and Ramdas, A. (2022). A permutation-free kernel two-sample test. Advances in Neural Information Processing Systems, 35:18168–18180

  66. [66]

    Shekhar, S., Kim, I., and Ramdas, A. (2023). A permutation-free kernel independence test. Journal of Machine Learning Research, 24(369):1–68

  67. [67]

    and Ramdas, A

    Shekhar, S. and Ramdas, A. (2024). Nonparametric two-sample testing by betting. IEEE Transactions on Information Theory, 70(2):1178–1203

  68. [68]

    and Onghena, P

    Solmi, F. and Onghena, P. (2014). Combining p-values in replicated single-case experiments with multivariate outcome.Neuropsychological Rehabilitation, 24(3-4):607–633

  69. [69]

    Stouffer, S.A., Suchman, E.A., DeVinney, L.C., Star, S.A., andWilliamsJr, R.M.(1949).The American Soldier: Adjustment during Army Life (Vol. 1). Princeton University Press, Princeton, NJ

  70. [70]

    Tansey, W., Veitch, V., Zhang, H., Rabadan, R., and Blei, D. M. (2022). The holdout random- ization test for feature selection in black box models.Journal of Computational and Graphical Statistics, 31(1):151–162

  71. [71]

    (2005).Algorithmic Learning in a Random World

    Vovk, V., Gammerman, A., and Shafer, G. (2005).Algorithmic Learning in a Random World. Springer. 31

  72. [72]

    Vovk, V., Wang, B., and Wang, R. (2022). Admissible ways of merging p-values under arbitrary dependence. The Annals of Statistics, 50(1):351–375

  73. [73]

    and Wang, R

    Vovk, V. and Wang, R. (2020). Combining p-values via averaging.Biometrika, 107(4):791–808

  74. [74]

    and Wang, R

    Vovk, V. and Wang, R. (2021). E-values: Calibration, combination, and applications. The Annals of Statistics, 49(3):1736–1754

  75. [75]

    Westfall, P. H. and Young, S. S. (1993). Resampling-based multiple testing: Examples and methods for p-value adjustment. John Wiley & Sons

  76. [76]

    and Kuchibhotla, A

    Yang, Y. and Kuchibhotla, A. K. (2025). Selection and Aggregation of Conformal Prediction Sets. Journal of the American Statistical Association, 120(549):435–447

  77. [77]

    qDf6pzroMMJvMcIkHWdR/tYPv+o=

    Zhou, Z., Tian, X., Peng, L., Lei, C., Schrab, A., Sutherland, D. J., and Liu, F. (2025). Dual: Learning diverse kernels for aggregated two-sample and independence testing. InAdvances in Neural Information Processing Systems, volume 38. 32 Supplementary material for Aggregation of Statistical Evidence under Exchangeability (a) SB aggregation X g1X g2X gBX...

  78. [78]

    qDf6pzroMMJvMcIkHWdR/tYPv+o=

    Using the equivalent formulations explained in Section 2.2, the decision rule in (2) can be written as min k∈[K] p(T k 0 ) ≤ ˜uα. Hence the MaxT procedure can be interpreted as a minimum p-value test with a Monte Carlo- calibrated correction factor. This viewpoint makes the comparison with SB and TB minimum 35 aggregation transparent: all three procedures...

  79. [79]

    By contrast, TB calibration p-values can still hit the smallest grid point with non-vanishing probability, creating ties that obstruct rejection under a strict comparison

    When B is small, this distinction can be decisive: under a strong signal, the inclusion ofT k 0 systematically inflates the SB calibration p-values away from the smallest grid 56 point 1/(B + 1), thereby increasing the SB critical value and facilitating rejection. By contrast, TB calibration p-values can still hit the smallest grid point with non-vanishin...

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.