REVIEW 5 minor 1 cited by
Calibrating aggregated test statistics on the permutation distribution itself is finite-sample valid and at least as powerful as Bonferroni-type worst-case calibration, with sequential and two-batch extensions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Calibrating aggregated test statistics on the permutation distribution itself is finite-sample valid and at least as powerful as Bonferroni-type worst-case calibration, with sequential and two-batch extensions.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection A clean, correct paper on permutation-based aggregation with a genuine dominance theorem; the main caveat is the exact-exchangeability assumption, which the authors state plainly.
Aggregation of Statistical Evidence under Exchangeability
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
This paper works with a construction that is more careful. It applies all K statistics to the original data and to B randomly shuffled copies. Within each copy, every statistic is converted to a permutation p-value, and the K p-values are merged with a function of the user's choice (minimum, mean, median, anything). Because every copy is processed in exactly the same way, the B+1 merged scores are exchangeable under the null hypothesis: the original data is just one of B+1 symmetric siblings. Ranking the original score among its shuffled siblings therefore gives an exact, finite-sample p-value (Theorem 1).
The central power result (Theorem 2) is a dominance argument: any threshold that is valid for every possible dependence pattern must in particular be valid for the specific pattern displayed by the permutation copies, so the data-calibrated threshold is always at least as large as the worst-case one. The paper then extends the idea in three ways: a sequential version that spends the significance budget across ordered stages and can stop early; a two-batch version that standardizes on one batch of shuffles and calibrates on another, allowing the merging rule itself to be learned; and a conformal-prediction application where, for absolute-residual scores with minimum merging, the merged prediction set is an expli
Core claim
Theorem 2: 'Suppose that c_α,K is a deterministic constant satisfying sup_ν P_ν(f(U_1,...,U_K) ≤ c_α,K) ≤ α. Then it holds almost surely that c_α,K < û^SB_α. Consequently, ... P(f_0 ≥ û^SB_α) ≤ P(f_0 > c_α,K).' Together with Theorem 1 (P(p_SB ≤ α) ≤ α for all α, B, any merge function f), the load-bearing assertion is that single-batch permutation-calibrated aggregation is finite-sample valid under the group-invariance null and at least as powerful as any deterministic calibration valid under arbitrary dependence (Bonferroni, O/M-family merging), while adapting to the realized dependence (Propositions 2-3). If correct, SB/TB aggregation can replace worst-case calibration in adaptive tests and conformal merging without losing level control.
Load-bearing premise
Exact joint exchangeability of the transformed rows under the null: the transformations g_1,...,g_B must be chosen independently of the data (i.i.d. uniform on G or without-replacement sampling) so that the row vectors (T^1_b,...,T^K_b) are exchangeable in b, and the standardization, merge, and tie-breaking maps must be row-wise equivariant. Location: Section 2.1 (Definition 1 and 'Unless stated otherwise, we assume that the transformations are chosen so that the transformed data g_0(X),...,g_B(X) are exchangeable'); Section 3.2 (row-wise application of f); Section 5.3 (learning the rule on the testing batch 'would generally destroy the exchangeability'); Section 8 ('The current guarantees assume that the transformed rows are exactly exchangeable'). If the null holds only approximately, or transformations are chosen with the data, every finite-sample guarantee in Theorems 1-3 and Proposition 5 is void.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies aggregation of multiple test statistics under the group-invariance (randomization) null. In single-batch (SB) aggregation, the same B+1 transformed datasets are used to compute permutation p-values for each of K statistics, merge them row-wise through an arbitrary function f, and calibrate the merged value by its rank among all rows. Theorem 1 establishes finite-sample type I error control for any f. Theorem 2 shows that the SB critical value almost surely exceeds any deterministic threshold that is valid under arbitrary dependence among super-uniform p-values, yielding power dominance over worst-case calibrated rules such as O- and M-family merging (Corollary 1). Propositions 2 and 3 show adaptivity to perfect rank alignment and asymptotic oracle calibration. Section 4 adds a sequential alpha-spending version (Proposition 5), and Section 5 adds a two-batch (TB) construction that allows reference-batch-dependent aggregation rules while preserving validity (Theorem 3). Applications to adaptive nonparametric testing and conformal prediction are given, including a closed-form interval intersection for minimum-merging TB in Proposition 6. Numerical experiments illustrate finite-sample power, level, and computational gains.
Significance. If the results hold, they provide a clean and useful unification: exact finite-sample validity for essentially arbitrary merging functions, a simple argument showing uniform dominance over deterministic worst-case calibrations, and explicit dependence adaptivity. The paper's strengths include elementary, checkable proofs, no fitted constants or circular calibration, explicit credit to the prior permutation-combination literature, and honest discussion of limitations in Section 8. The main scope caveat is the exact-exchangeability assumption on the transformed rows: without it, Theorems 1–3 and Proposition 5 have no finite-sample content. The paper states this condition in Section 2.1 and acknowledges the approximate-exchangeability gap in Section 8, and the primary applications (permutation and split-conformal methods) satisfy the condition. I therefore regard the assumption as a stated scope limitation rather than an internal inconsistency. The empirical work supports, rather than establishes, the theoretical claims, which is appropriate here.
minor comments (5)
- [Sections 2.1 and 8] The exact-exchangeability condition on the transformed rows is the single most important scope condition. Since Section 8 correctly notes that approximate exchangeability is open, I suggest adding one sentence in the abstract or introduction making explicit that all finite-sample guarantees require exact exchangeability, so that the phrase 'potentially complex dependence' is not over-read as relaxing this structural assumption.
- [Section 3.3, Theorem 2] The 'almost surely' qualifier in c_{α,K} < û^SB_α appears unnecessary: Lemma 1 is deterministic for any realized statistic matrix, so the inequality holds for every realization. Making this pointwise nature explicit would simplify the statement and clarify that no additional probabilistic regularity is needed.
- [Sections 3.5 and Proposition 1] The strict inequality in the SB minimum rejection rule is essential for the 1/(B+1) shift in Proposition 4. Since the threshold definition in Eq. (6) uses a supremum over {F(u) ≤ α}, the open/closed half-line distinction is easy to misread. A short remark after Proposition 1 explaining the strict-vs-non-strict convention would prevent errors.
- [Figures 1 and 2] The figure captions are compressed. In Figure 1, the left panel uses L1, L4, L∞ while the right panel labels L1, L2, L3, and the 'SB SB ties TB TB ties' entries do not map unambiguously to line types. In Figure 2, the decision-stage legend is unclear. Expanding the captions or adding explicit legends would materially improve readability.
- [Section 3.2, Eq. (4)] The permutation p-value definition uses the non-strict inequality 1(T^i_k ≥ T^b_k). This is standard, but the paper later emphasizes the importance of tie-breaking; connecting Eq. (4) explicitly to the tie-breaking discussion in Appendix B.6 in the main text would help readers understand why the conservative definition is used before tie-breaking is introduced.
Circularity Check
No significant circularity; the validity and dominance theorems are proved from first principles, and self-citations serve as comparison baselines rather than load-bearing support.
full rationale
The paper's central claims are derived within the manuscript. Theorem 1 and Theorem 3 rest on the standard exchangeability rank argument (row vectors (T^1_b,...,T^K_b) are exchangeable under the group-invariance null, so the rank statistic p_SB is super-uniform), stated and justified in Sections 3.2 and 5.2, with proofs in the supplement. Theorem 2 is not an identity: the SB threshold û^SB_α is defined directly from the realized merged values (Proposition 1), while the deterministic constant c_{α,K} is defined by a supremum over all joint laws with super-uniform margins; the link between them is Lemma 1, a deterministic count bound on empirical super-uniformity of permutation p-values, from which the feasibility of c for the empirical constraint and the strict inequality c < û follow. This is a genuine two-step argument, not a re-statement of an input. The sequential (Proposition 5) and TB results likewise reduce to row-removal counting and conditional exchangeability. No parameter is fitted to data and then reported as a prediction; thresholds are internally calibrated. Self-citations ([26], [61]–[63], [34], [38], [77], etc.) are used as historical attribution ([47], [49], [64], [75]) or as comparison baselines (MaxT with Monte Carlo calibration, rank-transformed subsampling), and are not the foundation of any load-bearing theorem; the alleged liberality of MaxT is proved in-appendix (Proposition S.8), not imported. No uniqueness theorem, no ansatz-smuggling citation, and no renaming of a known result is performed: Section 3.5 explicitly acknowledges that minimum merging ``recovers the Westfall--Young single-step method.'' The paper's own stated limitations (Section 8: exact exchangeability of transformed rows; Section 5.3: learning the aggregation rule on the calibration batch would destroy exchangeability) are honest scope conditions, not circular steps. Under the stated exact-exchangeability assumption, the derivation chain is self-contained, so the circularity score is minimal.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption Group-invariance hypothesis (Definition 1): under H0, g(X) has the same distribution as X for every g ∈ G.
- domain assumption The transformations g_1,...,g_B are drawn independently of X and jointly so that g_0(X),...,g_B(X) are exchangeable (i.i.d. uniform on G, or without-replacement sampling).
- domain assumption The standardization map (permutation p-values), the merging function f, and any tie-breaking augmentation are applied row-wise and are equivariant under row permutations.
- domain assumption TB/conformal: conditional exchangeability of testing-batch rows (or aggregation-batch scores plus test point) given the reference batch; the aggregation rule may depend only on reference-batch information.
- standard math Rank-p-value super-uniformity of exchangeable tuples (Lemma 1 and its generalization Lemma S.8, extending Harrison 2012, Lemma A1).
- domain assumption Asymptotic regime condition for Proposition 3: (f_{1,n}, f_{2,n}) converge jointly to i.i.d. limits with a well-separated oracle quantile Q*_α.
Cite this review
Pith. "Pith review of Aggregation of Statistical Evidence under Exchangeability." pith.science (2026). https://pith.science/paper/ODIEGRT3
@misc{pith2026260715823,
author = {Pith},
title = {Pith review of: Aggregation of Statistical Evidence under Exchangeability},
year = {2026},
howpublished = {\url{https://pith.science/paper/ODIEGRT3}},
note = {Machine review of arXiv:2607.15823}
}
read the original abstract
We study aggregation of statistical evidence under unknown and potentially complex dependence using group-invariance. Building on permutation-based constructions that treat transformed datasets as exchangeable units, we aggregate evidence across statistics for each transformed dataset and calibrate the resulting aggregates across transformations. We develop a finite-sample power and adaptivity theory for this framework, together with extensions to sequential and data-dependent aggregation that preserve validity. For single-batch aggregation, which uses one collection of transformed datasets for both standardization and calibration, we show that the critical values uniformly improve on deterministic calibrations valid under arbitrary dependence, including Bonferroni correction, while adapting to the unknown dependence structure. We also introduce a sequential alpha-spending version that permits early rejection when evidence is strong, and a two-batch extension that separates standardization from calibration to accommodate learned aggregation rules and reduce computation. Applications to adaptive nonparametric testing and conformal prediction illustrate how these results sharpen existing aggregation methods.
Figures
Forward citations
Cited by 1 Pith paper
-
Sharp Minimax Rates for Smooth Two-Sample Testing under Central Differential Privacy
Under central differential privacy, the sharp L1 separation radius for two-sample testing of Hölder-smooth densities is the maximum of the classical rate and three privacy barriers, and adapting to unknown smoothness ...
Reference graph
Works this paper leans on
-
[1]
(2015).Tests of independence by bootstrap and permutation: an asymptotic and non- asymptotic study
Albert, M. (2015).Tests of independence by bootstrap and permutation: an asymptotic and non- asymptotic study. Application to neurosciences.PhD thesis, Université Nice Sophia Antipolis
2015
-
[2]
Albert, M., Laurent, B., Marrel, A., and Meynaoui, A. (2022). Adaptive test of independence based on HSIC measures.The Annals of Statistics, 50(2):858–879
2022
-
[3]
Angelopoulos, A. N., Barber, R. F., and Bates, S. (2024). Theoretical Foundations of Conformal Prediction. arXiv preprint arXiv:2411.11824
Pith/arXiv arXiv 2024
-
[4]
Baraud, Y., Huet, S., and Laurent, B. (2003). Adaptive tests of linear hypotheses by model selection. The Annals of Statistics, 31(1):225–251
2003
-
[5]
B., Kontoyiannis, I., and Samworth, R
Berrett, T. B., Kontoyiannis, I., and Samworth, R. J. (2021). Optimal rates for independence testing via U-statistic permutation tests.The Annals of Statistics, 49(5):2457–2490
2021
-
[6]
B., Wang, Y., Barber, R
Berrett, T. B., Wang, Y., Barber, R. F., and Samworth, R. J. (2020). The conditional permu- tation test for independence while controlling for confounders.Journal of the Royal Statistical Society Series B: Statistical Methodology, 82(1):175–197. 27
2020
-
[7]
Biggs, F., Schrab, A., and Gretton, A. (2023). MMD-FUSE: Learning and combining kernels for two-sample testing without data splitting.Advances in Neural Information Processing Systems, 36
2023
-
[8]
Candes, E., Fan, Y., Janson, L., and Lv, J. (2018). Panning for gold: ‘Model-X’ knockoffs for high dimensional controlled variable selection.Journal of the Royal Statistical Society Series B: Statistical Methodology, 80(3):551–577
2018
-
[9]
Caughey, D., Dafoe, A., and Seawright, J. (2017). Nonparametric combination (NPC): A frame- work for testing elaborate theories.The Journal of Politics, 79(2):688–701
2017
-
[10]
Cha, S., Lee, S., Schrab, A., and Kim, I. (2026). More Permutations Do Not Always Increase Power: Non-monotonicity in Monte Carlo Permutation Tests.arXiv preprint arXiv:2605.03886
Pith/arXiv arXiv 2026
-
[11]
L., Schrab, A., Gretton, A., Sejdinovic, D., and Muandet, K
Chau, S. L., Schrab, A., Gretton, A., Sejdinovic, D., and Muandet, K. (2025). Credal two- sample tests of epistemic uncertainty. InProceedings of The 28th International Conference on Artificial Intelligence and Statistics, volume 258 ofProceedings of Machine Learning Research, pages 127–135. PMLR
2025
-
[12]
and Kim, I
Choi, W. and Kim, I. (2023). Averaging p-values under exchangeability.Statistics & Probability Letters, 194:109748
2023
-
[13]
and Romano, J
Chung, E. and Romano, J. P. (2013). Exact and asymptotically robust permutation tests.The Annals of Statistics, 41(2):484–507
2013
-
[14]
Cox, D. R. (1975). A note on data-splitting for the evaluation of significance levels.Biometrika, 62(2):441–444
1975
-
[15]
Domingo-Enrich, C., Dwivedi, R., and Mackey, L. (2025). Cheap permutation testing.arXiv preprint arXiv:2502.07672
Pith/arXiv arXiv 2025
-
[16]
Fisher, R. A. (1925).Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh
1925
-
[17]
Fisher, R. A. (1935).The Design of Experiments. Oliver and Boyd, Edinburgh
1935
-
[18]
and Laurent, B
Fromont, M. and Laurent, B. (2006). Adaptive goodness-of-fit tests in a density model.The Annals of Statistics, 34(2):680–720
2006
-
[19]
Fromont, M., Laurent, B., and Reynaud-Bouret, P. (2013). The two-sample problem for pois- son processes: Adaptive tests with a nonasymptotic wild bootstrap approach. The Annals of Statistics, 41(3):1431–1461
2013
-
[20]
Gasparin, M. and Ramdas, A. (2024). Merging uncertainty sets via majority vote. arXiv preprint arXiv:2401.09379
Pith/arXiv arXiv 2024
-
[21]
Gasparin, M., Wang, R., and Ramdas, A. (2025). Combining exchangeable p-values.Proceed- ings of the National Academy of Sciences, 122(11):e2410849122
2025
-
[22]
(2005).Permutation, parametric and bootstrap tests of hypotheses
Good, P. (2005).Permutation, parametric and bootstrap tests of hypotheses. Springer
2005
-
[23]
Gretton, A. (2015). A simpler condition for consistency of a kernel independence test.arXiv preprint arXiv:1501.06103. 28
Pith/arXiv arXiv 2015
-
[24]
M., Rasch, M
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A. (2012). A kernel two-sample test.Journal of Machine Learning Research, 13(25):723–773
2012
-
[25]
Gretton, A., Herbrich, R., Smola, A., Bousquet, O., and Schölkopf, B. (2005). Kernel methods for measuring independence.Journal of Machine Learning Research, 6:2075–2129
2005
-
[26]
Guo, F. R. and Shah, R. D. (2025). Rank-transformed subsampling: inference for multiple data splitting and exchangeable p-values.Journal of the Royal Statistical Society Series B: Statistical Methodology, 87(1):256–286
2025
-
[27]
Hagrass, O., Sriperumbudur, B., and Li, B. (2024). Spectral regularized kernel two-sample tests. The Annals of Statistics, 52(3):1076–1101
2024
-
[28]
Conservativehypothesistestsandconfidenceintervalsusingimportance sampling
Harrison, M.T.(2012). Conservativehypothesistestsandconfidenceintervalsusingimportance sampling. Biometrika, 99(1):57–69
2012
-
[29]
I., and Dieuleveut, A
Hegazy, M., Aolaritei, L., Jordan, M. I., and Dieuleveut, A. (2025). Valid selection among conformal sets. InAdvances in Neural Information Processing Systems, volume 38
2025
-
[30]
and Goeman, J
Hemerik, J. and Goeman, J. (2018). Exact testing with random permutations.Test, 27(4):811– 825
2018
-
[31]
and Suslina, I
Ingster, Y. and Suslina, I. A. (2012). Nonparametric goodness-of-fit testing under Gaussian models, volume 169. Springer Science & Business Media
2012
-
[32]
D., Bühlmann, P., and Samworth, R
Janková, J., Shah, R. D., Bühlmann, P., and Samworth, R. J. (2020). Goodness-of-fit testing in high dimensional generalized linear models.Journal of the Royal Statistical Society Series B: Statistical Methodology, 82(3):773–795
2020
-
[33]
B., and Yu, Y
Kent, A., Berrett, T. B., and Yu, Y. (2026). Locally Differentially Private Two-Sample Testing. Biometrika, page asag034
2026
-
[34]
Kim, I., Balakrishnan, S., and Wasserman, L. (2022). Minimax optimality of permutation tests. The Annals of Statistics, 50(1):225–251
2022
-
[35]
Kim, I., Neykov, M., Balakrishnan, S., and Wasserman, L. (2024). Conditional indepen- dence testing for discrete distributions: Beyondχ2- and G-tests.Electronic Journal of Statistics, 18(2):4767–4794
2024
-
[36]
and Ramdas, A
Kim, I. and Ramdas, A. (2024). Dimension-agnostic inference using cross U-statistics. Bernoulli, 30(1):683–711
2024
-
[37]
Kim, I., Ramdas, A., Singh, A., and Wasserman, L. (2021). Classification accuracy as a proxy for two-sample testing.The Annals of Statistics, 49(1):411–434
2021
-
[38]
and Schrab, A
Kim, I. and Schrab, A. (2026). Differentially Private Permutation Tests.Journal of the Amer- ican Statistical Association, pages 1–13
2026
-
[39]
and Romano, J
Lehmann, E. and Romano, J. P. (2022). Testing Statistical Hypotheses. Springer Texts in Statistics. Springer, 4th edition. 29
2022
-
[40]
J., and Wasserman, L
Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J., and Wasserman, L. (2018). Distribution-free predictiveinferenceforregression. Journal of the American Statistical Association, 113(523):1094– 1111
2018
-
[41]
Liu, F., Xu, W., Lu, J., Zhang, G., Gretton, A., and Sutherland, D. J. (2020). Learning deep kernels for non-parametric two-sample tests. InInternational Conference on Machine Learning, pages 6316–6326
2020
-
[42]
R., Kim, I., Shah, R
Lundborg, A. R., Kim, I., Shah, R. D., and Samworth, R. J. (2024). The projected covariance measure for assumption-lean variable significance testing.The Annals of Statistics, 52(6):2851– 2878
2024
-
[43]
H., and Bühlmann, P
Meinshausen, N., Maathuis, M. H., and Bühlmann, P. (2011). Asymptotic optimality of the westfall–young permutation procedure for multiple testing under dependence. The Annals of Statistics, 39(6):3369–3391
2011
-
[44]
Meng, X.-L. (1994). Posterior predictivep-values. The Annals of Statistics, 22(3):1142–1160
1994
-
[45]
Moran, P. A. (1973). Dividing a sample into two parts a statistical dilemma.Sankhy¯ a: The Indian Journal of Statistics, Series A, pages 329–333
1973
-
[46]
Mun, J., Kwak, S., and Kim, I. (2025). Minimax optimal two-sample testing under local differential privacy.Journal of Machine Learning Research, 26(252):1–79
2025
-
[47]
Paik, S., Celentano, M., Green, A., and Tibshirani, R. J. (2025). Integral Probability Metrics Meet Neural Networks: The Radon-Kolmogorov-Smirnov Test. Journal of Machine Learning Research, 26(86):1–57
2025
-
[48]
Papadopoulos, H. (2008). Inductive conformal prediction: Theory and application to neural networks. INTECH Open Access Publisher Rijeka
2008
-
[49]
and Salmaso, L
Pesarin, F. and Salmaso, L. (2010).Permutation Tests for Complex Data: Theory, Applications and Software. Wiley Series in Probability and Statistics. John Wiley & Sons
2010
-
[50]
Pitman, E. J. (1937). Significance tests which may be applied to samples from any populations. Supplement to the Journal of the Royal Statistical Society, 4(1):119–130
1937
-
[51]
Pogodin, R., Schrab, A., Li, Y., Sutherland, D. J., and Gretton, A. (2024). Practical Kernel Tests of Conditional Independence.arXiv preprint arXiv:2402.13196
arXiv 2024
-
[52]
F., Candès, E
Ramdas, A., Barber, R. F., Candès, E. J., and Tibshirani, R. J. (2023). Permutation tests using arbitrary permutation distributions.Sankhya A, 85(2):1156–1177
2023
-
[53]
and Wang, R
Ramdas, A. and Wang, R. (2025). Hypothesis testing with E-values.Foundations and Trends® in Statistics, 1(1-2):1–390
2025
-
[54]
Ribero, M., Schrab, A., and Gretton, A. (2026). Regularizedf-divergence kernel tests. InThe 29th International Conference on Artificial Intelligence and Statistics
2026
-
[55]
Romano, J. P. and Wolf, M. (2005). Exact and approximate stepdown methods for multiple hypothesis testing. Journal of the American Statistical Association, 100(469):94–108. 30
2005
-
[56]
Rüschendorf, L. (1982). Random Variables with Maximum Sums.Advances in Applied Proba- bility, 14(3):623–632
1982
-
[57]
Rüger, B. (1978). Das maximale Signifikanzniveau des Tests: „LehneH0 ab, wenn k unter n gegebenen Tests zur Ablehnung führen.".Metrika, 25:171–178
1978
-
[58]
Schrab, A. (2025a). A practical introduction to kernel discrepancies: MMD, HSIC & KSD. arXiv preprint arXiv:2503.04820
-
[59]
(2025b).Optimal Kernel Hypothesis Testing
Schrab, A. (2025b).Optimal Kernel Hypothesis Testing. PhD thesis, UCL (University College London)
-
[60]
Schrab, A. (2025c). A unified view of optimal kernel hypothesis testing. arXiv preprint arXiv:2503.07084
-
[61]
Schrab, A., Guedj, B., and Gretton, A. (2022a). KSD Aggregated Goodness-of-fit Test. InAd- vances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022
2022
-
[62]
Schrab, A., Kim, I., Albert, M., Laurent, B., Guedj, B., and Gretton, A. (2023). MMD aggregated two-sample test.Journal of Machine Learning Research, 24(194):1–81
2023
-
[63]
Schrab, A., Kim, I., Guedj, B., and Gretton, A. (2022b). Efficient aggregated kernel tests using incomplete U-statistics. Advances in Neural Information Processing Systems, 35:18793–18807
-
[64]
Shah, R. D. and Bühlmann, P. (2018). Goodness-of-fit tests for high dimensional linear models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 80(1):113–135
2018
-
[65]
Shekhar, S., Kim, I., and Ramdas, A. (2022). A permutation-free kernel two-sample test. Advances in Neural Information Processing Systems, 35:18168–18180
2022
-
[66]
Shekhar, S., Kim, I., and Ramdas, A. (2023). A permutation-free kernel independence test. Journal of Machine Learning Research, 24(369):1–68
2023
-
[67]
and Ramdas, A
Shekhar, S. and Ramdas, A. (2024). Nonparametric two-sample testing by betting. IEEE Transactions on Information Theory, 70(2):1178–1203
2024
-
[68]
and Onghena, P
Solmi, F. and Onghena, P. (2014). Combining p-values in replicated single-case experiments with multivariate outcome.Neuropsychological Rehabilitation, 24(3-4):607–633
2014
-
[69]
Stouffer, S.A., Suchman, E.A., DeVinney, L.C., Star, S.A., andWilliamsJr, R.M.(1949).The American Soldier: Adjustment during Army Life (Vol. 1). Princeton University Press, Princeton, NJ
1949
-
[70]
Tansey, W., Veitch, V., Zhang, H., Rabadan, R., and Blei, D. M. (2022). The holdout random- ization test for feature selection in black box models.Journal of Computational and Graphical Statistics, 31(1):151–162
2022
-
[71]
(2005).Algorithmic Learning in a Random World
Vovk, V., Gammerman, A., and Shafer, G. (2005).Algorithmic Learning in a Random World. Springer. 31
2005
-
[72]
Vovk, V., Wang, B., and Wang, R. (2022). Admissible ways of merging p-values under arbitrary dependence. The Annals of Statistics, 50(1):351–375
2022
-
[73]
and Wang, R
Vovk, V. and Wang, R. (2020). Combining p-values via averaging.Biometrika, 107(4):791–808
2020
-
[74]
and Wang, R
Vovk, V. and Wang, R. (2021). E-values: Calibration, combination, and applications. The Annals of Statistics, 49(3):1736–1754
2021
-
[75]
Westfall, P. H. and Young, S. S. (1993). Resampling-based multiple testing: Examples and methods for p-value adjustment. John Wiley & Sons
1993
-
[76]
and Kuchibhotla, A
Yang, Y. and Kuchibhotla, A. K. (2025). Selection and Aggregation of Conformal Prediction Sets. Journal of the American Statistical Association, 120(549):435–447
2025
-
[77]
qDf6pzroMMJvMcIkHWdR/tYPv+o=
Zhou, Z., Tian, X., Peng, L., Lei, C., Schrab, A., Sutherland, D. J., and Liu, F. (2025). Dual: Learning diverse kernels for aggregated two-sample and independence testing. InAdvances in Neural Information Processing Systems, volume 38. 32 Supplementary material for Aggregation of Statistical Evidence under Exchangeability (a) SB aggregation X g1X g2X gBX...
2025
-
[78]
qDf6pzroMMJvMcIkHWdR/tYPv+o=
Using the equivalent formulations explained in Section 2.2, the decision rule in (2) can be written as min k∈[K] p(T k 0 ) ≤ ˜uα. Hence the MaxT procedure can be interpreted as a minimum p-value test with a Monte Carlo- calibrated correction factor. This viewpoint makes the comparison with SB and TB minimum 35 aggregation transparent: all three procedures...
2000
-
[79]
By contrast, TB calibration p-values can still hit the smallest grid point with non-vanishing probability, creating ties that obstruct rejection under a strict comparison
When B is small, this distinction can be decisive: under a strong signal, the inclusion ofT k 0 systematically inflates the SB calibration p-values away from the smallest grid 56 point 1/(B + 1), thereby increasing the SB critical value and facilitating rejection. By contrast, TB calibration p-values can still hit the smallest grid point with non-vanishin...
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.