REVIEW 2 major objections 2 minor
Hash-augmented adaptive multilevel splitting Monte Carlo accurately estimates arbitrarily small two-sample permutation p-values with valid confidence intervals.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 02:50 UTC pith:OQ4Q74GI
load-bearing objection Solid applied Monte Carlo tooling for small permutation p-values; abstract-only so the generalization claim is still un-audited. the 2 major comments →
Hash-augmented adaptive multilevel splitting Monte Carlo algorithm for accurate estimation of two-sample permutation test p-values
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A hash-augmented adaptive multilevel splitting Monte Carlo algorithm yields accurate estimates of arbitrarily small p-values for two-sample permutation tests, together with valid confidence intervals, even when the test statistic is discrete; the procedure is validated on Kolmogorov–Smirnov and Mann–Whitney U statistics against an exact reference and is available for arbitrary user-defined statistics via the hamstest package.
What carries the argument
Hash-augmented adaptive multilevel splitting: successive levels of rare-event sampling are combined with a hash function that correctly accounts for the discrete support of the permutation distribution, thereby producing unbiased p-value estimates and properly calibrated confidence intervals.
Load-bearing premise
That the same hash-augmentation and multilevel-splitting corrections remain unbiased for completely general user-defined statistics, not only for the two classical tests that were checked against an exact algorithm.
What would settle it
Run the algorithm on a custom discrete statistic whose exact permutation p-value is known by exhaustive enumeration; if the reported point estimate falls systematically outside the claimed confidence interval, or the empirical coverage of the intervals is far from nominal, the central claim fails.
If this is right
- Arbitrarily small permutation p-values can be estimated without exhaustive enumeration, removing a common barrier to multiple-testing correction.
- Any user-defined two-sample statistic can be analysed with the same Monte-Carlo guarantees via the hamstest package.
- Discreteness-induced bias that previously plagued small-p Monte-Carlo estimates is systematically eliminated.
- Confidence intervals supplied by the method can be used directly for rigorous significance statements.
Where Pith is reading between the lines
- The same multilevel-splitting-plus-hash idea could be transferred to one-sample, multi-sample or paired permutation tests with only modest redesign of the proposal kernel.
- Because the method scales to extreme rarity, it may become the default engine for genome-wide or high-throughput screening pipelines that currently rely on asymptotic approximations.
- Empirical coverage studies on a broader suite of discontinuous statistics would strengthen the claim of generality beyond KS and Mann–Whitney.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a hash-augmented adaptive multilevel splitting Monte Carlo algorithm for estimating arbitrarily small p-values in two-sample permutation tests. It highlights pitfalls arising from discreteness of the test-statistic distribution (illustrated on Kolmogorov–Smirnov and Mann–Whitney U), claims that the proposed remedy yields accurate point estimates and valid confidence intervals by comparison with an exact algorithm, and supplies a Python package (hamstest) that accepts user-defined statistics.
Significance. If the claims hold, the work would supply a practical, general-purpose tool for reliable small-p-value estimation in nonparametric two-sample settings where exact enumeration is infeasible—directly relevant to multiple-testing correction and custom test statistics. Explicit treatment of discreteness and release of an open-source package are concrete strengths that would increase utility to practitioners, provided the accuracy and CI validity claims are substantiated beyond the two named examples.
major comments (2)
- [Abstract] The abstract asserts accuracy of p-value estimates and validity of associated confidence intervals by comparison with an exact algorithm for the Kolmogorov–Smirnov and Mann–Whitney U tests. With only the abstract available, no theorems, error analyses, simulation tables, or proofs can be examined; the central empirical claim is therefore unverifiable from the material in hand and remains load-bearing for acceptance.
- [Abstract] The claim that the method extends to arbitrary user-defined statistics rests on the premise that hash augmentation correctly resolves discreteness of the test-statistic distribution without distorting the target tail probability or the confidence intervals. The abstract reports validation only for the two named tests; generalization is therefore unsupported by the available text and constitutes a load-bearing assumption that requires either theoretical guarantees or broader empirical evidence.
minor comments (2)
- [Abstract] The abstract alludes to free parameters of the adaptive multilevel splitting schedule and of the hash function but does not indicate whether sensitivity analyses, default recommendations, or diagnostics for users of hamstest are supplied.
- [Abstract] No references to prior multilevel-splitting or rare-event Monte Carlo literature appear in the abstract; the full manuscript should situate the hash-augmentation contribution relative to existing AMS variants.
Circularity Check
No significant circularity; algorithmic Monte Carlo estimator validated against external exact algorithms
full rationale
The abstract describes a hash-augmented adaptive multilevel splitting Monte Carlo procedure for estimating small two-sample permutation-test p-values, with confidence intervals, and reports empirical accuracy by direct comparison to an exact algorithm on the Kolmogorov–Smirnov and Mann–Whitney U statistics. The method is offered for arbitrary user-defined statistics via the hamstest package. Nothing in the available text indicates that any reported accuracy, confidence-interval coverage, or p-value estimate is obtained by fitting a parameter to the same quantity later called a prediction, by self-definition of the estimator in terms of the target p-value, or by a load-bearing self-citation that merely renames a prior ansatz. The only residual caveat (handling of discreteness for fully general custom statistics) is an empirical-scope limitation, not a circular derivation. With only the abstract available, no equation-level reduction of output to input can be exhibited; the paper is therefore scored as free of circularity.
Axiom & Free-Parameter Ledger
free parameters (2)
- adaptive multilevel splitting schedule / intermediate levels
- hash function / hash-augmentation parameters
axioms (3)
- domain assumption Two-sample permutation tests under exchangeability yield valid p-values when the full permutation distribution is used.
- domain assumption Adaptive multilevel splitting produces unbiased or consistently biased-controllable rare-event probability estimates under stated Markov-chain and level conditions.
- ad hoc to paper Hash augmentation correctly resolves discreteness of the test-statistic distribution without distorting the target tail probability.
read the original abstract
Nonparametric permutation tests are widely used for statistical analysis. However, exact computation of test p-values can be algorithmically challenging, particularly for custom tests with complex test statistics. In contrast, Monte Carlo sampling can be easily applied to any test statistic, but it suffers from poor relative accuracy when estimating small p-values, interfering with multiple hypothesis testing correction and leading to other issues. In this work, we present a hash-augmented adaptive multilevel splitting Monte Carlo algorithm that enables accurate estimation of arbitrarily small p-values in two-sample permutation tests. Using the Kolmogorov-Smirnov and the Mann-Whitney U tests as examples, we highlight potential pitfalls related to the discreteness of the test statistic distribution and show how to address them. By comparing with an exact algorithm, we demonstrate the accuracy of the p-value estimates provided by the proposed algorithm and the validity of the associated confidence intervals. We provide a reference implementation of the proposed algorithm in the Python package hamstest, which allows p-value estimation for a user-defined statistic.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.