Pith. sign in

REVIEW 2 major objections 5 minor 2 cited by

This paper establishes that Gaffke's interval for a bounded mean is first-order asymptotically efficient, while Gaffke's p-value, as an e-to-p merger, is inadmissible for every n≥2.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 14:45 UTC pith:SL72F7Y7

load-bearing objection Gaffke is genuinely new and mostly right, but the n>=3 CI inadmissibility claim outruns the proof; the fixed-vector inadmissibility and asymptotic efficiency results are the real contributions. the 2 major comments →

arxiv 2607.18661 v1 pith:SL72F7Y7 submitted 2026-07-21 math.ST cs.ITeess.SPmath.ITmath.PRstat.MEstat.TH

Gaffke's confidence interval for the mean of bounded data is inadmissible but asymptotically efficient

classification math.ST cs.ITeess.SPmath.ITmath.PRstat.MEstat.TH MSC 62F2562G0562G20
keywords bounded meanconfidence intervalDirichlet averagee-valuep-value mergingadmissibilityasymptotic efficiencyelementary symmetric polynomial
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks two separate questions about Gaffke's nonparametric procedure for testing and estimating a mean from independent nonnegative observations. As a confidence interval for data in [0,1], the procedure is excellent: it has distribution-free finite-sample coverage, reduces to the exact binomial interval for Bernoulli data, and its equal-tail width matches the oracle normal interval in the large-sample limit, with √n times the width converging almost surely to 2σz_{1−α/2}. As a function that converts independent e-values into a p-value, however, the same statistic is inadmissible for every n≥2: an explicit rule K_ad_2 strictly dominates it in two dimensions, and a neutral-face extension proves dominance in all higher dimensions, while a randomized improvement acts on the full upper orthant. The paper also proves the Gaffke p-value never exceeds the SymPol p-value, so within its validity domain it is uniformly at least as powerful. The reason to care is that the two findings do not conflict: asymptotic interval efficiency and finite-sample test admissibility are different criteria, and the paper makes that distinction precise.

Core claim

The central claim is a two-sided portrait of one object. Gaffke's statistic K_n(x)=P_D(Σ x_i D_i ≤ 1), with Dirichlet weights, underlies a bounded-mean confidence interval that is first-order asymptotically efficient: for iid [0,1] data with σ²>0, √n Width(I_n)→2σ z_{1−α/2} almost surely and coverage tends to the nominal level, while preserving finite-sample validity. The same statistic, viewed as an e-to-p merger, is inadmissible for every n≥2: K_ad_2 is the unique admissible dominator of K_2, and neutral-face embedding extends strict domination to all n≥2. The paper also proves K_n e_k ≤ C(n,k) for every elementary symmetric polynomial, so Gaffke pointwise dominates the SymPol p-value, and

What carries the argument

The load-bearing object is Gaffke's e-to-p merger K_n(x)=P_D(Σ x_i D_i ≤ 1) with (D_0,...,D_n)∼Dirichlet(1,...,1), a function that maps independent e-values (nonnegative variables with mean at most one) to a valid p-value. Supporting the argument are three identities: the elementary-symmetric bound K_n(x)e_k(x)≤C(n,k); the neutral-coordinate identity K_n(x_1,...,x_m,1,...,1)=K_m(x), which lets a two-input improvement be embedded in every dimension; and the explicit two-input dominator K_ad_2 whose level sets are governed by the larger root τ(a,b) of t²−b(1+a)t+ab=0. The asymptotic efficiency is carried by a conditional Lindeberg central limit theorem for uniform Dirichlet averages, yielding

Load-bearing premise

The whole argument depends on the externally supplied theorem that K_n(X) is a valid p-value whenever X_1,...,X_n are independent nonnegative variables with means at most one; the paper does not reprove that theorem, and a flaw there would invalidate the interval's coverage and the validity of every dominator, while the asymptotic efficiency statement also needs positive variance σ²>0.

What would settle it

Compute the equal-tail Gaffke interval for large n from iid Beta(20,20) data, a low-variance case: if at n=1000 the empirical coverage departs from 0.95 beyond Monte Carlo error, or if √n times the mean width exceeds 2σz_{1−α/2} by more than noise, the asymptotic efficiency theorem is wrong. Separately, search numerically for a valid e-to-p merger F with F(x)<K_ad_2(x) at some 0<a<1<b, for instance near the golden-ratio point (1/2,2); existence would refute the claimed uniqueness of the admissible dominator.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • At every nondegenerate iid distribution on [0,1], the equal-tail Gaffke interval is first-order equivalent to the normal interval, so practitioners get Gaussian efficiency with distribution-free finite-sample coverage.
  • The Gaffke p-value is pointwise no larger than the SymPol p-value, so under independence it is at least as powerful; the corresponding confidence interval is pointwise contained in the SymPol interval.
  • K_2 is not admissible: K_ad_2 is the unique admissible merger dominating it, and neutral-face embedding makes K_n inadmissible for every n≥2.
  • Allowing one independent uniform random variable yields a randomized p-value equal to U/∏x_i on the upper orthant, a strictly smaller and locally optimal improvement over Gaffke.
  • For Bernoulli samples the Gaffke interval is exactly the classical binomial interval, and its finite-sample coverage guarantee does not require identical distributions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If a globally coordinatewise monotone dominator of K_n exists, it would likely invert to a uniformly shorter interval; the neutral-face dominator does not move interval endpoints, so that search is the natural next step.
  • The interpretation of K_n as interpolating between the product e-value and a Sidak-like correction for the maximum suggests the abstract e-to-p merger is useful beyond bounded-mean problems, wherever independent e-values can be constructed for a composite null.
  • A high-precision comparison of Gaffke to the empirical Berry–Esseen method at low variance would likely shrink the reported width gap, since the paper's grid approximation is explicitly coarse and conservative.
  • The exact two-input dominator can in principle be inverted for n=2 to produce a finite-sample shorter interval; the paper does not pursue this, making it a concrete testable extension.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper studies Gaffke's e-to-p merger K_n and the confidence interval for a bounded mean obtained by inverting it. It proves that K_n pointwise dominates the SymPol p-value (Theorem 2.2), constructs an explicit admissible dominator K_2^ad for n=2 and shows it is the unique valid rule below K_2 (Theorems 3.6 and 3.7), extends inadmissibility of the e-to-p merger to all n via a neutral-face construction (Theorem 4.1), introduces a randomized improvement on the upper orthant (Theorem 4.13), and proves a conditional quantile CLT showing that the equal-tail interval has the oracle Gaussian width asymptotically (Theorem 6.2, Eq. (63)). Simulations compare the Gaffke interval with several existing procedures, reporting favorable finite-sample widths.

Significance. If the results hold, the paper makes several solid contributions: new deterministic inequalities for Gaffke's statistic (Theorem 2.2), a complete characterization of admissible two-input e-to-p mergers below K_2, a new proof of first-order asymptotic efficiency for the Gaffke interval, and extensive reproducible simulations. The authors are appropriately careful to separate test admissibility from interval performance and include a limitation section. However, the headline claim that the confidence interval is inadmissible is broader than what is actually proved, as detailed below.

major comments (2)
  1. [§5.5, §4.4, Title/Abstract] The claim that the Gaffke confidence interval is inadmissible for n≥3 is not supported. Theorem 4.1 constructs a valid e-to-p merger K̃_n that is pointwise smaller than K_n only on neutral faces, but this does not imply an interval improvement. Section 4.4 explicitly states that K̃_n is 'not globally coordinatewise monotone' and that the improvement 'may leave the closure of the confidence set unchanged.' Inverting a non-monotone test need not produce an interval; the nonrejection set can be disconnected. Thus the sentence in §5.5, 'The same statements can be made for n≥3 based on results in Section 4,' is not justified. The title and abstract should be qualified to n=2, or rephrased to refer to the e-to-p merger inadmissibility.
  2. [§5.5] Even for n=2, the assertion that the Gaffke confidence interval is strictly improved by the interval generated by K_2^ad for some data points is not proved. The paper states in §4.4 that it does not study the two-input interval ('We do not study that special small-sample interval here'). Pointwise domination of the test K_2^ad ≤ K_2 does not automatically imply that the quantile endpoints L and U move; an explicit sample or a general argument showing strict interval containment is required.
minor comments (5)
  1. [§3.2, §5.5] Cross-references are inconsistent: in §3.2 'Theorem 3.4' should be 'Lemma 3.4', and in §5.5 'Theorem 3.5' should be 'Lemma 3.5'.
  2. [§4, Eq. (27)] The notation for the dominator is introduced as 'Define Kn(x) = ...' without a tilde in the displayed equation, while the surrounding text uses K̃_n. Please ensure consistent notation.
  3. [Abstract] The abstract says 'A neutral-face extension proves inadmissibility of K_n for every n≥2,' which is correct for the e-to-p merger, but the title claims inadmissibility of the confidence interval. Please reconcile the wording so the scope is unambiguous.
  4. [§7.4] The EBE comparison uses a coarse grid for the standard deviation; the authors acknowledge this, but the abstract's claim that Gaffke is the shortest among comparators with the same first-order target should be tempered by this numerical approximation.
  5. [Throughout] Minor typos: 'iid' should be 'i.i.d.' in several places; 'Sidak' should be 'Šidák'.

Circularity Check

0 steps flagged

No significant circularity: core derivations are proved in-paper; self-citations are comparators, not load-bearing.

full rationale

The paper's main results are self-contained rather than circular. Theorem 2.2 proves the elementary-symmetric bound K_n e_k <= C(n,k) by induction; Theorem 3.6 proves validity of the two-input improvement K^ad_2 directly using the two-point decomposition; Theorem 3.7 proves the unique-dominator envelope with an explicit least-favorable construction; Theorem 4.1 extends inadmissibility via the neutral-coordinate identity; and Theorem 6.2 derives the asymptotic endpoint expansion from a pathwise CLT for Dirichlet averages. No parameter is fitted to a subset of data and then reported as a predicted or derived quantity. The only external load-bearing validity fact is the Vlassis-Thomas (2026) proof of Gaffke's conjecture, cited in Theorem 2.1, and that is not a self-citation. The self-citations (Ming et al. for SymPol, Ramdas and Manole for randomized Markov bounds, Vovk and Wang for e-to-p merger terminology) are used as comparators or context, not as the justification for the paper's own derivations. The n>=3 confidence-interval inadmissibility claim in Section 5.5 is weaker than stated, since Section 4.4 itself concedes that the higher-dimensional dominator is not globally coordinatewise monotone and may leave inverted endpoints unchanged; however, that is an unsupported extrapolation or correctness gap, not a circular reduction to the paper's inputs. Accordingly, no circular step is present.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

The central claims rest on the external Vlassis–Thomas validity theorem, standard CLT machinery, and the iid bounded-data model. No fitted free parameters or newly invented entities are used; the numerical comparisons use only user-specified confidence levels and disclosed approximation grids.

axioms (3)
  • domain assumption Validity of Gaffke's statistic as an e-to-p merger (Theorem 2.1, Vlassis–Thomas [2026]): for independent nonnegative X_i with E[X_i]≤1, P{K_n(X)≤α}≤α.
    Assumed rather than proved; every validity and coverage claim in the paper relies on it (Sections 2.1, 5.1, 6).
  • standard math Standard CLT, Slutsky, and Lindeberg tools for the conditional quantile CLT (Lemma 6.1).
    Used to derive the endpoint expansions and the asymptotic efficiency result; no alternative proof is supplied.
  • domain assumption For the CI asymptotics, X_1,X_2,... are iid on [0,1] with σ^2>0 and empirical mean/variance converge a.s.
    Invoked in Theorem 6.2; the degenerate σ=0 case is handled separately in Section 6.3.

pith-pipeline@v1.3.0-alltime-deepseek · 22170 in / 34817 out tokens · 262536 ms · 2026-08-01T14:45:33.283565+00:00 · methodology

0 comments
read the original abstract

Given observations $\mathbf x=(x_1,\dots,x_n)$, Gaffke (2005) defined \[ K_n(\mathbf x)=\mathbb{P}_{\mathbf D}\!\left\{\sum_{i=1}^n x_iD_i\le 1\right\}, \qquad (D_0,D_1,\ldots,D_n)\sim\mathrm{Dirichlet}(1,\ldots,1), \] and conjectured that it is a $p$-value whenever the inputs are independent e-values. Recently, Vlassis and Thomas (2026) proved this conjecture. Inverting the tests for observations in $[0,1]$ gives the confidence interval studied by Learned-Miller and Thomas (2020), which reduces to Clopper--Pearson for Bernoulli data. We give a finite- and large-sample account of Gaffke's test and interval. First, for every $\mathbf x\in[0,\infty)^n$ and every elementary symmetric polynomial $e_k$, \( K_n(\mathbf x)e_k(\mathbf x)\le {n\choose k}, \) so the Gaffke $p$-value never larger than the SymPol $p$-value of Ming et al. (2026). However, Gaffke's p-value is inadmissible. For $n=2$, we construct a valid rule that is strictly smaller on mixed configurations and is the unique admissible rule that dominates $K_2$. A neutral-face extension proves inadmissibility of $K_n$ for every $n\ge2$. If one independent uniform random variable is allowed, there is an even simpler full-dimensional improvement: on the upper orthant, where $K_n(\mathbf x)=1/\prod_i x_i$, replace it by $U/\prod_i x_i$. The equal-tail Gaffke confidence interval $I_n$ is nevertheless first-order asymptotically efficient: for iid observations on $[0,1]$ with unknown variance $\sigma^2>0$, \[ \sqrt n\,\operatorname{Width}(I_n)\longrightarrow 2\sigma z_{1-\alpha/2}\qquad\text{almost surely}. \] Our simulations also find that, among a variety of bounded-mean intervals considered, the Gaffke interval is the shortest, including comparisons with a recent empirical Berry--Esseen procedure having the same first-order Gaussian target.

Figures

Figures reproduced from arXiv: 2607.18661 by Aaditya Ramdas, Ian Waudby-Smith, Jiahao Ming, Ruodu Wang, Yi Shen.

Figure 1
Figure 1. Figure 1: Left: at a fixed level, Kad 2 rejects on a strictly larger part of the mixed quadrant, with the two boundaries agreeing at a = 0 and a = 1. Right: the pointwise reduction K2 − Kad 2 in the mixed quadrant. Now we are ready to prove the validity of Kad 2 . Theorem 3.6. For every pair of independent e-variables E1, E2, P{Kad 2 (E1, E2) ≤ α} ≤ α, 0 ≤ α ≤ 1. That is, Kad 2 is a merger. Proof. The cases α = 0 an… view at source ↗
Figure 2
Figure 2. Figure 2: Empirical coverage of the equal-tail Gaffke interval (left) and the grid-approximated [PITH_FULL_IMAGE:figures/full_fig_p030_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Ratio of mean EBE width to mean Gaffke width. Values above one favor Gaffke. [PITH_FULL_IMAGE:figures/full_fig_p030_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. On Feige's conjecture

    math.PR 2026-07 accept novelty 7.0

    For independent nonnegative mean-one random variables, P(sum < n+1) is at least (n/(n+1))^n ≥ 1/e, proving Feige's conjecture with a matching extremal example.

  2. On the Order-Conditional Optimality of Gaffke's Bound

    math.ST 2026-07 conditional novelty 6.0

    Gaffke's bound is Buehler-optimal within the class of lower confidence bounds that induce its own sample ordering, for the maximum marginal mean of independent nonnegative variables.

Reference graph

Works this paper leans on

21 extracted references · cited by 2 Pith papers

  1. [1]

    The Annals of Statistics , volume=

    E-values: Calibration, combination and applications , author=. The Annals of Statistics , volume=. 2021 , publisher=

  2. [2]

    Carlson, B. C. , title =. Journal of Approximation Theory , year =

  3. [3]

    Clopper, C. J. and Pearson, E. S. , title =. Biometrika , year =

  4. [4]

    Mathematical Methods of Statistics , year =

    Gaffke, Norbert , title =. Mathematical Methods of Statistics , year =

  5. [5]

    , title =

    Learned-Miller, Erik and Thomas, Philip S. , title =. 2020 , howpublished =

  6. [6]

    2026 , howpublished =

    Ming, Jiahao and Ramdas, Aaditya and Shen, Yi and Wang, Ruodu and Waudby-Smith, Ian , title =. 2026 , howpublished =

  7. [7]

    , title =

    Vlassis, Nikos and Thomas, Philip S. , title =. 2026 , howpublished =

  8. [8]

    and Zhao, L

    Wang, W. and Zhao, L. H. , title =. Journal of Statistical Planning and Inference , year =

  9. [9]

    Journal of the Royal Statistical Society Series B: Statistical Methodology , year =

    Waudby-Smith, Ian and Ramdas, Aaditya , title =. Journal of the Royal Statistical Society Series B: Statistical Methodology , year =

  10. [10]

    Rocky Mountain Journal of Mathematics , year =

    zu Castell, Wolfgang , title =. Rocky Mountain Journal of Mathematics , year =

  11. [11]

    Empirical

    Maurer, Andreas and Pontil, Massimiliano , journal=. Empirical

  12. [12]

    Journal of the American statistical association , volume=

    Probability inequalities for sums of bounded random variables , author=. Journal of the American statistical association , volume=. 1963 , publisher=

  13. [13]

    Confidence limits for the expected value of an arbitrary bounded random variable with a continuous distribution function , author=

  14. [14]

    IEEE Transactions on Information Theory , volume=

    Tight concentrations and confidence sequences from the regret of universal portfolio , author=. IEEE Transactions on Information Theory , volume=. 2024 , doi=

  15. [15]

    Bentkus, Vidmantas , journal=. On. 2004 , publisher=

  16. [16]

    Efficient concentration with

    Austern, Morgane and Mackey, Lester , journal=. Efficient concentration with

  17. [17]

    Foundations and Trends

    Hypothesis testing with e-values , author=. Foundations and Trends. 2025 , publisher=

  18. [18]

    Randomized and exchangeable improvements of

    Ramdas, Aaditya and Manole, Tudor , journal=. Randomized and exchangeable improvements of. 2026 , publisher=

  19. [19]

    The Annals of Statistics , volume=

    Admissible ways of merging p-values under arbitrary dependence , author=. The Annals of Statistics , volume=. 2022 , publisher=

  20. [20]

    Advances in Neural Information Processing Systems , volume=

    Vor. Advances in Neural Information Processing Systems , volume=

  21. [21]

    International Conference on Machine Learning , pages=

    Towards practical mean bounds for small samples , author=. International Conference on Machine Learning , pages=. 2021 , organization=