REVIEW 2 major objections 4 minor 5 references
Sharp constants turn large-alphabet uniformity testing into a Gaussian SNR that also fixes how many bins calibration tests should use.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 08:25 UTC pith:TB4R5U22
load-bearing objection Clean, usable perspective note that turns the author’s concurrent SNR formula into an explicit binning rule; novelty is modest and the math is imported, but the design advice is solid and immediately applicable. the 2 major comments →
Why Constants Matter in Distribution Testing: From Uniformity to Calibration
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In large-alphabet uniformity testing against ℓ1-separated alternatives, the minimax risk converges to the Gaussian risk 2Φ(−u/2) whose signal-to-noise ratio is u=√(2N)sinh(nε²/(2N)); inverting that formula supplies the largest number of bins that can still achieve a prescribed risk level when the same theory is applied to binned calibration testing.
What carries the argument
The effective SNR u=√(2N)sinh(nε²/(2N)) (and its small-signal reduction nε²/√(2N)) that converts the high-dimensional multinomial problem into a one-dimensional Gaussian testing problem, together with the occupancy-histogram statistic that attains the limiting risk.
Load-bearing premise
The Gaussian risk formula and the claimed optimal statistic are taken from a Poissonized intermediate-regime analysis and are assumed to approximate the finite-sample multinomial problem of practical interest.
What would settle it
Simulate the multinomial uniformity problem at the (n,N,ε) triples where u is order-1, compute the empirical Type-I+Type-II risk of the occupancy-histogram test, and check whether it tracks 2Φ(−u/2) within the claimed o(1) error; a systematic gap falsifies the design rule for N_max.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The note argues that sharp constants in distribution testing, beyond rate-level sample complexity, are essential for ranking procedures, identifying effective signal-to-noise ratios, and guiding practical design choices. Using large-alphabet uniformity testing as the model problem, it recalls that the intermediate-regime minimax risk takes the Gaussian form 2\Phi(-u/2)+o(1) with effective SNR u=\sqrt(2N)sinh(n\epsilon^{2}/(2N)), attained by a linear functional of the occupancy histogram. It then reduces continuous calibration assessment of predictive models to binned uniformity testing of PIT values and inverts the SNR formula to obtain an explicit upper bound N_max on the number of bins that preserves a target risk (or Type-I/II error pair). An oscillatory miscalibration example and a Monte-Carlo risk curve illustrate the resulting bias–variance trade-off.
Significance. If the imported intermediate-regime asymptotics hold for the multinomial problems of interest, the note supplies a clean, operational design rule that turns abstract uniformity-testing theory into a concrete recommendation for calibration diagnostics. The Gaussian analogy, the explicit inversion to N_max, and the Monte-Carlo illustration of the U-shaped risk curve are transparent and immediately usable. The contribution is primarily expository and applicative rather than a new theorem; its value lies in making the constant-level viewpoint accessible and in linking two literatures (distribution testing and ML calibration). Strengths include the elementary algebraic derivation of the binning rule once the SNR is granted, the clear statement of the competing discretization and sampling errors, and the reproducible Monte-Carlo check for a concrete oscillatory density.
major comments (2)
- Section 4 and the display for u_\epsilon,n,N: the central risk formula R^*=2\Phi(-u/2)+o(1) and the claim that the linear occupancy statistic is asymptotically minimax are imported wholesale from Kipnis (2026b) without re-derivation or even a sketch of the Poissonized intermediate-regime argument. The note’s quantitative claims (including the subsequent N_max formula) therefore stand or fall with that external result; a short self-contained statement of the precise asymptotic regime (n,N,\epsilon scaling) under which the o(1) vanishes would make the dependence transparent and allow readers to judge applicability to finite-n multinomial calibration problems.
- Section 8, the N_max formula: the inversion assumes that the binned separation \epsilon remains fixed while N varies, yet for any fixed continuous alternative the binned ECE itself depends on N (as the oscillatory example later shows via the sinc factor). The design rule therefore needs an explicit caveat that \epsilon must be interpreted as a lower bound on the resolvable binned distance, or the formula should be restated in terms of continuous ECE attenuated by a resolution-dependent factor; otherwise the claimed N_max can be optimistic for alternatives whose mass cancels inside bins.
minor comments (4)
- Figure 1 caption and surrounding text: the Monte-Carlo curve is informative, but the precise definition of the test statistic used in the simulations (T_N or a thresholded version) and the number of Monte-Carlo repetitions should be stated in the caption itself for reproducibility.
- Section 6: the reduction from continuous calibration to binned uniformity is standard once PIT values are formed; a brief pointer to earlier literature on PIT-based calibration tests would help situate the contribution.
- Throughout: the notation switches between \epsilon for continuous and binned ℓ1 distances without always flagging the distinction; a consistent subscript or a short glossary would reduce ambiguity.
- References: Kipnis (2026a,b) are listed as forthcoming or under review; if they remain unpublished at the time of decision, the note should either include the essential statements as appendices or clearly mark the dependence as conditional.
Circularity Check
Load-bearing SNR, risk formula, optimal statistic and N_max all reduce to algebraic rearrangement of concurrent self-citations (Kipnis 2026a,b); the note supplies no independent derivation.
specific steps
-
self citation load bearing
[Section 4 (Sharp minimax risk)]
"In large-alphabet uniformity testing against ℓp-separated alternatives, Kipnis [2026b] characterizes this minimax risk at the constant level. ... the limiting risk takes the Gaussian form R∗(ϵ, n, N) = 2Φ(−u/2) + o(1), ... u = uϵ,n,N = √(2N) sinh(nϵ²/(2N)). ... The analysis also identifies the asymptotically minimax test. The optimal statistic is a linear functional of the occupancy histogram Zm ..."
The paper's sole source for the claimed limiting risk, the explicit SNR formula, and the optimality of the occupancy statistic is a concurrent self-citation. No derivation is supplied; every later quantitative claim rests on this imported characterization.
-
self citation load bearing
[Section 8 (Sharp constants give an optimal binning rule) and display for Nmax]
"Given a target total risk R, the condition 2Φ(−uϵ,n,N /2) ≤ R is equivalent to uϵ,n,N ≥ 2Φ−1(1−R/2). Solving this inequality gives the maximum number of bins Nmax = ⌊ n²ϵ⁴ / 8[Φ−1(1−R/2)]² ⌋. ... This statistic is closely related to chi-squared ..."
Nmax is obtained by purely algebraic rearrangement of the SNR formula imported from Kipnis [2026b]; the claimed optimality of TN is likewise taken from the same self-citation. The design rule therefore contains no independent content beyond the concurrent paper.
-
self citation load bearing
[Section 6 and Section 9 (calibration reduction and oscillatory example)]
"This formulation is developed in Kipnis [2026a], which uses sharp minimax uniformity testing to derive optimal binning and minimax calibration tests ... For example, Kipnis [2026a] considers n=5000, k=50, a=0.2, for which ϵ∞=2a/π≈0.127. ... Nmax=303. The resulting risk curve is illustrated in Figure 1."
The continuous-to-binned ECE reduction, the concrete numerical example, and the Monte-Carlo risk curve that 'validates' Nmax are all taken from the author's concurrent paper Kipnis [2026a]. The present note supplies no independent verification.
full rationale
The note is explicit that it does not re-derive the intermediate-regime asymptotics; it imports the Gaussian risk R^*=2Φ(-u/2)+o(1) with the precise SNR u=√(2N)sinh(nε²/(2N)) and the occupancy-histogram optimality from the author's concurrent paper Kipnis [2026b], then obtains the calibration binning rule by elementary inversion of that imported inequality. The continuous-to-binned reduction and the Monte-Carlo illustration are likewise taken from Kipnis [2026a]. Because those concurrent works are the sole source of the quantitative claims and are not independently verified inside the present manuscript, the central design formulas are load-bearing self-citations rather than first-principles results. The surrounding prose arguing that 'constants matter' is independent commentary, so the circularity is partial (score 6) rather than total. No self-definitional tautology or data-fitting loop is present; the reduction is purely by citation-plus-algebra.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Poissonized sampling model is asymptotically equivalent to the multinomial model for the risk quantities of interest
- domain assumption There exists an intermediate asymptotic regime in which the minimax risk converges to a nontrivial constant of the form 2Φ(−u/2)
- domain assumption Equal-width binning of the probability-integral-transform values reduces continuous calibration assessment to multinomial uniformity testing with separation measured by binned ECE
- ad hoc to paper The continuous ℓ1 calibration error is upper-bounded by the binned ECE and the attenuation factor |sinc(k/N)| adequately describes oscillatory alternatives
read the original abstract
Distribution goodness-of-fit testing has developed a powerful rate-level theory: we often know how the required sample size scales with the alphabet size, the separation from the null, and the target error probability. Uniformity testing is the canonical example. One can distinguish the uniform distribution on $N$ categories from alternatives at total-variation distance at least $\epsilon$ with far fewer than $N$ samples, and the optimal scaling is now well understood. But rate-level theory leaves an important question unresolved: among several tests with the same sample-complexity order, which one actually gives the best risk or power? This is a constant-level question. It is especially relevant in modern applications where distribution testing is used not merely as an asymptotic abstraction, but as a practical design tool. This note argues that sharp constants in distribution testing play a role analogous to Fisher information in parametric estimation and Pinsker's constant in nonparametric estimation. First, they distinguish between tests that are all rate-optimal but not equally powerful. Second, they reveal the effective signal-to-noise ratio governing the testing problem. Third, they can guide tuning-parameter choices in downstream applications. We illustrate this perspective through large-alphabet uniformity testing and then explain why the same logic matters for choosing the number of bins in calibration testing.
Figures
Reference graph
Works this paper leans on
-
[1]
Sharp constants in uniformity testing via the huber statistic
Shivam Gupta and Eric Price. Sharp constants in uniformity testing via the huber statistic. In Conference on Learning Theory, pages 3113--3192. PMLR, 2022
work page 2022
-
[2]
Alon Kipnis. Calibrating the calibration tester: Optimal binning and minimax calibration testing for continuous predictive models. In Towards Trustworthy Predictions: Theory and Applications of Calibration for Modern AI, 2026 a . URL https://openreview.net/forum?id=dy7XNC3W0g
work page 2026
-
[3]
The minimax risk in testing uniformity over large alphabets under missing-ball alternatives
Alon Kipnis. The minimax risk in testing uniformity over large alphabets under missing-ball alternatives. IEEE Transactions on Information Theory, 72 0 (3): 0 1831--1849, 2026 b . doi:10.1109/TIT.2025.3646804
-
[4]
Michael Nussbaum. Minimax risk: Pinsker bound. Encyclopedia of Statistical Sciences, 3: 0 451--460, 1999
work page 1999
-
[5]
A coincidence-based test for uniformity given very sparsely sampled discrete data
Liam Paninski. A coincidence-based test for uniformity given very sparsely sampled discrete data. IEEE Transactions on Information Theory, 54 0 (10): 0 4750--4755, 2008
work page 2008
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.