REVIEW 4 major objections 4 minor 1 cited by
A Comprehensive Comparison of the Wald, Wilson, and adjusted Wilson Confidence Intervals for Proportions
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Measured by mean coverage over a dense grid of n and p, adjusted Wilson intervals with 3, 4, and 6 pseudo-observations are optimal at 90%, 95%, and 99% confidence, beating Wald and Wilson.
desk verdict A readable abstract, an unreadable body, and a headline 3/4/6 result that is plausible but contingent on the chosen coverage criterion and grid. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the adjusted Wilson interval of type $\varepsilon$: take the ordinary Wilson (score) interval for a proportion and first add $\varepsilon/2$ successes and $\varepsilon/2$ failures to the observed counts, so that $\varepsilon$ pseudo-observations are added in total. $\varepsilon=0$ is the plain Wilson interval, and $\varepsilon=4$ is the Agresti-Coull-style correction. The argument is carried by exhaustive enumeration: for each $\varepsilon$, the paper computes coverage for all $n$ and $p$ in the stated grid at each confidence level, averages over the grid, and compares the resulting mean coverage across the Wald, Wilson, and adjusted-Wilson families. The pixel plots are
What would settle it
Recompute the same mean-coverage comparison on an extended grid that includes $p=0.001,\dots,0.999$; if the maximum moves away from $\varepsilon=3/4/6$, the paper's optimality claim depends on the omitted tails. A cheaper check: use the same grid but minimize maximum coverage shortfall instead of mean coverage, and see whether the winning $\varepsilon$ changes.
Extended reading notes
Core claim
On a finite grid of all sample sizes $n=1,\dots,1000$ and all population proportions $p=0.01,0.02,\dots,0.99$, the paper computes the coverage probability of the Wald, Wilson, and adjusted-Wilson-of-type-$\varepsilon$ intervals at the 90%, 95%, and 99% confidence levels. It reports that the mean coverage probability is maximized at $\varepsilon=3$ for 90%, $\varepsilon=4$ for 95%, and $\varepsilon=6$ for 99%, and that each of these three adjusted intervals also dominates the corresponding Wald and Wilson intervals on the same grid. The result is a finite census rather than an asymptotic theorem: every point in the grid is evaluated directly, and the comparison is displayed as color-coded pix
Load-bearing premise
The entire ranking rests on one definition of "best": coverage averaged uniformly over $p=0.01,0.02,\dots,0.99$ and $n=1,\dots,1000$; if a different loss or grid is used, the winning number of pseudo-observations can change.
Editorial extensions
If this is right
- At 95% confidence, the exhaustive census reproduces the Agresti-Coull finding: epsilon = 4 adjusted Wilson has higher mean coverage than plain Wilson, and both beat Wald.
- At 90% and 99%, the same ranking holds with epsilon = 3 and epsilon = 6, so the pseudo-observation trick is not a 95%-only phenomenon.
- Because the comparison is a finite census over $n=1,\dots,1000$ and $p=0.01,\dots,0.99$, the ranking is not an asymptotic approximation and covers small-sample cases that asymptotics miss.
- The pixel-color displays give instructors and practitioners a direct visual map of where Wald coverage collapses, making the case for abandoning it easy to show in teaching.
Reading between the lines
- The 3/4/6 optima are optima for the uniform mean-coverage loss on this particular grid; under a worst-case-coverage loss, or with p weighted toward realistic values, a different epsilon is likely to win.
- The grid stops at p = 0.01 and p = 0.99; near the boundary, all three intervals behave poorly, so the paper's ranking should not be read as holding in the extreme tails.
- The same enumeration could be run for interval length or for other families (e.g., Jeffreys or Clopper-Pearson); those extensions are not in the paper but follow naturally from its design.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript compares three families of confidence intervals for a binomial proportion—Wald, Wilson, and adjusted Wilson obtained by adding epsilon pseudo-observations—using mean coverage probability over n = 1,...,1000 and p = 0.01,...,0.99 at the 90%, 95%, and 99% confidence levels. It proposes pixel-color plots and a rainbow color code for visualizing coverage over the (p, n) grid. The headline finding is that epsilon = 3, 4, 6 is optimal for adjusted Wilson at 90%, 95%, and 99%, respectively, and that these adjusted intervals also have higher mean coverage than the corresponding Wald and Wilson intervals.
Significance. The paper's exhaustive finite-sample census is a useful empirical complement to the asymptotic literature; if the computations are exact and reproducible, the 90% and 99% pseudo-count recommendations (3 and 6) are new numerical results beyond the well-known Agresti-Coull epsilon = 4 at 95%. The proposed visualization may be instructive. However, the optimality claim is currently in-sample: epsilon is chosen by maximizing the same mean-coverage criterion on the same grid used for the final comparison. No sensitivity analysis, expected-length comparison, or out-of-sample check is visible in the supplied text. Therefore the significance of the headline constants is limited until the criterion dependence is quantified.
major comments (4)
- [Abstract] The headline epsilon values are selected by optimizing the mean coverage criterion on the same finite grid that is then used to declare the adjusted Wilson interval the winner. This is circular/in-sample. Report the mean-coverage curve as a function of epsilon for each level and test robustness to omitting small n, changing p-grid resolution, weighting, and exact vs simulated evaluation. In particular, the 99% optimum epsilon = 6 needs a sensitivity check.
- [Abstract] The coverage-only criterion is insufficient: a longer interval tends to have higher coverage, so 'performs better' is ambiguous. Report expected length or a coverage-length trade-off metric for the Wald, Wilson, and adjusted Wilson intervals at each level.
- [Full text (as supplied)] The full text supplied for review is corrupted (mojibake), so definitions, equations, tables, and figures cannot be checked. The abstract alone does not state whether coverage is exact or Monte Carlo. Provide a clean version and state the computation method; if simulated, give Monte Carlo standard errors.
- [Abstract / grid definition] The grid p = 0.01,...,0.99 and n = 1,...,1000 is treated as exhaustive, but it excludes extreme tails and weights all n equally, so the mean is dominated by small-n behavior. Qualify all 'best' claims as conditional on this grid and justify the grid choice or add a sensitivity analysis.
minor comments (4)
- [Abstract] 'Type 4' and 'type epsilon' are used without a formal definition; define the adjusted Wilson estimator explicitly.
- [Visualization] The rainbow color code should include a colorblind-safe option; pixel plots need clear axis labels and a legend.
- [References] Citations to Agresti and Coull (1998) and Brown, Cai, and DasGupta (2001) should be included in the reference list; the supplied text does not show them.
- [Abstract] The phrase 'comprehensively compare ... across all sample sizes' is misleading given the discrete grid; consider 'over the evaluated grid'.
Circularity Check
No significant circularity: the headline epsilon values are the output of an explicitly stated exhaustive enumeration, not a prediction derived from fitted inputs.
full rationale
The paper's central claim is an exhaustive finite census: for each confidence level, it computes mean coverage of the Wald, Wilson, and adjusted-Wilson(ε) intervals over the stated grid (n = 1,...,1000; p = 0.01,...,0.99), then reports which ε maximizes the criterion and that this ε also beats Wald and Wilson. This is a direct report of the computed ordering, not a derivation that assumes its conclusion. The 3/4/6 values are selected by the same mean-coverage criterion used to declare them best, which is exactly what an optimization over a one-parameter family looks like; it would be circular only if the paper claimed out-of-sample prediction or derived the criterion from the intervals. No such claim appears. There are no self-citations, no imported uniqueness theorems, and no equation reduces to its own input. The choice of uniform weighting on a discrete grid is a substantive evaluation-design assumption, and a different loss or grid could change the ranking, but that is a correctness/robustness concern, not a circularity concern under the stated rules.
Assumptions & free parameters
free parameters (3)
- Pseudo-observation count epsilon per confidence level =
3 (90%), 4 (95%), 6 (99%)
- Evaluation grid and weighting =
p = 0.01, ..., 0.99 (99 points, uniform), n = 1, ..., 1000
- Coverage loss function =
mean coverage probability (no tail weighting, no width penalty)
assumptions (4)
- domain assumption Observed count X follows a binomial(n, p) distribution and coverage is computed from exact binomial probabilities (or Monte Carlo; the abstract does not state which).
- domain assumption Mean coverage probability, uniformly weighted over the grid p = 0.01, ..., 0.99, is the appropriate notion of 'best' for a confidence interval.
- domain assumption The discrete grid with step 0.01 in p and unit steps in n up to 1000 is representative of the continuous parameter space.
- standard math Standard normal quantiles and the standard interval formulas (Wald, Wilson, adjusted Wilson of type epsilon) define the compared objects.
Cite this review
Pith. "Pith review of A Comprehensive Comparison of the Wald, Wilson, and adjusted Wilson Confidence Intervals for Proportions." pith.science (2026). https://pith.science/paper/Y7JUXPC4
@misc{pith2026250810223,
author = {Pith},
title = {Pith review of: A Comprehensive Comparison of the Wald, Wilson, and adjusted Wilson Confidence Intervals for Proportions},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y7JUXPC4}},
note = {Machine review of arXiv:2508.10223}
}
abstract
The standard confidence interval for a population proportion covered in the overwhelming majority of introductory and intermediate statistics textbooks surprisingly remains the Wald confidence interval despite having a poor coverage probability, especially for small sample sizes or when the unknown population proportion is close to either 0 or 1. Using the mean coverage probability, and for some sample sizes, Agresti and Coull showed not only that the 95\% Wilson confidence interval performs better, but also showed that 95\% adjusted Wilson of type 4 confidence interval, obtained by simply adding four pseudo-observations, outperforms both the Wald and the Wilson confidence intervals. In this paper, we introduce a rainbow color code and pixel-color plots as ways to comprehensively compare the Wald, Wilson, and adjusted-Wilson of type $\epsilon$ confidence intervals across all sample sizes $n=1, 2, \dots, 1000$, population proportion values $p=0.01, 0.02, \dots, 0.99$, and for the three typical confidence levels. We show not only that adding 3 (resp., 4 and 6) pseudo-observations is the best for the 90\% (resp., 95\% and 99\%) adjusted Wilson confidence interval, but it also performs better than both the 90\% (resp., 95\% and 99\%) Wald and Wilson confidence intervals.
Forward citations
Cited by 1 Pith paper
-
Understanding Textual Emotion Through Emoji Prediction
On the TweetEval emoji task, BERT scores best overall but a CNN handles rare emoji classes better, with focal loss used to counter class imbalance.
Reference graph
Works this paper leans on
-
[1]
������� ������������ ������� ��������� �������� �� ������� ��� �������� ��������� ��������� �� ������ ���� ������ ���� ������ ��� ���� ��������� ������� ��� �������� ��� ���������� �� ������� ��� ��������� ���������� ������� ������ ����������������������������� �������� �������������� ������ ��� ���� ��������� ������� ��� �������� ��� ���������� �� ������...
work page Pith review arXiv 2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.