Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

A Comprehensive Comparison of the Wald, Wilson, and adjusted Wilson Confidence Intervals for Proportions

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Measured by mean coverage over a dense grid of n and p, adjusted Wilson intervals with 3, 4, and 6 pseudo-observations are optimal at 90%, 95%, and 99% confidence, beating Wald and Wilson.

desk verdict A readable abstract, an unreadable body, and a headline 3/4/6 result that is plausible but contingent on the chosen coverage criterion and grid. read the letter →

arxiv 2508.10223 v1 pith:Y7JUXPC4 submitted 2025-08-13 stat.ME stat.CO

classification stat.MEstat.CO MSC 62F25
keywords binomialproportioncoverageprobabilityWaldintervalWilsonadjustedpseudo-observationsmeanfinite-samplecomparison
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper takes on a well-known problem in statistics teaching: the most commonly taught confidence interval for a proportion, the Wald interval (sample proportion plus or minus z times the standard error), has poor coverage for small samples and for proportions near 0 or 1. Setting out to find a simple fix, the paper evaluates a family of "adjusted Wilson" intervals in which a small number of pseudo-observations is added to the data before constructing the interval. Using mean coverage probability over every sample size $n=1,\dots,1000$ and every proportion $p=0.01,\dots,0.99$, it claims that the optimal number of pseudo-observations is 3 at 90% confidence, 4 at 95%, and 6 at 99%, and that at those settings the adjusted Wilson interval has higher mean coverage than either the plain Wald interval or the plain Wilson interval at the same level. If this holds, the practical prescription is nearly free: add a level-specific small number of fake successes and failures and use the Wilson formula. The paper also introduces rainbow pixel plots that show at a glance where each interval's coverage falls short.

What carries the argument

The central object is the adjusted Wilson interval of type $\varepsilon$: take the ordinary Wilson (score) interval for a proportion and first add $\varepsilon/2$ successes and $\varepsilon/2$ failures to the observed counts, so that $\varepsilon$ pseudo-observations are added in total. $\varepsilon=0$ is the plain Wilson interval, and $\varepsilon=4$ is the Agresti-Coull-style correction. The argument is carried by exhaustive enumeration: for each $\varepsilon$, the paper computes coverage for all $n$ and $p$ in the stated grid at each confidence level, averages over the grid, and compares the resulting mean coverage across the Wald, Wilson, and adjusted-Wilson families. The pixel plots are

What would settle it

Recompute the same mean-coverage comparison on an extended grid that includes $p=0.001,\dots,0.999$; if the maximum moves away from $\varepsilon=3/4/6$, the paper's optimality claim depends on the omitted tails. A cheaper check: use the same grid but minimize maximum coverage shortfall instead of mean coverage, and see whether the winning $\varepsilon$ changes.

Watch

Extended reading notes

Core claim

On a finite grid of all sample sizes $n=1,\dots,1000$ and all population proportions $p=0.01,0.02,\dots,0.99$, the paper computes the coverage probability of the Wald, Wilson, and adjusted-Wilson-of-type-$\varepsilon$ intervals at the 90%, 95%, and 99% confidence levels. It reports that the mean coverage probability is maximized at $\varepsilon=3$ for 90%, $\varepsilon=4$ for 95%, and $\varepsilon=6$ for 99%, and that each of these three adjusted intervals also dominates the corresponding Wald and Wilson intervals on the same grid. The result is a finite census rather than an asymptotic theorem: every point in the grid is evaluated directly, and the comparison is displayed as color-coded pix

Load-bearing premise

The entire ranking rests on one definition of "best": coverage averaged uniformly over $p=0.01,0.02,\dots,0.99$ and $n=1,\dots,1000$; if a different loss or grid is used, the winning number of pseudo-observations can change.

Editorial extensions

If this is right

  • At 95% confidence, the exhaustive census reproduces the Agresti-Coull finding: epsilon = 4 adjusted Wilson has higher mean coverage than plain Wilson, and both beat Wald.
  • At 90% and 99%, the same ranking holds with epsilon = 3 and epsilon = 6, so the pseudo-observation trick is not a 95%-only phenomenon.
  • Because the comparison is a finite census over $n=1,\dots,1000$ and $p=0.01,\dots,0.99$, the ranking is not an asymptotic approximation and covers small-sample cases that asymptotics miss.
  • The pixel-color displays give instructors and practitioners a direct visual map of where Wald coverage collapses, making the case for abandoning it easy to show in teaching.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 3/4/6 optima are optima for the uniform mean-coverage loss on this particular grid; under a worst-case-coverage loss, or with p weighted toward realistic values, a different epsilon is likely to win.
  • The grid stops at p = 0.01 and p = 0.99; near the boundary, all three intervals behave poorly, so the paper's ranking should not be read as holding in the extreme tails.
  • The same enumeration could be run for interval length or for other families (e.g., Jeffreys or Clopper-Pearson); those extensions are not in the paper but follow naturally from its design.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript compares three families of confidence intervals for a binomial proportion—Wald, Wilson, and adjusted Wilson obtained by adding epsilon pseudo-observations—using mean coverage probability over n = 1,...,1000 and p = 0.01,...,0.99 at the 90%, 95%, and 99% confidence levels. It proposes pixel-color plots and a rainbow color code for visualizing coverage over the (p, n) grid. The headline finding is that epsilon = 3, 4, 6 is optimal for adjusted Wilson at 90%, 95%, and 99%, respectively, and that these adjusted intervals also have higher mean coverage than the corresponding Wald and Wilson intervals.

Significance. The paper's exhaustive finite-sample census is a useful empirical complement to the asymptotic literature; if the computations are exact and reproducible, the 90% and 99% pseudo-count recommendations (3 and 6) are new numerical results beyond the well-known Agresti-Coull epsilon = 4 at 95%. The proposed visualization may be instructive. However, the optimality claim is currently in-sample: epsilon is chosen by maximizing the same mean-coverage criterion on the same grid used for the final comparison. No sensitivity analysis, expected-length comparison, or out-of-sample check is visible in the supplied text. Therefore the significance of the headline constants is limited until the criterion dependence is quantified.

major comments (4)
  1. [Abstract] The headline epsilon values are selected by optimizing the mean coverage criterion on the same finite grid that is then used to declare the adjusted Wilson interval the winner. This is circular/in-sample. Report the mean-coverage curve as a function of epsilon for each level and test robustness to omitting small n, changing p-grid resolution, weighting, and exact vs simulated evaluation. In particular, the 99% optimum epsilon = 6 needs a sensitivity check.
  2. [Abstract] The coverage-only criterion is insufficient: a longer interval tends to have higher coverage, so 'performs better' is ambiguous. Report expected length or a coverage-length trade-off metric for the Wald, Wilson, and adjusted Wilson intervals at each level.
  3. [Full text (as supplied)] The full text supplied for review is corrupted (mojibake), so definitions, equations, tables, and figures cannot be checked. The abstract alone does not state whether coverage is exact or Monte Carlo. Provide a clean version and state the computation method; if simulated, give Monte Carlo standard errors.
  4. [Abstract / grid definition] The grid p = 0.01,...,0.99 and n = 1,...,1000 is treated as exhaustive, but it excludes extreme tails and weights all n equally, so the mean is dominated by small-n behavior. Qualify all 'best' claims as conditional on this grid and justify the grid choice or add a sensitivity analysis.
minor comments (4)
  1. [Abstract] 'Type 4' and 'type epsilon' are used without a formal definition; define the adjusted Wilson estimator explicitly.
  2. [Visualization] The rainbow color code should include a colorblind-safe option; pixel plots need clear axis labels and a legend.
  3. [References] Citations to Agresti and Coull (1998) and Brown, Cai, and DasGupta (2001) should be included in the reference list; the supplied text does not show them.
  4. [Abstract] The phrase 'comprehensively compare ... across all sample sizes' is misleading given the discrete grid; consider 'over the evaluated grid'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline epsilon values are the output of an explicitly stated exhaustive enumeration, not a prediction derived from fitted inputs.

full rationale

The paper's central claim is an exhaustive finite census: for each confidence level, it computes mean coverage of the Wald, Wilson, and adjusted-Wilson(ε) intervals over the stated grid (n = 1,...,1000; p = 0.01,...,0.99), then reports which ε maximizes the criterion and that this ε also beats Wald and Wilson. This is a direct report of the computed ordering, not a derivation that assumes its conclusion. The 3/4/6 values are selected by the same mean-coverage criterion used to declare them best, which is exactly what an optimization over a one-parameter family looks like; it would be circular only if the paper claimed out-of-sample prediction or derived the criterion from the intervals. No such claim appears. There are no self-citations, no imported uniqueness theorems, and no equation reduces to its own input. The choice of uniform weighting on a discrete grid is a substantive evaluation-design assumption, and a different loss or grid could change the ranking, but that is a correctness/robustness concern, not a circularity concern under the stated rules.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim depends on three effectively free modeling choices: the pseudo-observation count epsilon (optimized by the same criterion used to score the winner), the discrete uniform p-grid with extremes excluded, and the mean-coverage loss function. The binomial model and standard interval formulas are background assumptions from prior literature.

free parameters (3)
  • Pseudo-observation count epsilon per confidence level = 3 (90%), 4 (95%), 6 (99%)
    Chosen by optimizing mean coverage probability over the same n/p grid used to evaluate the winner, so these are fitted values within the add-epsilon family.
  • Evaluation grid and weighting = p = 0.01, ..., 0.99 (99 points, uniform), n = 1, ..., 1000
    The uniform grid and the exclusion of extreme p regions are modeling choices that determine which epsilon wins; no justification is given in the abstract.
  • Coverage loss function = mean coverage probability (no tail weighting, no width penalty)
    The optimal epsilon depends on the chosen loss; mean coverage is one of several defensible criteria (minimax, average distance to nominal, interval width).
assumptions (4)
  • domain assumption Observed count X follows a binomial(n, p) distribution and coverage is computed from exact binomial probabilities (or Monte Carlo; the abstract does not state which).
    The entire coverage computation rests on the binomial model and on the unstated computational method; a simulation-based census would carry Monte Carlo noise that could shift the reported epsilon optima near ties.
  • domain assumption Mean coverage probability, uniformly weighted over the grid p = 0.01, ..., 0.99, is the appropriate notion of 'best' for a confidence interval.
    This criterion is inherited from Agresti-Coull, but the abstract does not justify uniform weighting or the exclusion of p below 0.01 and above 0.99; the optimal epsilon is a function of this choice.
  • domain assumption The discrete grid with step 0.01 in p and unit steps in n up to 1000 is representative of the continuous parameter space.
    Optimality is declared from a finite census; no continuity or robustness argument is given in the abstract, so the headline 3, 4, 6 values are only proven for the enumerated points.
  • standard math Standard normal quantiles and the standard interval formulas (Wald, Wilson, adjusted Wilson of type epsilon) define the compared objects.
    These are textbook definitions; the adjusted-Wilson parameterization itself is from the cited Agresti-Coull literature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Comprehensive Comparison of the Wald, Wilson, and adjusted Wilson Confidence Intervals for Proportions." pith.science (2026). https://pith.science/paper/Y7JUXPC4

@misc{pith2026250810223,
  author       = {Pith},
  title        = {Pith review of: A Comprehensive Comparison of the Wald, Wilson, and adjusted Wilson Confidence Intervals for Proportions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Y7JUXPC4}},
  note         = {Machine review of arXiv:2508.10223}
}
abstract

The standard confidence interval for a population proportion covered in the overwhelming majority of introductory and intermediate statistics textbooks surprisingly remains the Wald confidence interval despite having a poor coverage probability, especially for small sample sizes or when the unknown population proportion is close to either 0 or 1. Using the mean coverage probability, and for some sample sizes, Agresti and Coull showed not only that the 95\% Wilson confidence interval performs better, but also showed that 95\% adjusted Wilson of type 4 confidence interval, obtained by simply adding four pseudo-observations, outperforms both the Wald and the Wilson confidence intervals. In this paper, we introduce a rainbow color code and pixel-color plots as ways to comprehensively compare the Wald, Wilson, and adjusted-Wilson of type $\epsilon$ confidence intervals across all sample sizes $n=1, 2, \dots, 1000$, population proportion values $p=0.01, 0.02, \dots, 0.99$, and for the three typical confidence levels. We show not only that adding 3 (resp., 4 and 6) pseudo-observations is the best for the 90\% (resp., 95\% and 99\%) adjusted Wilson confidence interval, but it also performs better than both the 90\% (resp., 95\% and 99\%) Wald and Wilson confidence intervals.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Textual Emotion Through Emoji Prediction

    cs.CL 2025-08 unverdicted novelty 2.0 of 10

    On the TweetEval emoji task, BERT scores best overall but a CNN handles rare emoji classes better, with focal loss used to counter class imbalance.

Reference graph

Works this paper leans on

1 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    ������� ������������ ������� ��������� �������� �� ������� ��� �������� ��������� ��������� �� ������ ���� ������ ���� ������ ��� ���� ��������� ������� ��� �������� ��� ���������� �� ������� ��� ��������� ���������� ������� ������ ����������������������������� �������� �������������� ������ ��� ���� ��������� ������� ��� �������� ��� ���������� �� ������...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.