REVIEW 28 references
Investigation of finite-sample properties of robust location and scale estimators
T0 review · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Finite-sample breakdown points of the median, Hodges-Lehmann, MAD and Shamos estimators are derived in closed form, and Monte Carlo unbiasing factors and relative efficiencies are tabulated for sample sizes up to 100.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
This paper does two things. First, it derives exact formulas for the finite-sample breakdown point, the largest fraction of observations that can be corrupted before the estimator produces an arbitrarily large answer. For example, the median of 10 observations breaks down at 0.4, not the asymptotic 0.5. The paper gives similar closed forms for the three variants of the Hodges-Lehmann estimator and for the MAD and Shamos scale estimators.
Second, it uses 10 million simulations of normal data for each sample size from 2 to 100 to measure the bias and variance of these estimators. The MAD systematically underestimates the standard deviation at small n, while the Shamos estimator overestimates it. The paper provides correction factors, called unbiasing factors, and shows the corrections reduce bias and mean square error. It also fits simple curves to the simulated biases so the corrections can be extended to sample sizes above 100. A quality control chart example shows the bias-corrected Shamos estimator is nearly as efficient as the usual standard deviation when data are clean, and much more stable when one observation is corrupted.
Extended reading notes
Core claim
The finite-sample breakdown point of the HL1 and Shamos estimators is given in closed form by epsilon_n = floor(n - 1/2 - sqrt((n - 1/2)^2 - 2 floor((n^2 - n - 2)/4)))/n, with similar explicit formulas for HL2, HL3, median and MAD, and these values differ substantially from the asymptotic limits at small n.
Load-bearing premise
The finite-sample breakdown values are computed under the replacement-breakdown definition in Equation (1), which counts the largest m such that the bias remains finite. Under the alternative definition cited in the same section, Equation (1.6.1) of Hettmansperger and McKean (2010), the same estimators have different breakdown values (for example, the median at n = 3 would be 2/3 instead of 1/3), so the reported numbers are convention-dependent.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (5)
- Least-squares coefficients for A_n (MAD bias), Hayes model =
a1 = -0.76213, a2 = -0.86413
- Least-squares coefficients for A_n (MAD bias), Williams model =
a3 = -0.804168866, a4 = -1.008922
- Least-squares coefficients for B_n (Shamos bias), Hayes model =
b1 = 0.414253297, b2 = 0.442396799
- Least-squares coefficients for B_n (Shamos bias), Williams model =
b3 = 0.435760656, b4 = -1.0084443
- Least-squares coefficients for variance extrapolation formulas =
e.g., median odd: -0.6589, -0.943; median even: -2.1950, 1.929; MAD odd: 0.2996, -149.357; MAD even: -2.417, -153.010…
assumptions (4)
- standard math The replacement breakdown point of the median of N values is floor((N-1)/2)/N.
- domain assumption The supremum bias in the breakdown definition is achieved by sending corrupted observations to infinity, so the number of corrupted Walsh averages that become arbitrarily large determines breakdown.
- domain assumption Simulation draws from N(0,1) and scale equivariance extend biases and variances to N(mu, sigma^2).
- domain assumption Asymptotic relative efficiencies ARE(MAD|S_n) = 0.37 and ARE(Shamos|S_n) = 0.863 are correct values from the literature.
Cite this review
Pith. "Pith review of Investigation of finite-sample properties of robust location and scale estimators." pith.science (2026). https://pith.science/paper/436YLQNP
@misc{pith2026190800462,
author = {Pith},
title = {Pith review of: Investigation of finite-sample properties of robust location and scale estimators},
year = {2026},
howpublished = {\url{https://pith.science/paper/436YLQNP}},
note = {Machine review of arXiv:1908.00462}
}
read the original abstract
When the experimental data set is contaminated, we usually employ robust alternatives to common location and scale estimators such as the sample median and Hodges-Lehmann estimators for location and the sample median absolute deviation and Shamos estimators for scale. It is well known that these estimators have high positive asymptotic breakdown points and are Fisher-consistent as the sample size tends to infinity. To the best of our knowledge, the finite-sample properties of these estimators, depending on the sample size, have not well been studied in the literature. In this paper, we fill this gap by providing their closed-form finite-sample breakdown points and calculating the unbiasing factors and relative efficiencies of the robust estimators through the extensive Monte Carlo simulations up to the sample size 100. The numerical study shows that the unbiasing factor improves the finite-sample performance significantly. In addition, we provide the predicted values for the unbiasing factors obtained by using the least squares method which can be used for the case of sample size more than 100.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
ASQC and ANSI (1972). ASQC (ANSI Standards: A1-1971 (Z1.5-1971), Definitions, Symbols, Formulas, and Tables for Control Charts (approved Nov. 18, 1971) . American National Standards Institute, Milwaukee, Wisconsin
work page 1972
-
[2]
Manual on Presentation of Data and Control Chart Analysis
ASTM E11 (1976). Manual on Presentation of Data and Control Chart Analysis . American Society for Testing and Materials, Philadelphia, PA, 4th edition
work page 1976
-
[3]
Manual on Presentation of Data and Control Chart Analysis
ASTM E11 (2018). Manual on Presentation of Data and Control Chart Analysis . American Society for Testing and Materials, West Conshohocken, PA, 9th edition. S. N. Luko (Ed.)
work page 2018
-
[4]
Donoho, D. and Huber, P. J. (1983). The notion of breakdown point. In A F estschrift for E rich L . L ehmann , Wadsworth Statist./Probab. Ser., pages 157--184. Wadsworth, Belmont, CA
work page 1983
-
[5]
Fisher, R. A. (1922). On the mathematical foundations of theoretical statistics. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character , 222:309--368
work page 1922
-
[6]
Hampel, F. R. (1974). The influence curve and its role in robust estimation. Journal of the American Statistical Association , 69:383--393
work page 1974
-
[7]
R., Ronchetti, E., Rousseeuw, P
Hampel, F. R., Ronchetti, E., Rousseeuw, P. J., and Stahel, W. A. (1986). Robust Statistics: The Approach Based on Influence Functions . John Wiley & Sons, New York
work page 1986
-
[8]
Hayes, K. (2014). Finite-sample bias-correction factors for the median absolute deviation. Communications in Statistics: Simulation and Computation , 43:2205--2212
work page 2014
Show all 28 references
-
[9]
Hettmansperger, T. P. and McKean, J. W. (2010). Robust Nonparametric Statistical Methods . Chapman & Hall/CRC, Boca Raton, FL, 2nd edition
2010
-
[10]
Hodges, J. L. and Lehmann, E. L. (1963). Estimates of location based on rank tests. Annals of Mathematical Statistics , 34:598--611
1963
-
[11]
Hodges, Jr. , J. L. (1967). Efficiency in normal samples and tolerance of extreme values for some estimates of location. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability , volume 1, pages 163--186, Berkeley. University of California Press
1967
-
[12]
Huber, P. J. and Ronchetti, E. M. (2009). Robust Statistics . John Wiley & Sons, New York, 2nd edition
2009
-
[13]
S., and Reisen, V
L\` e vy-Leduc, C., Boistard, H., Moulines, E., Taqqu, M. S., and Reisen, V. A. (2011). Large sample behaviour of some well-known robust estimators under long-range dependence. Statistics , 45:59--71
2011
-
[14]
Montgomery, D. C. (2013). Statistical Quality Control: An Modern Introduction . John Wiley & Sons, 7th edition
2013
-
[15]
Ouyang, L., Park, C., Byun, J.-H., and Leeds, M. (2019). Robust design in the case of data contamination and model departure. In Lio, Y., Ng, H., Tsai, T. R., and Chen, D. G., editors, Statistical Quality Technologies: Theory and Practice (ICSA Book Series in Statistics) , pag...
2019
-
[16]
and Wang, M
Park, C. and Wang, M. (2019). rQCC : Robust quality control chart. https://cran.r-project.org/web/packages/rQCC/. R package version 0.19.8.2
2019
-
[17]
R: A Language and Environment for Statistical Computing
R Core Team (2019). R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing, Vienna, Austria
2019
-
[18]
Rousseeuw, P. (1991). Tutorial to robust statistics. Journal of Chemometrics , 5:1 -- 20
1991
-
[19]
and Croux, C
Rousseeuw, P. and Croux, C. (1993). Alternatives to the median absolute deviation. Journal of the American Statistical Association , 88:1273--1283
1993
-
[20]
Rousseeuw, P. J. and Croux, C. (1992). Explicit scale estimators with high breakdown point. In Dodge, Y., editor, L_1 -Statistical Analysis and Related Methods , pages 77--92. North-Holland,
1992
-
[21]
Serfling, R. J. (2011). Asymptotic relative efficiency in estimation. In Lovric, M., editor, Encyclopedia of Statistical Science, Part I , pages 68--82. Springer-Verlag, Berlin
2011
-
[22]
Shamos, M. I. (1976). Geometry and statistics: Problems at the interface. In Traub, J. F., editor, Algorithms and Complexity: New Directions and Recent Results , pages 251--280. Academic Press, New York
1976
-
[23]
Shewhart, W. A. (1926). Quality control charts. Bell Systems Technical Journal , pages 593--603
1926
-
[24]
Shewhart, W. A. (1931). Economic Control of Quality of Manufactured Product . Van Nostrand Reinhold, Princeton, NJ. Republished in 1981 by the American Society for Quality Control, Milwaukee, Wisconsin
1931
-
[25]
Vining, G. (2009). Technical advice: Phase I and Phase II control charts. Quality Engineering , 21(4):478--479
2009
-
[26]
Walsh, J. E. (1949). Some significance tests for the median which are valid under very general conditions. The Annals of Mathematical Statistics , 20:54--81
1949
-
[27]
Wilcox, R. R. (2016). Introduction to Robust Estimation and Hypothesis Testing . Academic Press, Loncon, United Kingdom, 4th edition
2016
-
[28]
Williams, D. C. (2011). Finite sample correction factors for several simple robust estimators of normal standard deviation. Journal of Statistical Computation and Simulation , 81:1697--1702
2011
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.