REVIEW 3 major objections 4 minor 34 references
On the Symmetry of Limiting Distributions of M-estimators
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper establishes that an M-estimator's limiting distribution is symmetric around zero whenever the constraint set, the deterministic drift, and the stochastic process are all even in law, and proves for linear-process limits that…
desk verdict Fix the statement of Theorem 3.2 and the proof of Lemma 3.1; the symmetry results are otherwise credible and useful. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the limiting stochastic process $Z(u)=D(u)+S(u)+X(u)$ from display (5), where $X$ is $0$ on the tangent cone of the constraint set at the true parameter and $+\infty$ outside it. A process is called even when it and its reflection have the same law; Lemma 3.1 and Theorem 3.2 transfer componentwise evenness plus almost-sure uniqueness of the argmin into $W \stackrel{d}{=} -W$. For the necessity direction, the argument shifts to convex duality: the minimizer $W$ is identified as a subgradient of the convex conjugate of $D+X$ evaluated at $-Y$, and uniqueness of optimal transport maps forces two such subgradient selections to agree almost everywhere, which yields evenness of $D+X$. Epi-convergence supplies the bridge from finite-sample minimizers to the limiting argmin.
What would settle it
Choose a one-dimensional linear-process limit $Z(u)=D(u)+uY$ with $Y$ standard normal and a superlinear convex $D$ that is not even, and compute the distribution of $W=\arg\min Z(u)$; Theorem 4.7 predicts this distribution is asymmetric whenever $D$ is not even, so an observed symmetric distribution in this setup would refute the necessity claim.
Extended reading notes
Core claim
The central claim is that the symmetry of a limiting argmin distribution is governed by the symmetry of the limit process's ingredients, not by the particular loss or estimator. Theorem 3.2 shows that if $Z(u)=D(u)+S(u)+X(u)$ and $Z(-u)$ have the same law and the argmin is almost surely unique, the random minimizer is symmetric about zero. Theorems 4.4 and 4.7 give the converse for the linear class $Z(u)=D(u)+\langle u,Y\rangle$ when $Y$ is symmetric, absolutely continuous, and dominates Lebesgue measure: then a symmetric $W$ forces $D+X$ to be even, so for Class I the paper obtains a complete characterisation. For Class II Gaussian-process limits and Class III Poisson-hyperplane limits, the same theorem supplies checkable sufficient conditions, which the paper verifies for shorth, least-median-of-squares, mode, bridge, LAD, and constrained regression examples.
Load-bearing premise
The load-bearing premise is that the normalized estimator converges in distribution to the almost surely unique minimizer of $D+S+X$, since uniqueness is what lets symmetry of the ingredients be transferred to the argmin, and for Classes II and III this uniqueness is assumed externally for the examples.
Editorial extensions
If this is right
- For cube-root-rate estimators such as shorth, least median of squares, and mode estimation with the tuning parameter below the boundary case, the Gaussian-process covariance is automatically even, so checking the drift and constraint cone decides symmetry of the limit.
- For LAD regression, symmetry holds when the error distribution's point-mass increments just above and below zero balance in the limit; the paper's examples show asymmetric error increments lead to asymmetric limits and nonzero median bias.
- For bridge and LASSO-type penalized regression, concave penalties ($\mu<1$) give even drift at every true parameter, while convex penalties ($\mu\ge 1$) give even drift only at $\theta_0=0$; uniform symmetric inference over the parameter space therefore cannot be certified for the convex case.
- For constrained mean estimation, the tangent cone must be a symmetric set; when the true parameter lies on a boundary with an asymmetric cone, the limiting distribution is asymmetric even though the errors are Gaussian.
Reading between the lines
- Because Theorem 3.2 is stated for any even-in-law process, the same symmetry argument should apply to resampling or bootstrap versions of these limits; any bootstrap scheme that preserves componentwise evenness would inherit the symmetric law.
- The paper leaves open whether a non-even drift outside the support of $Y$ can still produce a symmetric $W$; its own remarks suggest the true boundary is equality of two conjugates on the support of $Y$, which could be tested by constructing asymmetric $D$ that agree on that support.
- Median unbiasedness holds under weaker conditions than full symmetry (Theorem 4.5), so inference procedures built on vanishing median bias rather than distributional symmetry could cover irregular cases where the symmetric-limit condition fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies conditions under which the limiting distributions of M-estimators, represented as minimizers of processes Z(u)=D(u)+S(u)+X(u), are symmetric around zero. Symmetry of the limiting distribution is used to guarantee zero limiting median bias, which is the requirement for valid inference via the HulC method of Kuchibhotla et al. (2024). The paper gives a general sufficient condition (Theorem 3.2) based on evenness of the deterministic drift D, the constraint indicator X, and the stochastic process S, and verifies this condition for Pflug's three classes of limiting processes: linear stochastic processes, Gaussian processes, and generalized Poisson hyperplane processes. It then proves necessity of the drift/constraint evenness assumption (A1) for Class I limits in one dimension (Theorem 4.4) and in higher dimensions under stronger symmetry of Y (Theorem 4.7), using convex analysis and optimal transport uniqueness.
Significance. If the main results stand, the paper delivers a useful and nontrivial contribution: it identifies a simple, checkable sufficient condition for symmetry of non-normal M-estimator limits, and it shows in an important special case that the condition is also necessary. The verification of evenness for several classical irregular estimators (shorth, LMS, mode, LAD, bridge) and the explicit discussion of when (A1) fails (e.g., Example 2.7, mode estimation at the boundary tuning parameter) give practitioners concrete guidance for when HulC is or is not applicable. The necessity proofs through optimal transport are elegant and appear internally consistent. However, the central sufficiency theorem and its supporting lemma are not stated correctly, as detailed below.
major comments (3)
- [Lemma 3.1] The statement P(W in A) = P(inf_A Z < inf_{A^c} Z) is false even under the stated almost-sure uniqueness assumption. For the deterministic even process Z(u)=u^4 on R, the unique minimizer is W=0; taking A={0}, the left-hand side equals 1 while the right-hand side is P(0<0)=0. The conclusion W d= -W is nonetheless true and can be proved by the equivariance of the argmin map under u -> -u, but the printed proof is invalid as written.
- [Theorem 3.2] Theorem 3.2 omits the almost-sure uniqueness hypothesis that Lemma 3.1 requires. As stated, the theorem is false: take Omega=R, D(u)=0, S(u)=0, and X the indicator of R. Every real number is a minimizer, the process is even, and the deterministic selection W=1 is not symmetric around zero. The theorem should either assume that argmin(D+S+X) is almost surely unique or specify a symmetric tie-breaking rule that makes W a well-defined random variable; the current wording does not.
- [Section 4.1 and Example 2.7] The missing uniqueness hypothesis is not merely cosmetic: the paper proves uniqueness for parts of Class I in Theorem 4.3, but for Classes II and III it does not provide general uniqueness results, and Example 2.7 explicitly conditions on the limiting process attaining a unique minimum almost surely as an external assumption. Since Theorem 3.2 is invoked for these classes, the manuscript needs to state explicitly which uniqueness condition is being assumed in each application, or to restrict the sufficiency claim accordingly.
minor comments (4)
- [Section 1] There is a typographical artifact in the last paragraph of the introduction: 'min36imizers' should read 'minimizers'.
- [References] The spelling 'Rockafeller' is used inconsistently; the standard spelling is 'Rockafellar'.
- [Theorem 3.2] The phrase 'W = argmin_{u in Omega} D(u)+S(u)+X(u)' should specify that Omega is the effective domain and that the argmin is taken almost surely unique, to avoid ambiguity with the counterexample above.
- [Section 4.2] In the proof of Theorem 4.4, the sentence 'the derivatives here exist almost surely from almost sure uniqueness of minimizers' would benefit from a brief justification or a reference, since differentiability of the conjugate at the relevant point is what is being used.
Circularity Check
No significant circularity: the symmetry theorems are self-contained mathematical derivations; self-citations to HulC are motivational, not load-bearing.
full rationale
The paper's central derivation chain is self-contained. Theorem 3.2 takes as hypotheses that the constraint set is symmetric, the deterministic drift D is even, and the stochastic process S is even, and concludes that the unique argmin W is symmetric around zero. This is a direct implication; the conclusion is not used to define or fit the hypotheses. The necessity results (Theorems 4.4 and 4.7) assume symmetric, absolutely continuous random shifts and prove that (A1) must hold for W to be symmetric, using convex analysis and optimal transport rather than assuming the target symmetry. No parameter is fitted to data, and no closely related quantity is predicted from a fitted subset. The citations to Kuchibhotla et al. (2023, 2024) and Jain and Kuchibhotla (2023) introduce HulC and median-bias regularity as motivating applications and external validity conditions, but these results do not enter the proof of the symmetry theorems; removing them would leave Theorems 3.2, 4.3, 4.4, and 4.7 intact. The paper explicitly flags which minimizer-uniqueness conditions are imposed externally, as in Example 2.7 where it states 'provided that the limiting process attains a unique minimum almost surely.' Any concerns about the correctness of the Lemma 3.1 proof or about omitted uniqueness hypotheses in Theorem 3.2 are mathematical correctness issues, not circularity, because they do not involve the derivation reducing to its own inputs or to a self-citation chain. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption Epi-convergence and argmin convergence in (5): the normalized M-estimator converges in distribution to the argmin of D+S+X.
- domain assumption The argmin of the limit process is almost surely unique.
- standard math Legendre-Fenchel duality, greatest convex minorants, subdifferential calculus, and differentiability of convex conjugates.
- standard math Optimal transport uniqueness theorems for absolutely continuous measures, including Brenier-McCann and the one-dimensional result of Ambrosio.
- ad hoc to paper Assumption (A1): the deterministic drift D plus constraint indicator X is even.
Cite this review
Pith. "Pith review of On the Symmetry of Limiting Distributions of M-estimators." pith.science (2026). https://pith.science/paper/CMW42HYW
@misc{pith2026241117087,
author = {Pith},
title = {Pith review of: On the Symmetry of Limiting Distributions of M-estimators},
year = {2026},
howpublished = {\url{https://pith.science/paper/CMW42HYW}},
note = {Machine review of arXiv:2411.17087}
}
abstract
Many functionals of interest in statistics and machine learning can be written as minimizers of expected loss functions. Such functionals are called $M$-estimands, and can be estimated by $M$-estimators -- minimizers of empirical average losses. Traditionally, statistical inference (e.g., hypothesis tests and confidence sets) for $M$-estimands is obtained by proving asymptotic normality of $M$-estimators centered at the target. However, asymptotic normality is only one of several possible limiting distributions and (asymptotically) valid inference becomes significantly difficult with non-normal limits. In this paper, we provide conditions for the symmetry of three general classes of limiting distributions, enabling inference using HulC (Kuchibhotla et al. (2024)).
Reference graph
Works this paper leans on
-
[1]
Ambrosio, L. (2003). Lecture notes on optimal transport problems. In Mathematical Aspects of Evolving Interfaces , volume 1812 of Lecture Notes in Mathematics . Springer, Berlin, Heidelberg
work page 2003
-
[2]
Aubin, J.-P. and Frankowska, H. (2009). Set-valued analysis. modern birkh \"a user classics. Birkh \"a user Boston, Inc., Boston, MA. https://doi.org/10.1007/978-0-8176-4848-0 , pages XXI, 461
-
[3]
Bonnans, J. F. and Shapiro, A. (2013). Perturbation analysis of optimization problems . Springer Science & Business Media
2013
-
[4]
Burke, J. V. and Hoheisel, T. (2017). Epi-convergence properties of smoothing by infimal convolution. Set-Valued and Variational Analysis , 25:1--23
work page 2017
-
[5]
Chernozhukov, V. and Hong, H. (2002). Likelihood inference in a class of nonregular econometric models. Available at SSRN 301486
work page 2002
-
[6]
Geyer, C. J. (1994). On the asymptotics of constrained m-estimation. The Annals of statistics , pages 1993--2010
work page 1994
-
[7]
Hess, C. (1996). Epi-convergence of sequences of normal integrands and strong consistency of the maximum likelihood estimator. The Annals of Statistics , 24(3):1298--1315
work page 1996
-
[8]
Hiriart-Urruty, J.-B. and Lemar \'e chal, C. (2004). Fundamentals of convex analysis . Springer Science & Business Media
work page 2004
Show all 34 references
-
[9]
and Kuchibhotla, A
Jain, A. and Kuchibhotla, A. K. (2023). Rectangular hull confidence regions for multivariate parameters. arXiv preprint arXiv:2311.16598
2023 arXiv
-
[10]
Kall, P. (1986). Approximation to optimization problems: An elementary review. Mathematics of Operations Research , 11(1):9--18
1986
-
[11]
Kaufmann, H. (1988). On existence and uniqueness of a vector minimizing a convex function. Zeitschrift f \"u r Operations Research , 32:357--373
1988
-
[12]
and Pollard, D
Kim, J. and Pollard, D. (1990). Cube root asymptotics. The Annals of Statistics , pages 191--219
1990
-
[13]
Knight, K. (1998). Limiting distributions for l1 regression estimators under general conditions. The Annals of Statistics , 26(2):755--770
1998
-
[14]
Knight, K. (1999a). Asymptotics for l1-estimators of regression parameters under heteroscedasticityy. Canadian Journal of Statistics , 27(3):497--507
1999
-
[15]
Knight, K. (1999b). Epi-convergence in distribution and stochastic equi-semicontinuity. Unpublished manuscript , 37(7):14
1999
-
[16]
and Fu, W
Knight, K. and Fu, W. (2000). Asymptotics for lasso-type estimators. The Annals of statistics , 28(5):1356--1378
2000
-
[17]
K., Balakrishnan, S., and Wasserman, L
Kuchibhotla, A. K., Balakrishnan, S., and Wasserman, L. (2023). Median regularity and honest inference. Biometrika , 110(3):831--838
2023
-
[18]
K., Balakrishnan, S., and Wasserman, L
Kuchibhotla, A. K., Balakrishnan, S., and Wasserman, L. (2024). The HulC: confidence regions from convex hulls . Journal of the Royal Statistical Society Series B: Statistical Methodology , 86(3):586--622
2024
-
[19]
and Wright, S
Nocedal, J. and Wright, S. J. (2006). Numerical Optimization . Springer Series in Operations Research and Financial Engineering. Springer New York, NY, 2 edition
2006
-
[20]
Parthasarathy, T. (1972). General theorems on selectors , pages 49--57. Springer Berlin Heidelberg, Berlin, Heidelberg
1972
-
[21]
Pensia, A., Jog, V., and Loh, P.-L. (2019). Mean estimation for entangled single-sample distributions. In 2019 IEEE International Symposium on Information Theory (ISIT) , pages 3052--3056. IEEE
2019
-
[22]
Pfanzagl, J. (1970). On the asymptotic efficiency of median unbiased estimates. The Annals of Mathematical Statistics , pages 1500--1509
1970
-
[23]
Pflug, G. C. (1995). Asymptotic stochastic programs. Mathematics of Operations Research , 20(4):769--789
1995
-
[24]
and Wang, X
Planiden, C. and Wang, X. (2016a). Most convex functions have unique minimizers. Journal of Convex Analysis , 23(3):877--892
2016
-
[25]
and Wang, X
Planiden, C. and Wang, X. (2016b). Strongly convex functions, moreau envelopes, and the generic nature of convex functions with strong minimizers. SIAM Journal on Optimization , 26(2):1341--1364
2016
-
[26]
Rockafellar, R. T. (1970). Convex Analysis . Princeton University Press, Princeton
1970
-
[27]
and Wets, R
Rockafeller, R. and Wets, R. J. (1998). Variational analysis. Grundlehren der mathematischen Wissenschaften , 317
1998
-
[28]
Rousseeuw, P. J. (1984). Least median of squares regression. Journal of the American statistical association , 79(388):871--880
1984
-
[29]
and Wets, R
Salinetti, G. and Wets, R. J.-B. (1986). On the convergence in distribution of measurable multifunctions (random sets) normal integrands, stochastic processes and stochastic infima. Mathematics of Operations Research , 11(3):385--419
1986
-
[30]
Shapiro, A. (2000). On the asymptotics of constrained local m-estimators. Annals of statistics , pages 948--960
2000
-
[31]
van der Vaart, A. W. (2000). Asymptotic statistics , volume 3. Cambridge university press
2000
-
[32]
van der Vaart, A. W. and Wellner, J. A. (1996). Weak convergence . Springer
1996
-
[33]
Venter, J. H. (1967). On estimation of the mode. The Annals of Mathematical Statistics , pages 1446--1455
1967
-
[34]
Villani, C. (2021). Topics in optimal transportation , volume 58. American Mathematical Soc
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.