Pith. sign in

REVIEW 3 major objections 4 minor 34 references

On the Symmetry of Limiting Distributions of M-estimators

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper establishes that an M-estimator's limiting distribution is symmetric around zero whenever the constraint set, the deterministic drift, and the stochastic process are all even in law, and proves for linear-process limits that…

desk verdict Fix the statement of Theorem 3.2 and the proof of Lemma 3.1; the symmetry results are otherwise credible and useful. read the letter →

arxiv 2411.17087 v1 pith:CMW42HYW submitted 2024-11-26 math.ST stat.MEstat.TH

classification math.STstat.MEstat.TH MSC 62F1262F2562G2060F17
keywords M-estimatorslimitingdistributionssymmetrymedianbiasHulCevenstochasticprocesstangentconecube-rootasymptotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks when the limiting distribution of an M-estimator—the scaled error of a minimizer of an empirical loss—is symmetric about zero. Its main sufficient condition is that the constraint cone is symmetric, the deterministic drift $D$ is even, and the mean-zero stochastic process $S$ has the same law as its reflection, in which case the argmin $W$ of $D+S+X$ satisfies $W \stackrel{d}{=} -W$. For linear-process limits, with a symmetric, absolutely continuous, full-support random vector $Y$, the condition is also necessary, not just sufficient. Symmetric limits make under- and over-estimation equally likely, so the limiting median bias is zero—exactly what the HulC confidence-interval construction (a method that combines independent estimators) needs—and the paper extends valid inference to irregular, constrained, and cube-root-rate problems where normality and continuity-based methods do not apply.

What carries the argument

The load-bearing object is the limiting stochastic process $Z(u)=D(u)+S(u)+X(u)$ from display (5), where $X$ is $0$ on the tangent cone of the constraint set at the true parameter and $+\infty$ outside it. A process is called even when it and its reflection have the same law; Lemma 3.1 and Theorem 3.2 transfer componentwise evenness plus almost-sure uniqueness of the argmin into $W \stackrel{d}{=} -W$. For the necessity direction, the argument shifts to convex duality: the minimizer $W$ is identified as a subgradient of the convex conjugate of $D+X$ evaluated at $-Y$, and uniqueness of optimal transport maps forces two such subgradient selections to agree almost everywhere, which yields evenness of $D+X$. Epi-convergence supplies the bridge from finite-sample minimizers to the limiting argmin.

What would settle it

Choose a one-dimensional linear-process limit $Z(u)=D(u)+uY$ with $Y$ standard normal and a superlinear convex $D$ that is not even, and compute the distribution of $W=\arg\min Z(u)$; Theorem 4.7 predicts this distribution is asymmetric whenever $D$ is not even, so an observed symmetric distribution in this setup would refute the necessity claim.

Watch

Extended reading notes

Core claim

The central claim is that the symmetry of a limiting argmin distribution is governed by the symmetry of the limit process's ingredients, not by the particular loss or estimator. Theorem 3.2 shows that if $Z(u)=D(u)+S(u)+X(u)$ and $Z(-u)$ have the same law and the argmin is almost surely unique, the random minimizer is symmetric about zero. Theorems 4.4 and 4.7 give the converse for the linear class $Z(u)=D(u)+\langle u,Y\rangle$ when $Y$ is symmetric, absolutely continuous, and dominates Lebesgue measure: then a symmetric $W$ forces $D+X$ to be even, so for Class I the paper obtains a complete characterisation. For Class II Gaussian-process limits and Class III Poisson-hyperplane limits, the same theorem supplies checkable sufficient conditions, which the paper verifies for shorth, least-median-of-squares, mode, bridge, LAD, and constrained regression examples.

Load-bearing premise

The load-bearing premise is that the normalized estimator converges in distribution to the almost surely unique minimizer of $D+S+X$, since uniqueness is what lets symmetry of the ingredients be transferred to the argmin, and for Classes II and III this uniqueness is assumed externally for the examples.

Editorial extensions

If this is right

  • For cube-root-rate estimators such as shorth, least median of squares, and mode estimation with the tuning parameter below the boundary case, the Gaussian-process covariance is automatically even, so checking the drift and constraint cone decides symmetry of the limit.
  • For LAD regression, symmetry holds when the error distribution's point-mass increments just above and below zero balance in the limit; the paper's examples show asymmetric error increments lead to asymmetric limits and nonzero median bias.
  • For bridge and LASSO-type penalized regression, concave penalties ($\mu<1$) give even drift at every true parameter, while convex penalties ($\mu\ge 1$) give even drift only at $\theta_0=0$; uniform symmetric inference over the parameter space therefore cannot be certified for the convex case.
  • For constrained mean estimation, the tangent cone must be a symmetric set; when the true parameter lies on a boundary with an asymmetric cone, the limiting distribution is asymmetric even though the errors are Gaussian.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because Theorem 3.2 is stated for any even-in-law process, the same symmetry argument should apply to resampling or bootstrap versions of these limits; any bootstrap scheme that preserves componentwise evenness would inherit the symmetric law.
  • The paper leaves open whether a non-even drift outside the support of $Y$ can still produce a symmetric $W$; its own remarks suggest the true boundary is equality of two conjugates on the support of $Y$, which could be tested by constructing asymmetric $D$ that agree on that support.
  • Median unbiasedness holds under weaker conditions than full symmetry (Theorem 4.5), so inference procedures built on vanishing median bias rather than distributional symmetry could cover irregular cases where the symmetric-limit condition fails.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies conditions under which the limiting distributions of M-estimators, represented as minimizers of processes Z(u)=D(u)+S(u)+X(u), are symmetric around zero. Symmetry of the limiting distribution is used to guarantee zero limiting median bias, which is the requirement for valid inference via the HulC method of Kuchibhotla et al. (2024). The paper gives a general sufficient condition (Theorem 3.2) based on evenness of the deterministic drift D, the constraint indicator X, and the stochastic process S, and verifies this condition for Pflug's three classes of limiting processes: linear stochastic processes, Gaussian processes, and generalized Poisson hyperplane processes. It then proves necessity of the drift/constraint evenness assumption (A1) for Class I limits in one dimension (Theorem 4.4) and in higher dimensions under stronger symmetry of Y (Theorem 4.7), using convex analysis and optimal transport uniqueness.

Significance. If the main results stand, the paper delivers a useful and nontrivial contribution: it identifies a simple, checkable sufficient condition for symmetry of non-normal M-estimator limits, and it shows in an important special case that the condition is also necessary. The verification of evenness for several classical irregular estimators (shorth, LMS, mode, LAD, bridge) and the explicit discussion of when (A1) fails (e.g., Example 2.7, mode estimation at the boundary tuning parameter) give practitioners concrete guidance for when HulC is or is not applicable. The necessity proofs through optimal transport are elegant and appear internally consistent. However, the central sufficiency theorem and its supporting lemma are not stated correctly, as detailed below.

major comments (3)
  1. [Lemma 3.1] The statement P(W in A) = P(inf_A Z < inf_{A^c} Z) is false even under the stated almost-sure uniqueness assumption. For the deterministic even process Z(u)=u^4 on R, the unique minimizer is W=0; taking A={0}, the left-hand side equals 1 while the right-hand side is P(0<0)=0. The conclusion W d= -W is nonetheless true and can be proved by the equivariance of the argmin map under u -> -u, but the printed proof is invalid as written.
  2. [Theorem 3.2] Theorem 3.2 omits the almost-sure uniqueness hypothesis that Lemma 3.1 requires. As stated, the theorem is false: take Omega=R, D(u)=0, S(u)=0, and X the indicator of R. Every real number is a minimizer, the process is even, and the deterministic selection W=1 is not symmetric around zero. The theorem should either assume that argmin(D+S+X) is almost surely unique or specify a symmetric tie-breaking rule that makes W a well-defined random variable; the current wording does not.
  3. [Section 4.1 and Example 2.7] The missing uniqueness hypothesis is not merely cosmetic: the paper proves uniqueness for parts of Class I in Theorem 4.3, but for Classes II and III it does not provide general uniqueness results, and Example 2.7 explicitly conditions on the limiting process attaining a unique minimum almost surely as an external assumption. Since Theorem 3.2 is invoked for these classes, the manuscript needs to state explicitly which uniqueness condition is being assumed in each application, or to restrict the sufficiency claim accordingly.
minor comments (4)
  1. [Section 1] There is a typographical artifact in the last paragraph of the introduction: 'min36imizers' should read 'minimizers'.
  2. [References] The spelling 'Rockafeller' is used inconsistently; the standard spelling is 'Rockafellar'.
  3. [Theorem 3.2] The phrase 'W = argmin_{u in Omega} D(u)+S(u)+X(u)' should specify that Omega is the effective domain and that the argmin is taken almost surely unique, to avoid ambiguity with the counterexample above.
  4. [Section 4.2] In the proof of Theorem 4.4, the sentence 'the derivatives here exist almost surely from almost sure uniqueness of minimizers' would benefit from a brief justification or a reference, since differentiability of the conjugate at the relevant point is what is being used.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the symmetry theorems are self-contained mathematical derivations; self-citations to HulC are motivational, not load-bearing.

full rationale

The paper's central derivation chain is self-contained. Theorem 3.2 takes as hypotheses that the constraint set is symmetric, the deterministic drift D is even, and the stochastic process S is even, and concludes that the unique argmin W is symmetric around zero. This is a direct implication; the conclusion is not used to define or fit the hypotheses. The necessity results (Theorems 4.4 and 4.7) assume symmetric, absolutely continuous random shifts and prove that (A1) must hold for W to be symmetric, using convex analysis and optimal transport rather than assuming the target symmetry. No parameter is fitted to data, and no closely related quantity is predicted from a fitted subset. The citations to Kuchibhotla et al. (2023, 2024) and Jain and Kuchibhotla (2023) introduce HulC and median-bias regularity as motivating applications and external validity conditions, but these results do not enter the proof of the symmetry theorems; removing them would leave Theorems 3.2, 4.3, 4.4, and 4.7 intact. The paper explicitly flags which minimizer-uniqueness conditions are imposed externally, as in Example 2.7 where it states 'provided that the limiting process attains a unique minimum almost surely.' Any concerns about the correctness of the Lemma 3.1 proof or about omitted uniqueness hypotheses in Theorem 3.2 are mathematical correctness issues, not circularity, because they do not involve the derivation reducing to its own inputs or to a self-citation chain. Accordingly, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper's conclusions rest on standard variational-analysis machinery (epi-convergence, tangent cones, convex conjugates), the explicit evenness assumption (A1), and the almost-sure uniqueness of the limiting argmin. No free parameters are fitted to data, and no new physical or probabilistic entities are postulated beyond the generalized Poisson hyperplane process, which is presented as a definitional extension rather than an ad hoc construct.

assumptions (5)
  • domain assumption Epi-convergence and argmin convergence in (5): the normalized M-estimator converges in distribution to the argmin of D+S+X.
    This is the standard regularity bridge from finite-sample estimators to the limiting process studied throughout the paper, following Pflug (1995), Geyer (1994), and Knight (1999b).
  • domain assumption The argmin of the limit process is almost surely unique.
    Lemma 3.1 requires uniqueness to equate the event {W in A} with strict inequality of infima over A and A^c. Sufficient conditions for uniqueness are given only for the linear class in Theorem 4.3.
  • standard math Legendre-Fenchel duality, greatest convex minorants, subdifferential calculus, and differentiability of convex conjugates.
    Used throughout Section 4 and the supplement, including the representation W = (D+X)*(-Y) and the measurable selection argument in Lemma 4.6.
  • standard math Optimal transport uniqueness theorems for absolutely continuous measures, including Brenier-McCann and the one-dimensional result of Ambrosio.
    These results are the key step in Theorems 4.4 and 4.7, converting equality in distribution of two monotone or gradient maps into equality almost everywhere.
  • ad hoc to paper Assumption (A1): the deterministic drift D plus constraint indicator X is even.
    This is the paper's defining condition rather than a standard background result. The paper shows it is sufficient for symmetry in all three classes and necessary for the linear class under stated assumptions, and it fails for bridge estimators with nonzero target and mu >= 1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Symmetry of Limiting Distributions of M-estimators." pith.science (2026). https://pith.science/paper/CMW42HYW

@misc{pith2026241117087,
  author       = {Pith},
  title        = {Pith review of: On the Symmetry of Limiting Distributions of M-estimators},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CMW42HYW}},
  note         = {Machine review of arXiv:2411.17087}
}
abstract

Many functionals of interest in statistics and machine learning can be written as minimizers of expected loss functions. Such functionals are called $M$-estimands, and can be estimated by $M$-estimators -- minimizers of empirical average losses. Traditionally, statistical inference (e.g., hypothesis tests and confidence sets) for $M$-estimands is obtained by proving asymptotic normality of $M$-estimators centered at the target. However, asymptotic normality is only one of several possible limiting distributions and (asymptotically) valid inference becomes significantly difficult with non-normal limits. In this paper, we provide conditions for the symmetry of three general classes of limiting distributions, enabling inference using HulC (Kuchibhotla et al. (2024)).

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

34 extracted references · 31 canonical work pages

  1. [1]

    Ambrosio, L. (2003). Lecture notes on optimal transport problems. In Mathematical Aspects of Evolving Interfaces , volume 1812 of Lecture Notes in Mathematics . Springer, Berlin, Heidelberg

  2. [2]

    a user classics. Birkh \

    Aubin, J.-P. and Frankowska, H. (2009). Set-valued analysis. modern birkh \"a user classics. Birkh \"a user Boston, Inc., Boston, MA. https://doi.org/10.1007/978-0-8176-4848-0 , pages XXI, 461

  3. [3]

    Bonnans, J. F. and Shapiro, A. (2013). Perturbation analysis of optimization problems . Springer Science & Business Media

  4. [4]

    Burke, J. V. and Hoheisel, T. (2017). Epi-convergence properties of smoothing by infimal convolution. Set-Valued and Variational Analysis , 25:1--23

  5. [5]

    and Hong, H

    Chernozhukov, V. and Hong, H. (2002). Likelihood inference in a class of nonregular econometric models. Available at SSRN 301486

  6. [6]

    Geyer, C. J. (1994). On the asymptotics of constrained m-estimation. The Annals of statistics , pages 1993--2010

  7. [7]

    Hess, C. (1996). Epi-convergence of sequences of normal integrands and strong consistency of the maximum likelihood estimator. The Annals of Statistics , 24(3):1298--1315

  8. [8]

    and Lemar \'e chal, C

    Hiriart-Urruty, J.-B. and Lemar \'e chal, C. (2004). Fundamentals of convex analysis . Springer Science & Business Media

Show all 34 references
  1. [9]

    and Kuchibhotla, A

    Jain, A. and Kuchibhotla, A. K. (2023). Rectangular hull confidence regions for multivariate parameters. arXiv preprint arXiv:2311.16598

  2. [10]

    Kall, P. (1986). Approximation to optimization problems: An elementary review. Mathematics of Operations Research , 11(1):9--18

  3. [11]

    Kaufmann, H. (1988). On existence and uniqueness of a vector minimizing a convex function. Zeitschrift f \"u r Operations Research , 32:357--373

  4. [12]

    and Pollard, D

    Kim, J. and Pollard, D. (1990). Cube root asymptotics. The Annals of Statistics , pages 191--219

  5. [13]

    Knight, K. (1998). Limiting distributions for l1 regression estimators under general conditions. The Annals of Statistics , 26(2):755--770

  6. [14]

    Knight, K. (1999a). Asymptotics for l1-estimators of regression parameters under heteroscedasticityy. Canadian Journal of Statistics , 27(3):497--507

  7. [15]

    Knight, K. (1999b). Epi-convergence in distribution and stochastic equi-semicontinuity. Unpublished manuscript , 37(7):14

  8. [16]

    and Fu, W

    Knight, K. and Fu, W. (2000). Asymptotics for lasso-type estimators. The Annals of statistics , 28(5):1356--1378

  9. [17]

    K., Balakrishnan, S., and Wasserman, L

    Kuchibhotla, A. K., Balakrishnan, S., and Wasserman, L. (2023). Median regularity and honest inference. Biometrika , 110(3):831--838

  10. [18]

    K., Balakrishnan, S., and Wasserman, L

    Kuchibhotla, A. K., Balakrishnan, S., and Wasserman, L. (2024). The HulC: confidence regions from convex hulls . Journal of the Royal Statistical Society Series B: Statistical Methodology , 86(3):586--622

  11. [19]

    and Wright, S

    Nocedal, J. and Wright, S. J. (2006). Numerical Optimization . Springer Series in Operations Research and Financial Engineering. Springer New York, NY, 2 edition

  12. [20]

    Parthasarathy, T. (1972). General theorems on selectors , pages 49--57. Springer Berlin Heidelberg, Berlin, Heidelberg

  13. [21]

    Pensia, A., Jog, V., and Loh, P.-L. (2019). Mean estimation for entangled single-sample distributions. In 2019 IEEE International Symposium on Information Theory (ISIT) , pages 3052--3056. IEEE

  14. [22]

    Pfanzagl, J. (1970). On the asymptotic efficiency of median unbiased estimates. The Annals of Mathematical Statistics , pages 1500--1509

  15. [23]

    Pflug, G. C. (1995). Asymptotic stochastic programs. Mathematics of Operations Research , 20(4):769--789

  16. [24]

    and Wang, X

    Planiden, C. and Wang, X. (2016a). Most convex functions have unique minimizers. Journal of Convex Analysis , 23(3):877--892

  17. [25]

    and Wang, X

    Planiden, C. and Wang, X. (2016b). Strongly convex functions, moreau envelopes, and the generic nature of convex functions with strong minimizers. SIAM Journal on Optimization , 26(2):1341--1364

  18. [26]

    Rockafellar, R. T. (1970). Convex Analysis . Princeton University Press, Princeton

  19. [27]

    and Wets, R

    Rockafeller, R. and Wets, R. J. (1998). Variational analysis. Grundlehren der mathematischen Wissenschaften , 317

  20. [28]

    Rousseeuw, P. J. (1984). Least median of squares regression. Journal of the American statistical association , 79(388):871--880

  21. [29]

    and Wets, R

    Salinetti, G. and Wets, R. J.-B. (1986). On the convergence in distribution of measurable multifunctions (random sets) normal integrands, stochastic processes and stochastic infima. Mathematics of Operations Research , 11(3):385--419

  22. [30]

    Shapiro, A. (2000). On the asymptotics of constrained local m-estimators. Annals of statistics , pages 948--960

  23. [31]

    van der Vaart, A. W. (2000). Asymptotic statistics , volume 3. Cambridge university press

  24. [32]

    van der Vaart, A. W. and Wellner, J. A. (1996). Weak convergence . Springer

  25. [33]

    Venter, J. H. (1967). On estimation of the mode. The Annals of Mathematical Statistics , pages 1446--1455

  26. [34]

    Villani, C. (2021). Topics in optimal transportation , volume 58. American Mathematical Soc

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.