REVIEW 4 major objections 5 minor 2 cited by
Elements of asymptotic theory with outer probability measures
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper proves a Bernstein-von Mises theorem for posterior possibility functions, yielding asymptotically normal MAP estimates and chi-squared credibility tests.
desk verdict The possibilistic Bernstein-von Mises theorem is a real contribution, but the asymptotic normality and chi-squared results rest on an unproven Slutsky lemma that is false as stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the possibility function, a nonnegative function with supremum $1$, and the associated uncertain variable, a mapping from a sample space without a probability measure to the parameter space with a true state of nature. Possibility functions define outer measures by $\bar{P}(A)=\sup_{\theta\in A} f(\theta)$, and Bayes' rule takes the form $f_\theta(\theta\mid y_{1:n})=L_n(\theta)f_\theta(\theta)/\sup_{\psi}L_n(\psi)f_\theta(\psi)$. Expectation is the argmax, $E_*(x)=\arg\max f_x$, and variance is $V_*(x)=-1/f''_x(E_*(x))$, so the normal possibility function $N(\mu,\sigma^2)=\exp(-(\theta-\mu)^2/(2\sigma^2))$ replaces the normal density. Two limit theorems do the work: the law of large numbers for uncertain variables drives the sample mean's possibility function to the convex hull of the argmax, and the central limit theorem sends $n^{-1/2}\sum_i(x_i-\mu)$ to $N(0,1/(-f''_x(\mu)))$. These give the score and observed-information convergences that feed the Bernstein-von Mises proof.
What would settle it
Take the paper's own ratio-of-means setting, simulate many datasets from the stated normal laws, and compare the empirical distribution of $\sqrt{n}(\theta^*_n-\theta_0)$ with the predicted normal possibility function $N(0, V_*(s_{\theta_0}(y))/I_*(\theta_0)^2)$; if the mismatch persists as $n$ grows, Theorem 4.3 is wrong. More directly, exhibit two sequences of uncertain variables satisfying the convergence hypotheses of Proposition B.2 for which $x_n/z_n$ does not converge in outer probability to $x/\alpha$; since the lemma is asserted without proof, such a counterexample would invalidate the proof chain of Theorems 4.3 and 4.4.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that asymptotic statistics is not tied to the additivity of probability. Theorem 4.1 states that under Assumptions A.1 and A.2, for large $n$, $$f_\$\theta$(\$\theta$\mid y_{1:n})\approx N\left(\$\theta$;\theta_0+\frac{\Delta_n}{\sqrt{n}},\frac{1}{J^*_n}\right),$$ where $\Delta_n=\sqrt{n}\,\partial_\theta\ell_n(\theta_0)/J^*_n$ and $J^*_n=-\partial^2_\theta\ell_n(\hat\theta_n)$ is the observed information, with $\hat\theta_n$ the maximum likelihood estimator. Consequently (Theorem 4.3) $\sqrt{n}(\theta^*_n-\theta_0)$ converges in outer probability to $N(0,\sigma^2)$ with $\sigma^2=V_*(s_{\theta_0}(y))/I_*(\theta_0)^2$, where $s_{\theta_0}(y)=\partial_\theta\ell(\theta_0;y)$ is the score and $I_*(\theta_0)=E_*(-\partial^2_\theta\ell(\theta_0;y)\mid\theta)$ the Fisher-information analogue; and (Theorem 4.4) $-2\log f_\theta(\theta_0\mid y_{1:n})$ converges to the chi-squared possibility function $\chi^2(0,V_*(s_{\theta_0}(y))/I_*(\theta_0))$. The posterior forgets the prior's shape but not the prior's nature: the likelihood can be a probability or a possibility function, yet the posterior remains a possibility function.
Load-bearing premise
The load-bearing premise is that the paper's unproved analogue of Slutsky's lemma for uncertain variables (Proposition B.2) is true under the conditions used, including when one sequence is divided by another; the paper states that proof of this lemma is beyond its scope, so Theorems 4.3 and 4.4 collapse if the lemma fails.
Editorial extensions
If this is right
- Bayesian updating with possibility functions inherits the standard asymptotic shape of frequentist inference: large-sample posterior credibility regions centred at the MAP agree with normal intervals based on the observed information.
- A practitioner can report uncertainty about a parameter using only derivatives of the chosen log-likelihood, with no need to specify the true sampling distribution; the ratio $V_*(s_{\theta_0}(y))/I_*(\theta_0)^2$ plays the role of the asymptotic variance.
- Simple hypothesis tests $H_0:\theta=\theta_0$ can be calibrated from the chi-squared possibility limit of $-2\log f_\theta(\theta_0\mid y_{1:n})$, avoiding bootstrap or full probabilistic modelling.
- The prior's nature persists asymptotically: even after the prior's information is forgotten, a probability prior yields a probability posterior and a possibility prior yields a possibility posterior, so the choice of uncertainty representation must be made deliberately.
- Because constant possibility functions represent complete ignorance and single points can have positive credibility, the framework keeps posterior-like objects proper in cases where conventional Bayesian posteriors would be improper, such as the paper's ratio-of-means example.
Reading between the lines
- The unproved Slutsky-style lemma (Proposition B.2) is the main identifiable gap; a natural extension is to prove it under explicit regularity conditions, or to identify the extra assumptions needed for the quotient case used in normalising by observed information.
- The variance formula $V_*(s_{\theta_0}(y))/I_*(\theta_0)^2$ looks like a mode-based analogue of a sandwich variance, so the theory may extend to misspecified or quasi-likelihood settings where the true distribution is unknown; the paper does not explore this connection.
- Because the CLT for uncertain variables has a degenerate case ($V_*=\infty$ when the second derivative vanishes), mapping which parametric families hit that boundary would delimit where the Bernstein-von Mises conclusion breaks down.
- The ratio-of-means example suggests that non-integrable posterior possibility functions can still support credible-interval claims; a testable extension would check whether the predicted normal and chi-squared limits continue to hold for such heavy-tailed posteriors.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops asymptotic theory for inference with 'possibility functions', i.e. supremal outer probability measures, as an alternative to probability measures. It introduces expected value and variance for uncertain variables via the argmax and the inverse second derivative of the possibility function, proves a law of large numbers and a central limit theorem for sums of independent uncertain variables, and then states a Bernstein-von Mises theorem for posterior possibility functions. From the BvM approximation it derives asymptotic normality of the MAP estimator and a chi-squared limit for the credibility statistic of a simple hypothesis test. The paper closes with a numerical illustration for the ratio of two means and positions the framework as a bridge between frequentist and Bayesian inference.
Significance. If the main results held as stated, the paper would provide a useful asymptotic toolbox for possibility-theoretic Bayesian inference and connect the proposed formalism to Fisher information, observed information, and likelihood-ratio-style testing. The paper is clear in its motivation, contains several worked examples, and gives a detailed proof of the LLN and a substantial proof of the CLT. It is also transparent in flagging Proposition B.2 as unproved. However, the two headline results of Section 4 — asymptotic normality of the MAP and the chi-squared limit of the credibility statistic — rest on an unproved and, as stated, false Slutsky lemma, and the BvM statement is only an informal '≈' with uncontrolled remainders. These issues are load-bearing and currently prevent the paper from delivering its central claims.
major comments (4)
- [Appendix B, Proposition B.2] Proposition B.2 is stated without proof ('the proof ... is beyond the scope of this work') and is false as stated. A counterexample to the product assertion: let f_{x_n}(t)=max{e^{-t^2}, 1_{t=n}} and f_x(t)=e^{-t^2}; then x_n o.p.m.→x. Let f_{z_n}(t)=1_{t=1/n}; then z_n c→0. The product has possibility f_{x_n z_n}(s)=max{e^{-n^2 s^2}, 1_{s=1}}, which converges pointwise to a function equal to 1 at s=0 and s=1 and 0 elsewhere, not to 1_{s=0}. Hence x_n z_n does not converge to 0·x. Theorems 4.3 and 4.4 use exactly this product/quotient structure, as seen in the proofs around (B.5) and (B.6). A corrected, restricted version of the lemma with explicit conditions and a full proof is required before these theorems can be accepted.
- [Section 4.1, Theorem 4.1 and proof around (B.2)-(B.4)] Theorem 4.1 is stated only with the symbol '≈', and its proof is a Taylor expansion without control of the remainder terms or of uniformity in ψ. In particular, dropping the third-order term in the denominator of (B.2) requires more than Assumption A.1; one needs a bound on n^{-1}(θhat_n−θ0)∂^3_θ l_n(ψ_n), and if Assumption A.5 is invoked for that purpose it should be stated. Similarly, replacing the denominator by L_n(θhat_n) and neglecting the prior f_θ in (B.4) requires sup_{|ψ|≤M} |ψ|^3/n^{3/2} ∂^3 l = o(1) and a comparable prior contribution, none of which is proved. As written, the BvM result is an informal approximation, and Theorems 4.3 and 4.4 inherit this informality.
- [Section 4.2, Theorems 4.3 and 4.4] The convergence statements in Theorems 4.3 and 4.4 are in outer probability measure induced by the possibility function f_{y|θ}(·|θ0) for uncertain observations, not under the true sampling distribution p_Y of the random variables Y_i. The text motivates these results as 'frequentist-type guarantees', but the o.p.m. convergence of possibility functions does not by itself imply coverage properties or distributional convergence for the actual random data-generating mechanism. If the intended claim is the usual repeated-sampling behaviour, the relationship between the o.p.m. model and the true random law must be made explicit and proved; otherwise the terminology 'frequentist' is misleading.
- [Section 3.1, Eq. (3.2) and Theorem 3.3] The variance V*(x) is defined as −1/f_x''(E*(x)), and Theorem 3.3 then returns exactly this quantity as the limiting spread of the rescaled sum. The CLT is therefore close to a tautology: it shows that under strict log-concavity the limiting possibility function is determined by the local curvature of f_x at its mode. This does not invalidate the theorem, but the paper should acknowledge explicitly that the 'asymptotic variance' is a consequence of the definition of V*, not an independently derived quantity. The same observation applies to Theorem 4.3, where σ² = V*(s_θ0(y))/I*(θ0)² is a ratio of second derivatives by construction.
minor comments (5)
- [Section 4.2.2] The construction of the α-credible interval [a,b] assumes unimodality of the posterior possibility function, but this assumption is only mentioned in passing; it should be stated precisely in the theorem or proposition that uses it.
- [Figure 1] The caption does not explain how the 'truth' line is computed, what the averaging is over, or how the standard deviation bands are constructed; these details should be added for reproducibility.
- [Proof of Theorem 3.3] The claim that the bordered-Hessian check plus the stationarity condition identifies the global maximum is sketched rather than proved; the argument that any solution of (log f(x_i))' = (log f(x_j))' must have x_i = x_j under strict log-concavity deserves a short explicit proof.
- [Section 2, sufficient statistics] The decomposition (2.2) is stated as a factorisation of possibility functions, but the notation f_{y|t,θ} and f_{y|t} is introduced only informally; a precise definition of conditional possibility functions would help.
- [Various] There are minor typographical issues, including inconsistent use of 'Students model' for 'Student's model', and the phrase 'where the MLE is consistent' in Assumption A.1 is not a formal condition; it should be replaced by explicit regularity conditions.
Circularity Check
No significant circularity: the asymptotic variance statements are internal Taylor/Laplace identities, and the main unsupported lemma (Proposition B.2) is a correctness gap, not a circular derivation.
full rationale
The central results are internal asymptotic expansions rather than data-fitted predictions. Theorem 3.1 and Theorem 3.3 prove pointwise limiting behaviour for sums of uncertain variables; the paper then defines E* as the argmax and V* in (3.2) as the inverse curvature that those theorems identify. Because the definitions are introduced after, and explicitly motivated by, the LLN/CLT, the fact that the CLT limit has spread V* is a consistency of notation, not a circular prediction. Similarly, Theorem 4.1 is a second-order Taylor/Laplace expansion of the normalized likelihood: J*_n is defined in (4.1) as -d^2_theta l_n(theta_hat_n), and the posterior curvature is asymptotically the same quantity; the theorem does not fit a parameter and then predict it. Theorems 4.3 and 4.4 combine this expansion with the LLN/CLT and with Proposition B.2. That proposition is explicitly stated without proof in the appendix, with the paper conceding that 'the proof of Slutsky's lemma for uncertain variables is beyond the scope of this work,' and it is load-bearing for those theorems; the skeptical counterexample indicates it may be false as stated. This is a serious completeness and correctness issue, but it is not circularity: the conclusions are not assumed or fitted, merely unsupported. The self-citations (Houssineau 2018a, 2018b; Houssineau and Bishop 2018) are background references for uncertain variables, o.p.m.s, and filtering, and they do not carry the asymptotic argument. No circularity score above 0 is therefore warranted.
Assumptions & free parameters
assumptions (5)
- domain assumption A.1: The parameter theta0 is well defined and the MLE is consistent.
- domain assumption A.3: fy is strictly log-concave.
- domain assumption A.5: The third derivative of the log-likelihood is O(n) under the o.p.m.
- ad hoc to paper Proposition B.2: Slutsky's lemma for uncertain variables, stated without proof.
- standard math Taylor expansion of the log-likelihood with negligible remainder is valid.
Cite this review
Pith. "Pith review of Elements of asymptotic theory with outer probability measures." pith.science (2026). https://pith.science/paper/OTBOSCAW
@misc{pith2026190804331,
author = {Pith},
title = {Pith review of: Elements of asymptotic theory with outer probability measures},
year = {2026},
howpublished = {\url{https://pith.science/paper/OTBOSCAW}},
note = {Machine review of arXiv:1908.04331}
}
read the original abstract
Outer measures can be used for statistical inference in place of probability measures to bring flexibility in terms of model specification. The corresponding statistical procedures such as Bayesian inference, estimators or hypothesis testing need to be analysed in order to understand their behaviour, and motivate their use. In this article, we consider a class of outer measures based on the supremum of particular functions that we refer to as possibility functions. We then characterise the asymptotic behaviour of the corresponding Bayesian posterior uncertainties, from which the properties of the corresponding maximum a posteriori estimators can be deduced. These results are largely based on versions of both the law of large numbers and the central limit theorem that are adapted to possibility functions. Our motivation with outer measures is through the notion of uncertainty quantification, where verification of these procedures is of crucial importance. These introduced concepts shed a new light on some standard concepts such as the Fisher information and sufficient statistics and naturally strengthen the link between the frequentist and Bayesian approaches.
Figures
Forward citations
Cited by 2 Pith papers
-
Improving Active Learning with a Bayesian Representation of Epistemic Uncertainty
New active learning acquisition functions derived from a possibilistic representation of epistemic uncertainty match or beat standard GP-based baselines on several classification datasets.
-
Redesigning the ensemble Kalman filter with a dedicated model of epistemic uncertainty
A new ensemble Kalman filter built on possibility theory fits the tightest Gaussian possibility function to weighted particles, yielding better calibrated uncertainty estimates from small ensembles.
Reference graph
Works this paper leans on
-
[1]
J. Aldrich. R. A . F isher and the making of maximum likelihood 1912--1922. Statistical science, 12 0 (3): 0 162--176, 1997
work page 1912
-
[2]
R. L. Berger and G. Casella. Statistical inference. Duxbury, 2001
work page 2001
-
[3]
P. J. Bickel, Y. Ritov, and T. Ryden. Asymptotic normality of the maximum-likelihood estimator for general hidden M arkov models. The Annals of Statistics, 26 0 (4): 0 1614--1635, 1998
work page 1998
-
[4]
P. G. Bissiri, C. C. Holmes, and S. G. Walker. A general framework for updating belief distributions. Journal of the Royal Statistical Society: Series B, 78 0 (5): 0 1103--1130, 2016
2016
- [5]
-
[6]
C. Carlsson and R. Full \'e r. On possibilistic mean value and variance of fuzzy numbers. Fuzzy sets and systems, 122 0 (2): 0 315--326, 2001
work page 2001
-
[7]
Y. Y. Chen. Statistical inference based on the possibility and belief measures. Transactions of the American Mathematical Society, 347 0 (5): 0 1855--1863, 1995
work page 1995
-
[8]
P. Constantinou and A. P. Dawid. Extended conditional independence and applications in causal inference. The Annals of Statistics, 45 0 (6): 0 2618--2653, 2017
work page 2017
Show all 51 references
-
[9]
Dashti, K
M. Dashti, K. J. H. Law, A. M. Stuart, and J. Voss. MAP estimators and their consistency in B ayesian nonparametric inverse problems. Inverse Problems, 29 0 (9): 0 095017, 2013
2013
-
[10]
De Baets, E
B. De Baets, E. Tsiporkova, and R. Mesiar. Conditioning in possibility theory with strict order norms. Fuzzy Sets and Systems, 106 0 (2): 0 221--229, 1999
1999
-
[11]
Del Moral and M
P. Del Moral and M. Doisy. Maslov idempotent probability calculus, i. Theory of Probability & Its Applications, 43 0 (4): 0 562--576, 1999
1999
-
[12]
Del Moral and M
P. Del Moral and M. Doisy. Maslov idempotent probability calculus. ii. Theory of Probability & Its Applications, 44 0 (2): 0 319--332, 2000
2000
-
[13]
A. P. Dempster. A generalization of bayesian inference. Journal of the Royal Statistical Society: Series B (Methodological), 30 0 (2): 0 205--232, 1968
1968
-
[14]
Druilhet and J.-M
P. Druilhet and J.-M. Marin. Invariant HPD credible sets and MAP estimators. Bayesian Analysis, 2 0 (4): 0 681--691, 2007
2007
-
[15]
Dubois and H
D. Dubois and H. Prade. Additions of interactive fuzzy numbers. IEEE Transactions on Automatic Control, 26 0 (4): 0 926--936, 1981
1981
-
[16]
Dubois and H
D. Dubois and H. Prade. Possibility theory and its applications: Where do we stand? In Springer Handbook of Computational Intelligence, pages 31--60. Springer, 2015
2015
-
[17]
Dubois, S
D. Dubois, S. Moral, and H. Prade. A semantics for possibility theory based on likelihoods. In Proceedings of 1995 IEEE International Conference on Fuzzy Systems, volume 3, pages 1597--1604. IEEE, 1995
1995
-
[18]
H. W. Engl, M. Hanke, and A. Neubauer. Regularization of inverse problems, volume 375. Springer Science & Business Media, 1996
1996
-
[19]
Evans and H
M. Evans and H. Moshonov. Checking for prior-data conflict. Bayesian analysis, 1 0 (4): 0 893--914, 2006
2006
-
[20]
R. A. Fisher. On the mathematical foundations of theoretical statistics. Phil. Trans. R. Soc. Lond. A, 222 0 (594-604): 0 309--368, 1922
1922
-
[21]
R. A. Fisher. The fiducial argument in statistical inference. Annals of eugenics, 6 0 (4): 0 391--398, 1935
1935
-
[22]
Gelman, J
A. Gelman, J. B. Carlin, H. S. Stern, D. B. Dunson, A. Vehtari, and D. B. Rubin. Bayesian data analysis. CRC press, 2013
2013
-
[23]
V. P. Godambe. Estimating functions. Oxford University Press, 1991
1991
-
[24]
P. D. Gr \"u nwald and A. P. Dawid. Game theory, maximum entropy, minimum discrepancy and robust bayesian decision theory. the Annals of Statistics, 32 0 (4): 0 1367--1433, 2004
2004
-
[25]
A. Hald. On the history of maximum likelihood in relation to inverse probability and least squares. Statistical Science, 14 0 (2): 0 214--222, 1999
1999
-
[26]
M.-A. Henn, H. Gross, F. Scholze, M. Wurm, C. Elster, and M. B \"a r. A maximum likelihood approach to the inverse problem of scatterometry. Optics Express, 20 0 (12): 0 12771--12786, 2012
2012
-
[27]
Houssineau
J. Houssineau. Parameter estimation with a class of outer probability measures. arXiv preprint arXiv:1801.00569, 2018 a
2018 arXiv
-
[28]
Houssineau
J. Houssineau. Detection and estimation of partially-observed dynamical systems: an outer-measure approach. arXiv preprint arXiv:1801.00571, 2018 b
2018 arXiv
-
[29]
Houssineau and A
J. Houssineau and A. N. Bishop. Smoothing and filtering with a class of outer measures. SIAM/ASA Journal on Uncertainty Quantification, 6 0 (2): 0 845--866, 2018
2018
-
[30]
Langford
J. Langford. Tutorial on practical prediction theory for classification. Journal of machine learning research, 6 0 (Mar): 0 273--306, 2005
2005
-
[31]
L. Le Cam. Maximum likelihood: an introduction. International Statistical Review/Revue Internationale de Statistique, pages 153--171, 1990
1990
-
[32]
E. L. Lehmann. Elements of large-sample theory. Springer Science & Business Media, 2004
2004
-
[33]
Maccheroni, M
F. Maccheroni, M. Marinacci, et al. A strong law of large numbers for capacities. The Annals of Probability, 33 0 (3): 0 1171--1178, 2005
2005
-
[34]
Marinacci
M. Marinacci. Limit laws for non-additive probabilities and their frequentist interpretation. Journal of Economic Theory, 84 0 (2): 0 145--195, 1999
1999
-
[35]
V. P. Maslov. Idempotent analysis, volume 13. American Mathematical Soc., 1992
1992
-
[36]
McGoff, S
K. McGoff, S. Mukherjee, A. Nobel, N. Pillai, et al. Consistency of maximum likelihood estimation for some dynamical systems. The Annals of Statistics, 43 0 (1): 0 1--29, 2015
2015
-
[37]
S. A. Murphy and A. W. Van der Vaart. On profile likelihood. Journal of the American Statistical Association, 95 0 (450): 0 449--465, 2000
2000
-
[38]
A. Raue, C. Kreutz, T. Maiwald, J. Bachmann, M. Schilling, U. Klingm \"u ller, and J. Timmer. Structural and practical identifiability analysis of partially observed dynamical models by exploiting the profile likelihood. Bioinformatics, 25 0 (15): 0 1923--1929, 2009
1923
-
[39]
Robert and G
C. Robert and G. Casella. Monte C arlo statistical methods . Springer Science & Business Media, 2013
2013
-
[40]
G. L. S. Shackle. Decision order and time in human affairs. Cambridge University Press, 1961
1961
-
[41]
G. Shafer. A mathematical theory of evidence, volume 42. Princeton University Press, 1976
1976
-
[42]
Smets and R
P. Smets and R. Kennes. The transferable belief model. Artificial intelligence, 66 0 (2): 0 191--234, 1994
1994
-
[43]
J. Q. Smith. Bayesian decision analysis: principles and practice. Cambridge University Press, 2010
2010
-
[44]
A. M. Stuart. Inverse problems: a B ayesian perspective. Acta Numerica, 19: 0 451--559, 2010
2010
-
[45]
Ter \'a n
P. Ter \'a n. Law of large numbers for the possibilistic mean value. Fuzzy Sets and Systems, 245: 0 116--124, 2014
2014
-
[46]
Vanlier, C
J. Vanlier, C. A. Tiemann, P. A. J. Hilbers, and N. A. W. van Riel. An integrated strategy for prediction uncertainty analysis. Bioinformatics, 28 0 (8): 0 1130--1135, 2012
2012
-
[47]
A. Wald. Note on the consistency of the maximum likelihood estimate. The Annals of Mathematical Statistics, 20 0 (4): 0 595--601, 1949
1949
-
[48]
P. Walley. Statistical reasoning with imprecise probabilities. Chapman & Hall, 1991
1991
-
[49]
Walley and S
P. Walley and S. Moral. Upper probabilities based only on the likelihood function. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 61 0 (4): 0 831--847, 1999
1999
-
[50]
R. W. Wedderburn. Quasi-likelihood functions, generalized linear models, and the G auss- N ewton method. Biometrika, 61 0 (3): 0 439--447, 1974
1974
-
[51]
L. A. Zadeh. Fuzzy sets as a basis for a theory of possibility. Fuzzy sets and systems, 1 0 (1): 0 3--28, 1978
1978
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.