REVIEW 3 major objections 4 minor 42 references
Generalised Robust Bayes for Joint Inference of Model and Contamination
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Hölder-Bayes makes the contamination proportion an inferable parameter in robust Bayesian updating.
desk verdict Worth engaging with, but the headline finite-sample excess-risk theorem has a real quantifier gap and the paper needs a revision before the theory is trustworthy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Hölder divergence applied to a scaled model: the empirical risk H_n^(γ)(θ,α)=ϕ(α^{-1}S_{n,γ}(θ)/C_γ(θ)) α^{1+γ}C_γ(θ), where S_{n,γ}(θ) is the empirical average of pθ^γ and C_γ(θ) is its model expectation. The scaling parameter α acts as the inlier proportion and is inferred jointly with θ. The theoretical analysis is carried by the tail-overlap quantity ργ(θ)=∫pθ^γ dQ0, which controls the bias in both the population minimiser (Theorem 3) and the asymptotic distribution (Theorem 5).
What would settle it
Simulate data from a contamination model where Q0 has substantial density at the centre of Pθ0 (e.g., a Gaussian outlier centred at the true mean with the same scale) and check whether the Hölder posterior for α recovers the injected ϵ0; if the posterior is sharply biased away from ϵ0, the identifiability mechanism breaks.
Extended reading notes
Core claim
The paper introduces the Hölder posterior as a generalised posterior over (θ, α) formed by exponentiating the empirical Hölder divergence between the data and the scaled model αPθ. The key structural insight is that, when the outlier distribution Q0 has small tail overlap with the target model density pθ0—quantified by ργ(θ0)=∫pθ0^γ dQ0—the ratio Sn,γ(θ)/Cγ(θ) is dominated by the clean component, so the risk-minimising α approximates the true inlier proportion α0. The paper proves global bias-robustness via a bounded posterior influence function, establishes a finite-sample excess-risk bound at rate n^{-1/2}, and derives a Bernstein–von Mises theorem showing asymptotic normality with covaria
Load-bearing premise
The outliers must be concentrated in the low-density tails of the target model, so that the overlap term ργ(θ0) is negligible; otherwise the contamination proportion is not identifiable and the bias bound in Theorem 3 can be large.
Editorial extensions
If this is right
- Robust Bayesian inference can now report a coherent posterior over the contamination proportion rather than a single robustness guarantee, enabling principled uncertainty quantification of data quality.
- Outlier detection becomes part of the same probabilistic model that fits the clean signal: observation-level Frequency-of-Detection scores are produced by propagating posterior draws of (θ, α), with no external threshold required.
- The temperature parameter in generalised Bayes gains a concrete meaning as an affine scaling of the data space, suggesting that data standardisation can replace expensive temperature-tuning procedures.
- The bias bounds imply that the feasibility of heavy-contamination inference is governed by tail overlap between outliers and the model, not merely by the contamination fraction itself.
Reading between the lines
- The tail-overlap assumption could be tested in practice: on real datasets, one might estimate ργ(θ*) and compare the posterior for ϵ to an external robust estimate; large discrepancies would signal regime in which the framework is unreliable.
- The temperature-as-scaling interpretation suggests a possible extension: instead of the sample covariance, one could choose the affine transformation that minimises posterior predictive loss, making the temperature fully data-adaptive.
- The joint inference idea may extend beyond Hölder divergence to other scoring rules that are proper and sensitive to scaling; the bounded-derivative condition in Assumption 1 is what prevents scale-invariant losses from identifying α.
- For dependent data, the same scaled-model construction could be applied to conditional densities, potentially yielding a posterior over the fraction of temporally or spatially anomalous observations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Hölder-Bayes, a generalised Bayesian framework that performs joint inference over a model parameter θ and an inlier proportion α by applying the Hölder divergence to a scaled model density α p_θ. The stated contributions are: a bounded posterior influence function (Theorem 1), a finite-sample excess-risk bound at rate n^{-1/2} (Theorem 2), population bias control by the tail-overlap quantity ρ_γ(θ*) (Theorem 3), a Bernstein–von Mises theorem (Theorem 4), and a distributional bias bound (Theorem 5). The paper also interprets the posterior temperature as affine scaling of the data space and derives a posterior-based Frequency-of-Detection (FoD) score for outlier detection. Synthetic and real-data experiments are presented to illustrate parameter robustness, contamination recovery, and cleaning performance.
Significance. If the theoretical results are correct, Hölder-Bayes offers a meaningful advance over existing robust GBI methods by jointly quantifying model parameters and contamination without a separate outlier model or external threshold. The connection to the frequentist Hölder-divergence literature is natural, and the FoD mechanism is practically appealing. The paper provides detailed proofs and extensive experiments. However, the proof of the central excess-risk bound (Theorem 2) contains a quantifier error that invalidates the argument as written, and there are inconsistencies in the experimental reporting. The core idea is promising and repairable, but the manuscript currently overstates the strength of its finite-sample guarantees.
major comments (3)
- [Theorem 2 / Appendix B.2] The proof of Theorem 2 fixes a single η=(θ,α), applies McDiarmid to A_n(θ), and obtains a pointwise high-probability bound |H_n(η)−H(η)| ≤ LC√(log(2/δ))/√n. It then invokes Lemma 2 (Pacchiardi–Dutta), which requires a uniform deviation control: sup_η |H_n(η)−H(η)| ≤ ξ_n(δ) on a single event of probability ≥ 1−δ. The pointwise bound does not supply such an event; the intersection over uncountably many η can have probability zero. Consequently the chain of inequalities proving Theorem 2 does not follow. This is a load-bearing gap in the paper's main theoretical contribution. A repair requires uniform concentration, e.g., via covering/bracketing over a compact Θ with Lipschitz θ↦p_θ^γ.
- [Section 6.3, Table 6] Section 6.3 states that the Hölder-Bayes audit for the data-cleaning task uses γ=5, while Table 6 (and its caption) report all Hölder-Bayes audit rows at γ=0.5, and Table 5 ranges over γ∈{0.1,0.5,1.0}. The reported cleaning metrics are therefore ambiguous: it is unclear whether the benchmark corresponds to γ=0.5 or γ=5. Since the cleaning results are used to support the practical value of the method, this inconsistency must be resolved before publication.
- [Theorem 5 / Appendix B.6] The proof of Theorem 5 bounds the Hellinger distance between G* and G0 by C∥η*−η0∥² and then invokes the population bias bound of Theorem 3. However, Theorem 3's bias bound requires δ-strong convexity of H on the level set Uρ, whereas Theorem 5 only assumes eigenvalue bounds on J(η) on the neighbourhood U from Assumption 3, which need not contain η0 or the segment between η* and η0. As stated, the proof does not justify applying the strong-convexity inequality along that segment. The theorem needs either an explicit assumption that η0 (or the relevant segment) lies in the strongly convex region, or a separate argument.
minor comments (4)
- [Appendix B.3] The proof of Theorem 3 introduces ργ(θ*) = {Rγ(θ*)}^γ with an undefined Rγ; use ργ consistently throughout.
- [Introduction] The notation 'clean data ratio in the training dataset α1' should be α; the subscript appears to be a typo.
- [Appendix B.6] The proof of Theorem 5 also uses Rγ(θ*) in the final bound; replace with ργ(θ*) to match the theorem statement.
- [Section 3.4] The derivation of Eq. (16) would benefit from an explicit statement that the second, (θ,α)-independent term of the Hölder divergence also scales by β_T, so that only the first term matters for the posterior.
Circularity Check
No significant circularity: the Hölder posterior construction, α-identification, and FoD summaries are derived from stated assumptions rather than being restatements of their inputs.
full rationale
The paper's central claims are not circular. The identification of the contamination proportion in Eq. (12)-(13) is an explicit identification argument: under the small tail-overlap condition ργ(θ0)≈0, the risk-minimizing α approximates α0. This is a conditional statement whose failure is quantified in Theorem 3 via the bound proportional to ϵ0ργ(θ*); it does not define α0 into the posterior by construction. The FoD outlier scores are procedural summaries of posterior samples: the per-observation ranking depends on the fitted density and is evaluated against external outlier labels, while the fact that the average FoD equals the posterior mean contamination proportion is an accounting identity, not a prediction forced from a fitted parameter renamed as a result. The Hölder divergence, affine scaling property, and scaled-model estimation are imported from published work by Kanamori and Fujisawa and by Matsubara et al.; these are peer-reviewed mathematical results with stated assumptions, used as lemmas whose hypotheses the present paper verifies, rather than unverified self-citations carrying the entire argument. The finite-sample excess-risk proof may have a pointwise-versus-uniform concentration gap, and the heavy-tail-overlap assumption is strong, but these are technical correctness or assumption concerns, not circular reductions. No step was found where a claimed prediction reduces by the paper's own equations to a fitted input or to a self-citation chain.
Assumptions & free parameters
free parameters (3)
- γ (Hölder exponent) =
0.1, 0.5, 1.0 in reported experiments
- Temperature β =
|det Ω̂_n|^{γ/2} from sample covariance
- Saturation parameters κ_exp, κ_rat =
0.5 (ExpSatDPD), 0.25 (RatSatDPD)
assumptions (7)
- domain assumption Heavy-contamination model: P = (1−ε0)Pθ0 + ε0 Q0 with θ0 well-specified for the clean component (Eq. 4).
- domain assumption Small tail overlap: ργ(θ0)=∫pθ0^γ dQ0 is small (Section 2.2, Eq. 5).
- standard math Hölder divergence is a proper scoring rule and satisfies the affine-invariance scaling property (Section 3.1 and 3.4, citing [22,23]).
- domain assumption Assumption 1: model densities uniformly bounded; φ differentiable with −L ≤ φ′ ≤ 0.
- standard math Assumption 2: prior mass condition around population minimiser.
- domain assumption Assumptions 3 and 4: unique interior minimiser, compact local neighbourhood, smoothness/envelope conditions, α0∈(0,1).
- domain assumption Cγ(θ)=∫pθ^{1+γ}(x)dμ(x) is finite and computable (closed form or numerically).
Cite this review
Pith. "Pith review of Generalised Robust Bayes for Joint Inference of Model and Contamination." pith.science (2026). https://pith.science/paper/FQQCCCEK
@misc{pith2026260725665,
author = {Pith},
title = {Pith review of: Generalised Robust Bayes for Joint Inference of Model and Contamination},
year = {2026},
howpublished = {\url{https://pith.science/paper/FQQCCCEK}},
note = {Machine review of arXiv:2607.25665}
}
read the original abstract
Generalised Bayesian inference (GBI) has emerged as a compelling robust alternative to standard Bayesian inference, mitigating sensitivity to data contamination by replacing the log-likelihood with a robust loss or divergence. However, existing robust GBI frameworks typically provide only qualitative robustness: while they can make posterior inference less sensitive to contamination, they lack an intrinsic mechanism to quantify the contamination proportion or identify anomalous observations. This paper introduces H\"older-Bayes, a GBI framework for joint inference of the model parameter and the contamination proportion. We construct a generalised joint posterior over both model and contamination parameter by applying the H\"older divergence to a scaled model density. Theoretically, we establish global bias-robustness via the uniform boundedness of the posterior influence function, derive a finite-sample excess-risk bound, and prove a Bernstein--von Mises approximation together with interpretable contamination-induced bias bounds under a heavy-contamination regime. We further show that, for the H\"older posterior, temperature calibration admits a direct interpretation as affine volume scaling of the data space. The resulting posterior yields a self-contained probabilistic mechanism for outlier detection: posterior uncertainty in both the model parameter and the contamination proportion is propagated to observation-level Frequency-of-Detection scores, without requiring an external anomaly-score threshold. Empirical evaluations demonstrate that H\"older-Bayes provides robust parameter inference, contamination-level recovery, and uncertainty-aware outlier detection.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Basu, I. R. Harris, N. L. Hjort, and M. C. Jones. Robust and efficient estimation by minimising a density power divergence.Biometrika, 85(3):549–559, 1998
1998
-
[2]
J. O. Berger. Robust Bayesian analysis: sensitivity to the prior.Journal of Statistical Planning and Inference, 25 (3):303–328, 1990. ISSN 0378-3758
1990
-
[3]
Mark Berliner
James Berger and L. Mark Berliner. Robust Bayes and empirical Bayes analysis with ϵ-contaminated priors. The Annals of Statistics, 14(2):461 – 486, 1986
1986
-
[4]
P. G. Bissiri, C. C. Holmes, and S. G. Walker. A general framework for updating belief distributions.Journal of the Royal Statistical Society. Series B (Statistical Methodology), 78(5):1103–1130, 2016
2016
-
[5]
Breunig, Hans-Peter Kriegel, Raymond T
Markus M. Breunig, Hans-Peter Kriegel, Raymond T. Ng, and Jörg Sander. LOF: identifying density-based local outliers. InProceedings of the 2000 ACM SIGMOD International Conference on Management of Data, SIGMOD ’00, pages 93–104, New York, NY , USA, 2000. Association for Computing Machinery. ISBN 1581132174. doi: 10.1145/342009.335388. URLhttps://doi.org/1...
-
[6]
Cherief-Abdellatif and P
B.-E. Cherief-Abdellatif and P. Alquier. MMD-Bayes: Robust Bayesian estimation via maximum mean 18 discrepancy. InProceedings of The 2nd Symposium on Advances in Approximate Bayesian Inference, volume 118, pages 1–21, 2020
2020
-
[7]
D. K. Dey and L. R. Birmiwal. Robust Bayesian analysis using divergence measures.Statistics & Probability Letters, 20(4):287–294, 1994. ISSN 0167-7152
1994
-
[8]
Fujisawa and S
H. Fujisawa and S. Eguchi. Robust parameter estimation with a small bias against heavy contamination.Journal of Multivariate Analysis, 99(9):2053–2081, 2008
Show all 42 references
-
[9]
Scalable valuation of human feedback through provably robust model alignment
Masahiro Fujisawa, Masaki Adachi, and Michael A Osborne. Scalable valuation of human feedback through provably robust model alignment. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors,Advances in Neural Information Processing Systems,...
2025
-
[10]
Ghosh, and Aad W
Subhashis Ghosal, Jayanta K. Ghosh, and Aad W. van der Vaart. Convergence rates of posterior distributions. The Annals of Statistics, 28(2):500–531, 2000
2000
-
[11]
Ghosh and A
A. Ghosh and A. Basu. Robust Bayes estimation using the density power divergence.Annals of the Institute of Statistical Mathematics, 68:413–437, 2016
2016
-
[12]
Strictly proper scoring rules, prediction, and estimation.Journal of the American statistical Association, 102(477):359–378, 2007
Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation.Journal of the American statistical Association, 102(477):359–378, 2007
2007
-
[13]
Grünwald
P. Grünwald. Safe learning: bridging the gap between Bayes, MDL and statistical learning theory via empirical convexity. InProceedings of the 24th Annual Conference on Learning Theory, volume 19, pages 397–420, 2011
2011
-
[14]
Grünwald
P. Grünwald. The safe Bayesian. InAlgorithmic Learning Theory, pages 169–183, 2012
2012
-
[15]
Inconsistency of Bayesian Inference for Misspecified Linear Models, and a Proposal for Repairing It.Bayesian Analysis, 12(4):1069 – 1103, 2017
Peter Grünwald and Thijs van Ommen. Inconsistency of Bayesian Inference for Misspecified Linear Models, and a Proposal for Repairing It.Bayesian Analysis, 12(4):1069 – 1103, 2017
2017
-
[16]
F. R. Hampel, E. M. Ronchetti, P. J. Rousseeuw, and W. A. Stahel.Robust Statistics: The Approach Based on Influence Functions. Wiley Series in Probability and Statistics. John Wiley & Sons, 2011
2011
-
[17]
C. C. Holmes and S. G. Walker. Assigning a value to a power likelihood in a general Bayesian model.Biometrika, 104(2):497–503, 2017
2017
-
[18]
Hooker and A
G. Hooker and A. N. Vidyashankar. Bayesian model robustness via disparities.TEST, 23:556–584, 2014
2014
-
[19]
Rousseeuw
Mia Hubert, Michiel Debruyne, and Peter J. Rousseeuw. Minimum covariance determinant and extensions. WIREs Computational Statistics, 10(3), 2018. ISSN 1939-5108
2018
-
[20]
Lecture Notes in Statistics
David Ríos Insua and Fabrizio Ruggeri.Robust Bayesian Analysis. Lecture Notes in Statistics. Springer, 2000
2000
-
[21]
Jewson, J
J. Jewson, J. Q. Smith, and C. Holmes. Principles of bayesian inference using general divergence criteria. Entropy, 20(6), 2018
2018
-
[22]
Kanamori and H
T. Kanamori and H. Fujisawa. Robust estimation under heavy contamination using enlarged models.arXiv preprint arXiv:1311.5301, 2013
2013 arXiv
-
[23]
Kanamori and H
T. Kanamori and H. Fujisawa. Affine invariant divergences associated with proper composite scoring rules and their applications.Bernoulli, 20(4):2278–2304, 2014. 19
2014
-
[24]
Kanamori and H
T. Kanamori and H. Fujisawa. Robust estimation under heavy contamination using unnormalized models. Biometrika, 102(3):559–572, 2015
2015
-
[25]
Duncan, Jeremias Knoblauch, and Francois-Xavier Briol
William Laplante, Matias Altamirano, Andrew B. Duncan, Jeremias Knoblauch, and Francois-Xavier Briol. Robust and conjugate spatio-temporal Gaussian processes. InForty-second International Conference on Machine Learning, 2025. URLhttps://openreview.net/forum?id=YG84SWm7gn
2025
-
[26]
Isolation forest
Fei Tony Liu, Kai Ming Ting, and Zhi-Hua Zhou. Isolation forest. InIEEE International Conference on Data Mining, pages 413–422. IEEE Computer Society, 2008. URL http://dblp.uni-trier.de/db/conf/ icdm/icdm2008.html#LiuTZ08
2008
-
[27]
S. P. Lyddon, C. C. Holmes, and S. G. Walker. General Bayesian updating and the loss-likelihood bootstrap. Biometrika, 106(2):465–478, 2019
2019
-
[28]
Matsubara, J
T. Matsubara, J. Knoblauch, F.-X. Briol, and C. J. Oates. Robust generalised Bayesian inference for intractable likelihoods.Journal of the Royal Statistical Society: Series B (Statistical Methodology), 84(3):997–1022, 2022
2022
-
[29]
Takuo Matsubara, Jeremias Knoblauch, François-Xavier Briol, and Chris. J. Oates and. Generalized bayesian inference for discrete intractable likelihood.Journal of the American Statistical Association, 119(547):2345– 2355, 2024
2024
-
[30]
J. W. Miller and D. B. Dunson. Robust Bayesian inference via coarsening.Journal of the American Statistical Association, 114(527):1113–1125, 2019
2019
-
[31]
Jeffrey W. Miller. Asymptotic normality, concentration, and coverage of generalized posteriors.Journal of Machine Learning Research, 22(168):1–53, 2021
2021
-
[32]
Nakagawa and S
T. Nakagawa and S. Hashimoto. Robust Bayesian inference via γ-divergence.Communications in Statistics - Theory and Methods, 49(2):343–360, 2020
2020
-
[33]
Pacchiardi and R
L. Pacchiardi and R. Dutta. Generalized Bayesian likelihood-free inference using scoring rules estimators.arXiv preprint arXiv:2104.03889, 2021
2021 arXiv
-
[34]
Page and David B
Garritt L. Page and David B. Dunson. Bayesian local contamination models for multivariate outliers.Techno- metrics, 53(2):152–162, 2011. doi: 10.1198/TECH.2011.10041
2011 arXiv
-
[35]
Platt, John C
Bernhard Schölkopf, John C. Platt, John C. Shawe-Taylor, Alex J. Smola, and Robert C. Williamson. Estimating the support of a high-dimensional distribution.Neural Computation, 13(7):1443–1471, 2001. ISSN 0899-7667. doi: 10.1162/089976601750264965. URLhttps://doi.org/10.1162/08...
2001 doi
-
[36]
Shotwell and Elizabeth H
Matthew S. Shotwell and Elizabeth H. Slate. Bayesian Outlier Detection with Dirichlet Process Mixtures. Bayesian Analysis, 6(4):665 – 690, 2011. doi: 10.1214/11-BA625
2011 doi
-
[37]
A. W. van der Vaart.Asymptotic Statistics. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2000
2000
-
[38]
P. M. Williams. Bayesian conditionalisation and the principle of minimum information.British Journal for the Philosophy of Science, 31(2):131–144, 1980
1980
-
[39]
A Comparison of Learning Rate Selection Methods in Generalized Bayesian Inference.Bayesian Analysis, 18(1):105 – 132, 2023
Pei-Shien Wu and Ryan Martin. A Comparison of Learning Rate Selection Methods in Generalized Bayesian Inference.Bayesian Analysis, 18(1):105 – 132, 2023
2023
-
[40]
A. Zellner. Optimal information processing and Bayes’s theorem.The American Statistician, 42(4):278–280, 1988. 20
1988
-
[41]
From ϵ-entropy to KL-entropy: Analysis of minimum information complexity density estimation
Tong Zhang. From ϵ-entropy to KL-entropy: Analysis of minimum information complexity density estimation. The Annals of Statistics, 34(5):2180–2210, 2006
2006
-
[42]
Zografos and S
K. Zografos and S. Nadarajah. Expressions for Rényi and Shannon entropies for multivariate distributions. Statistics & Probability Letters, 71(1):71–84, 2005. 21 Supplementary Material This supplementary material contains proofs of all theoretical results presented in the main...
2005
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.