{"id":"91adc335-6b04-4634-adf3-5e085a951295","arxiv_id":"2607.25665","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Hölder-Bayes performs joint Bayesian inference over model parameters and contamination proportion by applying the Hölder divergence to scaled model densities, yielding robust estimates and observation-level outlier detection.","lead":"This paper introduces Hölder-Bayes, a robust Bayesian method that jointly infers model parameters and the clean-data fraction by evaluating the Hölder divergence against a scaled model density. It combines theory (influence bounds, concentration, Bernstein–von Mises) with experiments showing robust parameter recovery and uncertainty-aware outlier detection.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2's excess-risk bound is unproven as written: the proof applies a pointwise McDiarmid inequality as if it held uniformly over (θ, α), so Lemma 2 cannot be invoked.","rationale":"The reader's verdict (CONDITIONAL) and its rationale already list the pointwise-vs-uniform McDiarmid issue as one of the main weaknesses, but the reader's explicitly identified 'weakest assumption' is the tail-overlap ρ_γ(θ0) identifiability condition. In my stress-test, the most load-bearing concern is instead the proof gap in Theorem 2, because it directly undermines one of the paper's three headline theoretical guarantees and is a matter of logical validity, not just an acknowledged assumption. The tail-overlap assumption is explicitly stated and the paper's bias bounds transparently depend on it; the excess-risk proof gap is an internal inconsistency that the paper does not flag. That said, the concern is likely repairable: under compact Θ and Lipschitz continuity of p_θ^γ, a uniform McDiarmid/covering argument would restore a concentration bound, possibly with a log factor. The empirical FoD mechanism itself is not overturned by this gap, since it relies on the joint posterior construction and the bias-control results rather than on the finite-sample excess-risk theorem. Thus the central framework remains plausible, but the theoretical support is weaker than stated; the conditional verdict is appropriate. I set verdict_should_be to UNCHANGED because the reader's CONDITIONAL verdict already captures this: the paper should be accepted only after the proof is repaired or the theorem is restated with correct assumptions.","tokens_in":39912,"tokens_out":4393,"duration_ms":51554,"concrete_test":"Attempt to re-prove Theorem 2 under the paper's stated assumptions. Specifically: (1) State Lemma 2 precisely and verify whether its hypothesis is a uniform bound; (2) Show that the McDiarmid bound in Appendix B.2 holds for all η simultaneously with probability ≥1-δ, or locate the missing step. If it does not, add a uniform concentration argument (e.g., cover Θ × (0,1] with finitely many ε-balls and use a Lipschitz bound on η↦H_n(η), assuming such a bound exists) and state the corrected rate. If the corrected rate contains an extra log n factor or requires a new entropy condition, Theorem 2's stated O(n^{-1/2}) bound is not established by the current proof.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The finite-sample excess-risk guarantee (Theorem 2) rests on a quantifier error. In Appendix B.2, the authors fix a single η=(θ, α), apply bounded-difference concentration to A_n(θ)=n^{-1}Σ p_θ^γ(Y_i), and obtain |H_n(η)-H(η)| ≤ LC√(log(2/δ))/√n with probability ≥1-δ. This is a pointwise statement. They then feed this into Lemma 2 (the Pacchiardi–Dutta excess-risk lemma) to bound the posterior expectation ∫H(η)π_n(dη). Lemma 2 requires a uniform deviation control: sup_η |H_n(η)-H(η)| ≤ ξ_n(δ) on a single event of probability ≥1-δ, independent of η. A pointwise bound does not provide such an event; the intersection over uncountably many η may have probability zero. Consequently, the chain of inequalities proving Theorem 2 does not follow. This is not a minor constant issue: the proof needs a genuinely different argument, e.g., a uniform concentration inequality over a compact Θ using Lipschitz properties of θ↦p_θ^γ and a covering-number or bracketing argument. The same gap also affects the claim that the posterior concentrates at rate n^{-1/2} in the population Hölder risk. Since Theorem 2 is one of the headline theoretical contributions, this is a load-bearing flaw in the paper's argument as written. It is plausibly repairable under additional compactness and smoothness assumptions, which is why the appropriate verdict remains conditional rather than rejection.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Hölder-Bayes, a generalised Bayesian framework that performs joint inference over a model parameter θ and an inlier proportion α by applying the Hölder divergence to a scaled model density α p_θ. The stated contributions are: a bounded posterior influence function (Theorem 1), a finite-sample excess-risk bound at rate n^{-1/2} (Theorem 2), population bias control by the tail-overlap quantity ρ_γ(θ*) (Theorem 3), a Bernstein–von Mises theorem (Theorem 4), and a distributional bias bound (Theorem 5). The paper also interprets the posterior temperature as affine scaling of the data space and derives a posterior-based Frequency-of-Detection (FoD) score for outlier detection. Synthetic and real-data experiments are presented to illustrate parameter robustness, contamination recovery, and cleaning performance.","tokens_in":40253,"tokens_out":16064,"duration_ms":156078,"significance":"If the theoretical results are correct, Hölder-Bayes offers a meaningful advance over existing robust GBI methods by jointly quantifying model parameters and contamination without a separate outlier model or external threshold. The connection to the frequentist Hölder-divergence literature is natural, and the FoD mechanism is practically appealing. The paper provides detailed proofs and extensive experiments. However, the proof of the central excess-risk bound (Theorem 2) contains a quantifier error that invalidates the argument as written, and there are inconsistencies in the experimental reporting. The core idea is promising and repairable, but the manuscript currently overstates the strength of its finite-sample guarantees.","major_comments":[{"comment":"The proof of Theorem 2 fixes a single η=(θ,α), applies McDiarmid to A_n(θ), and obtains a pointwise high-probability bound |H_n(η)−H(η)| ≤ LC√(log(2/δ))/√n. It then invokes Lemma 2 (Pacchiardi–Dutta), which requires a uniform deviation control: sup_η |H_n(η)−H(η)| ≤ ξ_n(δ) on a single event of probability ≥ 1−δ. The pointwise bound does not supply such an event; the intersection over uncountably many η can have probability zero. Consequently the chain of inequalities proving Theorem 2 does not follow. This is a load-bearing gap in the paper's main theoretical contribution. A repair requires uniform concentration, e.g., via covering/bracketing over a compact Θ with Lipschitz θ↦p_θ^γ.","section":"Theorem 2 / Appendix B.2"},{"comment":"Section 6.3 states that the Hölder-Bayes audit for the data-cleaning task uses γ=5, while Table 6 (and its caption) report all Hölder-Bayes audit rows at γ=0.5, and Table 5 ranges over γ∈{0.1,0.5,1.0}. The reported cleaning metrics are therefore ambiguous: it is unclear whether the benchmark corresponds to γ=0.5 or γ=5. Since the cleaning results are used to support the practical value of the method, this inconsistency must be resolved before publication.","section":"Section 6.3, Table 6"},{"comment":"The proof of Theorem 5 bounds the Hellinger distance between G* and G0 by C∥η*−η0∥² and then invokes the population bias bound of Theorem 3. However, Theorem 3's bias bound requires δ-strong convexity of H on the level set Uρ, whereas Theorem 5 only assumes eigenvalue bounds on J(η) on the neighbourhood U from Assumption 3, which need not contain η0 or the segment between η* and η0. As stated, the proof does not justify applying the strong-convexity inequality along that segment. The theorem needs either an explicit assumption that η0 (or the relevant segment) lies in the strongly convex region, or a separate argument.","section":"Theorem 5 / Appendix B.6"}],"minor_comments":[{"comment":"The proof of Theorem 3 introduces ργ(θ*) = {Rγ(θ*)}^γ with an undefined Rγ; use ργ consistently throughout.","section":"Appendix B.3"},{"comment":"The notation 'clean data ratio in the training dataset α1' should be α; the subscript appears to be a typo.","section":"Introduction"},{"comment":"The proof of Theorem 5 also uses Rγ(θ*) in the final bound; replace with ργ(θ*) to match the theorem statement.","section":"Appendix B.6"},{"comment":"The derivation of Eq. (16) would benefit from an explicit statement that the second, (θ,α)-independent term of the Hölder divergence also scales by β_T, so that only the first term matters for the posterior.","section":"Section 3.4"}],"recommendation":"major_revision","confidential_remarks":"The paper leans on the authors' own prior work on Hölder divergences and unnormalized models, which is appropriate given the topic. The main novelty lies in the Bayesian joint formulation and the FoD mechanism. No scope concerns. The Theorem 2 gap is the key technical issue; the experimental inconsistency in γ is less severe but should be fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a serious attempt to make robust generalised Bayesian inference actually quantify contamination: a joint posterior over (θ, α) with a Hölder risk on the scaled model, plus observation-level Frequency-of-Detection scores. That is genuinely new relative to the existing frequentist contamination-ratio estimators. Second, the main finite-sample excess-risk theorem (Theorem 2) has a proof gap that is not a matter of tightening constants. The authors fix a single η, apply McDiarmid, get a pointwise deviation bound, and then feed it into the Pacchiardi–Dutta lemma, which needs a uniform deviation control over the whole parameter space. A pointwise bound does not give you a single probability-1−δ event on which the integration over the posterior goes through. This is a genuine quantifier error. It is plausibly repairable with compactness and Lipschitz arguments, but as written the theorem does not follow.\n\nThe paper has real strengths. The construction of the Hölder posterior and the FoD mechanism is clear and well motivated. The appendix is serious: PIF, BvM, bias control, and the affine-scaling interpretation of the temperature are all worked out in detail. The empirical sections show that FoD works on synthetic and real data, and the cleaning experiments are sensible. The authors are also honest that contamination recovery depends on small tail overlap between Q0 and the model; the identifiability argument is an approximation, not a restatement of the conclusion.\n\nSoft spots beyond Theorem 2: Theorem 1 assumes the posterior density is bounded without proving it, which is a gap but probably minor. There is an internal inconsistency where Section 6.3 says γ=5 while the tables and the rest of the text use γ=0.5; almost certainly a typo, but it should be fixed. No code is provided, which makes the empirical claims harder to check. The bias bound in Theorem 3 is proportional to ϵ0 ργ(θ*), so if outliers overlap the model tails, α is weakly identified; the paper acknowledges this but does not stress-test the regime where the overlap is moderate.\n\nWho is this for? People working on robust Bayes, outlier detection, and generalised posteriors. It deserves a serious referee, but the referee should demand a fix for Theorem 2, a clarification of the γ inconsistency, and ideally a code release. I would not cite the excess-risk bound until the proof is repaired, but I would cite the FoD construction and the temperature interpretation as useful ideas.\n\nRecommendation: send to peer review, but expect at least one round of major revision.","headline":"Worth engaging with, but the headline finite-sample excess-risk theorem has a real quantifier gap and the paper needs a revision before the theory is trustworthy.","tokens_in":40780,"tokens_out":2078,"would_cite":false,"duration_ms":26020,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62F35"],"pacs":[],"model":"deepseek-v4-flash","headline":"Hölder-Bayes makes the contamination proportion an inferable parameter in robust Bayesian updating.","keywords":["Hölder divergence","generalised Bayesian inference","heavy contamination","robustness","outlier detection","contamination proportion","posterior influence function","Bernstein–von Mises"],"falsifier":"Simulate data from a contamination model where Q0 has substantial density at the centre of Pθ0 (e.g., a Gaussian outlier centred at the true mean with the same scale) and check whether the Hölder posterior for α recovers the injected ϵ0; if the posterior is sharply biased away from ϵ0, the identifiability mechanism breaks.","tokens_in":39757,"feed_emoji":"🎯","tokens_out":2423,"duration_ms":26918,"temperature":0.7,"pith_summary":"The paper tries to show that generalised Bayesian inference can go beyond protecting parameter estimates from outliers and actually quantify the contamination itself. It builds a joint posterior over the model parameter θ and the inlier proportion α by scoring the data against a scaled version of the model, αPθ, using the Hölder divergence. The central claim is that this joint posterior simultaneously gives robust parameter inference, a calibrated posterior over the contamination level, and observation-level outlier probabilities (Frequency-of-Detection scores) without any external anomaly threshold. The paper supports this with theoretical guarantees: a uniformly bounded posterior influence function, a finite-sample excess-risk bound, a Bernstein–von Mises theorem, and bias bounds controlled by a tail-overlap quantity. If correct, this would give statisticians a single, self-contained tool for robust inference, contamination auditing, and outlier detection.","feed_headline":"One posterior learns the model and the contamination level","feed_subtitle":"A scaled-model Hölder divergence lets Bayesian updating estimate the outlier fraction and flag anomalies, no external threshold.","key_machinery":"The central object is the Hölder divergence applied to a scaled model: the empirical risk H_n^(γ)(θ,α)=ϕ(α^{-1}S_{n,γ}(θ)/C_γ(θ)) α^{1+γ}C_γ(θ), where S_{n,γ}(θ) is the empirical average of pθ^γ and C_γ(θ) is its model expectation. The scaling parameter α acts as the inlier proportion and is inferred jointly with θ. The theoretical analysis is carried by the tail-overlap quantity ργ(θ)=∫pθ^γ dQ0, which controls the bias in both the population minimiser (Theorem 3) and the asymptotic distribution (Theorem 5).","core_discovery":"The paper introduces the Hölder posterior as a generalised posterior over (θ, α) formed by exponentiating the empirical Hölder divergence between the data and the scaled model αPθ. The key structural insight is that, when the outlier distribution Q0 has small tail overlap with the target model density pθ0—quantified by ργ(θ0)=∫pθ0^γ dQ0—the ratio Sn,γ(θ)/Cγ(θ) is dominated by the clean component, so the risk-minimising α approximates the true inlier proportion α0. The paper proves global bias-robustness via a bounded posterior influence function, establishes a finite-sample excess-risk bound at rate n^{-1/2}, and derives a Bernstein–von Mises theorem showing asymptotic normality with covaria","pith_inferences":["The tail-overlap assumption could be tested in practice: on real datasets, one might estimate ργ(θ*) and compare the posterior for ϵ to an external robust estimate; large discrepancies would signal regime in which the framework is unreliable.","The temperature-as-scaling interpretation suggests a possible extension: instead of the sample covariance, one could choose the affine transformation that minimises posterior predictive loss, making the temperature fully data-adaptive.","The joint inference idea may extend beyond Hölder divergence to other scoring rules that are proper and sensitive to scaling; the bounded-derivative condition in Assumption 1 is what prevents scale-invariant losses from identifying α.","For dependent data, the same scaled-model construction could be applied to conditional densities, potentially yielding a posterior over the fraction of temporally or spatially anomalous observations."],"forward_implications":["Robust Bayesian inference can now report a coherent posterior over the contamination proportion rather than a single robustness guarantee, enabling principled uncertainty quantification of data quality.","Outlier detection becomes part of the same probabilistic model that fits the clean signal: observation-level Frequency-of-Detection scores are produced by propagating posterior draws of (θ, α), with no external threshold required.","The temperature parameter in generalised Bayes gains a concrete meaning as an affine scaling of the data space, suggesting that data standardisation can replace expensive temperature-tuning procedures.","The bias bounds imply that the feasibility of heavy-contamination inference is governed by tail overlap between outliers and the model, not merely by the contamination fraction itself."],"fun_headline_variants":["Hölder-Bayes: one posterior for model and contamination","No threshold: Hölder-Bayes jointly infers model and contamination","Hölder posterior: joint model and contamination inference","Hölder-Bayes: outliers and model from one posterior","Joint inference of model and contamination via Hölder divergence"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The outliers must be concentrated in the low-density tails of the target model, so that the overlap term ργ(θ0) is negligible; otherwise the contamination proportion is not identifiable and the bias bound in Theorem 3 can be large.","fun_headline_variants_meta":{"raw":{"variants":["Hölder-Bayes: one posterior for model and contamination","No threshold: Hölder-Bayes jointly infers model and contamination","Hölder posterior: joint model and contamination inference","Hölder-Bayes: outliers and model from one posterior","Joint inference of model and contamination via Hölder divergence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002082,"raw_usage":{"total_tokens":7961,"prompt_tokens":801,"completion_tokens":7160,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":7084}},"tokens_in":545,"tokens_out":7160,"duration_ms":48010,"temperature":1.0,"reasoning_tokens":7084,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:48:27.201206+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data from a contamination model where Q0 has substantial density at the centre of Pθ0 (e.g., a Gaussian outlier centred at the true mean with the same scale) and check whether the Hölder posterior for α recovers the injected ϵ0; if the posterior is sharply biased away from ϵ0, the identifiability mechanism breaks.","supporting_citations":[],"review_version":1}