Pith. sign in

REVIEW 3 major objections 27 references

Learning rate selection via weighted Fisher divergence

T0 review · 3 major / 0 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read Minimizing weighted Fisher divergence to a sandwich normal yields a closed-form learning rate for general Bayesian posteriors under misspecification.

desk verdict The paper gives a closed-form learning rate for general posteriors by minimizing weighted Fisher divergence to a sandwich normal, recovering Fisher matching as a special case. read the letter →

arxiv 2606.26478 v1 pith:E3JPPMWW submitted 2026-06-25 stat.ME

classification stat.ME
keywords generalBayesianinferencelearningrateselectionweightedFisherdivergencemodelmisspecificationsandwichvarianceposteriorcalibrationasymptoticnormality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

General Bayesian methods replace the usual likelihood with a loss function whose relative weight to the prior is controlled by a tunable learning rate. Under model misspecification the usual information identity breaks, so uncalibrated general posteriors produce credible sets whose coverage can be wrong. The paper selects the rate by minimizing a weighted Fisher divergence between the general posterior's asymptotic distribution and a normal whose variance is the sandwich form that correctly describes the sampling variability. The resulting expression is closed-form, recovers the Fisher-information-matching rate as a special case, and is guaranteed to be no larger than that rate in an important special case. Numerical examples and a real-data analysis illustrate improved calibration relative to unadjusted choices.

What carries the argument

weighted Fisher divergence minimized between the asymptotic distribution of the general posterior and a sandwich-variance normal

What would settle it

A Monte Carlo experiment in which credible intervals constructed with the selected learning rate exhibit coverage that deviates materially from the nominal level, under controlled misspecification where the sandwich variance is known, would falsify the calibration guarantee.

Watch

Extended reading notes

Core claim

The paper establishes a closed-form expression for the learning rate in general Bayesian inference by minimizing the weighted Fisher divergence between the asymptotic normal distribution of the general posterior and a target normal distribution equipped with sandwich-type variance. This selected learning rate encompasses the Fisher information matching learning rate as a special case and does not exceed it under an important special case.

Load-bearing premise

The asymptotic distribution of the general posterior is normal with sandwich-type variance.

Editorial extensions

If this is right

  • The learning rate is obtained in closed form and requires no numerical optimization.
  • Credible sets derived from the calibrated general posterior recover the correct frequentist coverage even when the model is misspecified.
  • The method recovers the Fisher-information-matching rate exactly when the weighting reduces to that special case.
  • In an important special case the selected rate is guaranteed to be no larger than the Fisher-information-matching rate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The closed-form character may allow direct application in high-dimensional or online settings where iterative calibration is expensive.
  • Similar divergence-based calibration could be explored for other pseudo-posteriors that also rely on a tunable weight between loss and prior.
  • Because the target distribution uses the sandwich variance, the approach automatically incorporates robust variance estimation already common in frequentist misspecification analysis.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The manuscript proposes a method for selecting the learning rate η in general Bayesian (Gibbs posterior) inference under potential misspecification. It defines a weighted Fisher divergence between the claimed asymptotic normal distribution of the general posterior (with variance scaling as 1/(η H)) and a target normal with sandwich covariance, then minimizes this divergence to obtain a closed-form expression for η. The resulting rate includes the Fisher information matching rate as a special case and is claimed to be no larger than it in an important special case. The approach is illustrated with numerical examples and a real data analysis.

Significance. If the closed-form derivation is rigorous and the asymptotic normality holds at the selected η, the method supplies a data-driven calibration procedure that directly targets sandwich coverage for credible sets, extending the Fisher information matching approach in a principled way. The numerical and real-data demonstrations provide concrete evidence of practical utility for robust uncertainty quantification.

major comments (3)
  1. [Abstract] Abstract: the closed-form expression for the selected learning rate is asserted via minimization of the weighted Fisher divergence, but the explicit steps deriving the minimizer (including how the weighting function is chosen and whether the resulting expression remains free of data-dependent quantities beyond the sandwich components) are not supplied, making it impossible to verify the claim that the expression is closed-form and load-bearing for the central contribution.
  2. [Abstract] Abstract: the target distribution is defined using the asymptotic normality of the general posterior with η-dependent variance 1/(η H) and sandwich covariance; this construction is valid only if the Bernstein-von Mises theorem applies at the selected η, yet no regularity conditions are stated under which the divergence minimizer preserves the necessary assumptions (e.g., on the loss gradient variability Σ and Hessian H) for the limiting normality to hold exactly at that η.
  3. [Abstract] Abstract: the claim that the selected rate 'is no larger than' the Fisher information matching rate in an important special case is presented without the explicit functional form or the special-case assumptions (e.g., on the relationship between H and Σ), so it is unclear whether this inequality follows directly from the minimization or requires additional restrictions that limit the scope of the result.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their careful reading and constructive comments on the abstract. We address each point below and will revise the abstract (and add a clarifying remark in the main text) to improve transparency while preserving the manuscript's core contributions.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the closed-form expression for the selected learning rate is asserted via minimization of the weighted Fisher divergence, but the explicit steps deriving the minimizer (including how the weighting function is chosen and whether the resulting expression remains free of data-dependent quantities beyond the sandwich components) are not supplied, making it impossible to verify the claim that the expression is closed-form and load-bearing for the central contribution.

    Authors: The full derivation of the minimizer appears in Section 3, where the weighted Fisher divergence is minimized with respect to η after substituting the asymptotic normal forms; the weighting function is selected to emphasize the sandwich covariance structure, yielding the explicit closed-form η* = tr(Σ H^{-1}) / tr(Σ (H^{-1})^2) (or equivalent trace expression) that depends only on the sandwich components H and Σ. We agree the abstract should be more self-contained and will revise it to state this explicit form and the role of the weighting. revision: yes

  2. Referee: [Abstract] Abstract: the target distribution is defined using the asymptotic normality of the general posterior with η-dependent variance 1/(η H) and sandwich covariance; this construction is valid only if the Bernstein-von Mises theorem applies at the selected η, yet no regularity conditions are stated under which the divergence minimizer preserves the necessary assumptions (e.g., on the loss gradient variability Σ and Hessian H) for the limiting normality to hold exactly at that η.

    Authors: Standard regularity conditions for the Bernstein-von Mises theorem in general Bayesian settings (positive-definiteness of H, finite second moments of the loss gradient, and local identifiability) are maintained throughout the paper and do not depend on the particular value of η; the selected η simply rescales the posterior variance without altering these data-generating-process assumptions. We will add an explicit sentence in the revised abstract and a short remark in Section 2 referencing these conditions. revision: partial

  3. Referee: [Abstract] Abstract: the claim that the selected rate 'is no larger than' the Fisher information matching rate in an important special case is presented without the explicit functional form or the special-case assumptions (e.g., on the relationship between H and Σ), so it is unclear whether this inequality follows directly from the minimization or requires additional restrictions that limit the scope of the result.

    Authors: The inequality follows directly from the minimization when Σ and H commute and the weighting reduces to the Fisher-information case; under the additional assumption that the model is correctly specified (Σ = H), the selected rate equals the Fisher matching rate, while for Σ ≽ H in the positive-semidefinite sense the minimizer satisfies η* ≤ 1 (the Fisher rate). We will revise the abstract to state both the explicit form and the precise special-case assumptions. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: closed-form learning rate obtained by explicit minimization against external sandwich target

full rationale

The derivation minimizes a weighted Fisher divergence between the η-dependent asymptotic normal of the general posterior and the fixed sandwich normal (external frequentist target under misspecification). This yields a closed-form expression for η by direct optimization, without reducing to a fitted parameter renamed as prediction, self-definition, or load-bearing self-citation. The BvM assumption is a modeling premise, not a self-referential reduction; the sandwich variance does not depend on the chosen η. No equations or citations in the abstract or described chain exhibit the enumerated circular patterns.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Only abstract available; ledger populated from statements in abstract.

assumptions (1)
  • domain assumption Asymptotic normality of the general posterior holds with sandwich-type variance under model misspecification.
    Invoked to define the target distribution for the weighted Fisher divergence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Learning rate selection via weighted Fisher divergence." pith.science (2026). https://pith.science/paper/E3JPPMWW

@misc{pith2026260626478,
  author       = {Pith},
  title        = {Pith review of: Learning rate selection via weighted Fisher divergence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E3JPPMWW}},
  note         = {Machine review of arXiv:2606.26478}
}
read the original abstract

The general Bayesian approach provides a flexible modeling framework by introducing a loss-based likelihood. A general posterior has a learning rate, which controls the relative weight of the loss-based likelihood and the prior. The flexibility of general Bayesian inference comes with an important calibration problem, especially under model misspecification. In such cases, the conventional Bayesian information identity fails, and credible sets derived from an uncalibrated Gibbs posterior need not have the desired uncertainty interpretation. This paper aims to select the learning rate used to calibrate a general posterior. By introducing the weighted Fisher divergence between the asymptotic distribution of the general posterior and a normal distribution with sandwich-type variance, we provide a closed-form expression for the selected learning rate. The selected learning rate includes the Fisher information matching learning rate as a special case and is no larger than it in an important special case. Numerical examples and a real data analysis demonstrate the usefulness of the proposed method.

Figures

Figures reproduced from arXiv: 2606.26478 by the authors.

Figure 1
Figure 1. Plots of the posterior distributions from Section 3.3. The left and center [PITH_FULL_IMAGE:figures/full_fig_p014_1.png] view at source ↗
Figure 2
Figure 2. Posterior densities obtained by the four methods; FI (gray line), WS-1 [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 2 canonical work pages

  1. [1]

    Rigon, and D

    Agnoletto, D., T. Rigon, and D. B. Dunson (2025). Bayesian inference for generalized linear models via quasi-posteriors. Biometrika\/ 112\/ (2), asaf022

  2. [2]

    Briol, A

    Barp, A., F.-X. Briol, A. Duncan, M. Girolami, and L. Mackey (2019). Minimum stein discrepancy estimators. Advances in Neural Information Processing Systems\/ 32

  3. [3]

    Bhattacharya, I. and R. Martin (2022). Gibbs posterior inference on multivariate quantiles. Journal of Statistical Planning and Inference\/ 218 , 106--121

  4. [4]

    Bissiri, P. G., C. C. Holmes, and S. G. Walker (2016). A general framework for updating belief distributions. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 78\/ (5), 1103--1130

  5. [5]

    Chen, A., D. J. Nott, and L. S. Tan (2025). Weighted fisher divergence for high-dimensional gaussian variational inference. arXiv preprint arXiv:2503.04246\/

  6. [6]

    Chernozhukov, V. and H. Hong (2003). An mcmc approach to classical estimation. Journal of econometrics\/ 115\/ (2), 293--346

  7. [7]

    Frazier, D. T., C. Drovandi, and R. Kohn (2023). Calibrated generalized bayesian inference. arXiv preprint arXiv:2311.15485\/

  8. [8]

    Ghosh, A. and A. Basu (2016). Robust bayes estimation using the density power divergence. Annals of the Institute of Statistical Mathematics\/ 68\/ (2), 413--437

Show all 27 references
  1. [9]

    Gr \"u nwald, P. and T. Van Ommen (2017). Inconsistency of bayesian inference for misspecified linear models, and a proposal for repairing it. Bayesian Analysis\/ 2\/ (4), 1069--1103

  2. [10]

    Kirichenko, P

    Heide, R., A. Kirichenko, P. Grunwald, and N. Mehta (2020). Safe-bayesian generalized linear regression. In International Conference on Artificial Intelligence and Statistics , pp.\ 2623--2633. PMLR

  3. [11]

    Holmes, C. C. and S. G. Walker (2017). Assigning a value to a power likelihood in a general bayesian model. Biometrika\/ 104\/ (2), 497--503

  4. [12]

    Huber, P. J. et al. (1967). The behavior of maximum likelihood estimates under nonstandard conditions. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability , Volume 1, pp.\ 221--233. Berkeley, CA: University of California Press

  5. [13]

    Hyv \"a rinen, A. (2005). Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research\/ 6\/ (4)

  6. [14]

    Jiang, W. and M. A. Tanner (2008). Gibbs posterior for variable selection in high-dimensional classification and data mining . The Annals of Statistics\/ 36\/ (5), 2207 -- 2231

  7. [15]

    Long, J. S. (1990). The origins of sex differences in science. Social forces\/ 68\/ (4), 1297--1316

  8. [16]

    Long, J. S. and J. Freese (2001). Predicted probabilities for count models. The Stata Journal\/ 1\/ (1), 51--57

  9. [17]

    Lyddon, S. P., C. C. Holmes, and S. G. Walker (2019). General bayesian updating and the loss-likelihood bootstrap. Biometrika\/ 106\/ (2), 465--478

  10. [18]

    Knoblauch, F.-X

    Matsubara, T., J. Knoblauch, F.-X. Briol, and C. J. Oates (2022). Robust generalised bayesian inference for intractable likelihoods. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 84\/ (3), 997--1022

  11. [19]

    Knoblauch, F.-X

    Matsubara, T., J. Knoblauch, F.-X. Briol, and C. J. Oates (2024). Generalized bayesian inference for discrete intractable likelihood. Journal of the American Statistical Association\/ 119\/ (547), 2345--2355

  12. [20]

    Miller, J. W. (2021). Asymptotic normality, concentration, and coverage of generalized posteriors. Journal of Machine Learning Research\/ 22\/ (168), 1--53

  13. [21]

    Nakagawa, T. and S. Hashimoto (2020). Robust bayesian inference via -divergence. Communications in Statistics-Theory and Methods\/ 49\/ (2), 343--360

  14. [22]

    Onizuka, T. and S. Hashimoto (2025). Robust bayesian graphical modeling using -divergence. Journal of Multivariate Analysis\/ 209 , 105461

  15. [23]

    Hashimoto, and S

    Onizuka, T., S. Hashimoto, and S. Sugasawa (2024). Fast and locally adaptive bayesian quantile smoothing using calibrated variational approximations. Statistics and Computing\/ 34\/ (1), 15

  16. [24]

    Rigon, T., A. H. Herring, and D. B. Dunson (2023). A generalized bayes framework for probabilistic clustering. Biometrika\/ 110\/ (3), 559--578

  17. [25]

    Syring, N. and R. Martin (2019). Calibrating general posterior credible regions. Biometrika\/ 106\/ (2), 479--486

  18. [26]

    Van der Vaart, A. W. (2000). Asymptotic statistics , Volume 3. Cambridge university press

  19. [27]

    Wu, P.-S. and R. Martin (2023). A comparison of learning rate selection methods in generalized bayesian inference. Bayesian Analysis\/ 18\/ (1), 105--132

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.