REVIEW 3 major objections 27 references
Learning rate selection via weighted Fisher divergence
T0 review · 3 major / 0 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Minimizing weighted Fisher divergence to a sandwich normal yields a closed-form learning rate for general Bayesian posteriors under misspecification.
desk verdict The paper gives a closed-form learning rate for general posteriors by minimizing weighted Fisher divergence to a sandwich normal, recovering Fisher matching as a special case. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
weighted Fisher divergence minimized between the asymptotic distribution of the general posterior and a sandwich-variance normal
What would settle it
A Monte Carlo experiment in which credible intervals constructed with the selected learning rate exhibit coverage that deviates materially from the nominal level, under controlled misspecification where the sandwich variance is known, would falsify the calibration guarantee.
Extended reading notes
Core claim
The paper establishes a closed-form expression for the learning rate in general Bayesian inference by minimizing the weighted Fisher divergence between the asymptotic normal distribution of the general posterior and a target normal distribution equipped with sandwich-type variance. This selected learning rate encompasses the Fisher information matching learning rate as a special case and does not exceed it under an important special case.
Load-bearing premise
The asymptotic distribution of the general posterior is normal with sandwich-type variance.
Editorial extensions
If this is right
- The learning rate is obtained in closed form and requires no numerical optimization.
- Credible sets derived from the calibrated general posterior recover the correct frequentist coverage even when the model is misspecified.
- The method recovers the Fisher-information-matching rate exactly when the weighting reduces to that special case.
- In an important special case the selected rate is guaranteed to be no larger than the Fisher-information-matching rate.
Reading between the lines
- The closed-form character may allow direct application in high-dimensional or online settings where iterative calibration is expensive.
- Similar divergence-based calibration could be explored for other pseudo-posteriors that also rely on a tunable weight between loss and prior.
- Because the target distribution uses the sandwich variance, the approach automatically incorporates robust variance estimation already common in frequentist misspecification analysis.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a method for selecting the learning rate η in general Bayesian (Gibbs posterior) inference under potential misspecification. It defines a weighted Fisher divergence between the claimed asymptotic normal distribution of the general posterior (with variance scaling as 1/(η H)) and a target normal with sandwich covariance, then minimizes this divergence to obtain a closed-form expression for η. The resulting rate includes the Fisher information matching rate as a special case and is claimed to be no larger than it in an important special case. The approach is illustrated with numerical examples and a real data analysis.
Significance. If the closed-form derivation is rigorous and the asymptotic normality holds at the selected η, the method supplies a data-driven calibration procedure that directly targets sandwich coverage for credible sets, extending the Fisher information matching approach in a principled way. The numerical and real-data demonstrations provide concrete evidence of practical utility for robust uncertainty quantification.
major comments (3)
- [Abstract] Abstract: the closed-form expression for the selected learning rate is asserted via minimization of the weighted Fisher divergence, but the explicit steps deriving the minimizer (including how the weighting function is chosen and whether the resulting expression remains free of data-dependent quantities beyond the sandwich components) are not supplied, making it impossible to verify the claim that the expression is closed-form and load-bearing for the central contribution.
- [Abstract] Abstract: the target distribution is defined using the asymptotic normality of the general posterior with η-dependent variance 1/(η H) and sandwich covariance; this construction is valid only if the Bernstein-von Mises theorem applies at the selected η, yet no regularity conditions are stated under which the divergence minimizer preserves the necessary assumptions (e.g., on the loss gradient variability Σ and Hessian H) for the limiting normality to hold exactly at that η.
- [Abstract] Abstract: the claim that the selected rate 'is no larger than' the Fisher information matching rate in an important special case is presented without the explicit functional form or the special-case assumptions (e.g., on the relationship between H and Σ), so it is unclear whether this inequality follows directly from the minimization or requires additional restrictions that limit the scope of the result.
Simulated Author's Rebuttal
We thank the referee for their careful reading and constructive comments on the abstract. We address each point below and will revise the abstract (and add a clarifying remark in the main text) to improve transparency while preserving the manuscript's core contributions.
read point-by-point responses
-
Referee: [Abstract] Abstract: the closed-form expression for the selected learning rate is asserted via minimization of the weighted Fisher divergence, but the explicit steps deriving the minimizer (including how the weighting function is chosen and whether the resulting expression remains free of data-dependent quantities beyond the sandwich components) are not supplied, making it impossible to verify the claim that the expression is closed-form and load-bearing for the central contribution.
Authors: The full derivation of the minimizer appears in Section 3, where the weighted Fisher divergence is minimized with respect to η after substituting the asymptotic normal forms; the weighting function is selected to emphasize the sandwich covariance structure, yielding the explicit closed-form η* = tr(Σ H^{-1}) / tr(Σ (H^{-1})^2) (or equivalent trace expression) that depends only on the sandwich components H and Σ. We agree the abstract should be more self-contained and will revise it to state this explicit form and the role of the weighting. revision: yes
-
Referee: [Abstract] Abstract: the target distribution is defined using the asymptotic normality of the general posterior with η-dependent variance 1/(η H) and sandwich covariance; this construction is valid only if the Bernstein-von Mises theorem applies at the selected η, yet no regularity conditions are stated under which the divergence minimizer preserves the necessary assumptions (e.g., on the loss gradient variability Σ and Hessian H) for the limiting normality to hold exactly at that η.
Authors: Standard regularity conditions for the Bernstein-von Mises theorem in general Bayesian settings (positive-definiteness of H, finite second moments of the loss gradient, and local identifiability) are maintained throughout the paper and do not depend on the particular value of η; the selected η simply rescales the posterior variance without altering these data-generating-process assumptions. We will add an explicit sentence in the revised abstract and a short remark in Section 2 referencing these conditions. revision: partial
-
Referee: [Abstract] Abstract: the claim that the selected rate 'is no larger than' the Fisher information matching rate in an important special case is presented without the explicit functional form or the special-case assumptions (e.g., on the relationship between H and Σ), so it is unclear whether this inequality follows directly from the minimization or requires additional restrictions that limit the scope of the result.
Authors: The inequality follows directly from the minimization when Σ and H commute and the weighting reduces to the Fisher-information case; under the additional assumption that the model is correctly specified (Σ = H), the selected rate equals the Fisher matching rate, while for Σ ≽ H in the positive-semidefinite sense the minimizer satisfies η* ≤ 1 (the Fisher rate). We will revise the abstract to state both the explicit form and the precise special-case assumptions. revision: yes
Circularity Check
No circularity: closed-form learning rate obtained by explicit minimization against external sandwich target
full rationale
The derivation minimizes a weighted Fisher divergence between the η-dependent asymptotic normal of the general posterior and the fixed sandwich normal (external frequentist target under misspecification). This yields a closed-form expression for η by direct optimization, without reducing to a fitted parameter renamed as prediction, self-definition, or load-bearing self-citation. The BvM assumption is a modeling premise, not a self-referential reduction; the sandwich variance does not depend on the chosen η. No equations or citations in the abstract or described chain exhibit the enumerated circular patterns.
Assumptions & free parameters
assumptions (1)
- domain assumption Asymptotic normality of the general posterior holds with sandwich-type variance under model misspecification.
Cite this review
Pith. "Pith review of Learning rate selection via weighted Fisher divergence." pith.science (2026). https://pith.science/paper/E3JPPMWW
@misc{pith2026260626478,
author = {Pith},
title = {Pith review of: Learning rate selection via weighted Fisher divergence},
year = {2026},
howpublished = {\url{https://pith.science/paper/E3JPPMWW}},
note = {Machine review of arXiv:2606.26478}
}
read the original abstract
The general Bayesian approach provides a flexible modeling framework by introducing a loss-based likelihood. A general posterior has a learning rate, which controls the relative weight of the loss-based likelihood and the prior. The flexibility of general Bayesian inference comes with an important calibration problem, especially under model misspecification. In such cases, the conventional Bayesian information identity fails, and credible sets derived from an uncalibrated Gibbs posterior need not have the desired uncertainty interpretation. This paper aims to select the learning rate used to calibrate a general posterior. By introducing the weighted Fisher divergence between the asymptotic distribution of the general posterior and a normal distribution with sandwich-type variance, we provide a closed-form expression for the selected learning rate. The selected learning rate includes the Fisher information matching learning rate as a special case and is no larger than it in an important special case. Numerical examples and a real data analysis demonstrate the usefulness of the proposed method.
Figures
Reference graph
Works this paper leans on
-
[1]
Rigon, and D
Agnoletto, D., T. Rigon, and D. B. Dunson (2025). Bayesian inference for generalized linear models via quasi-posteriors. Biometrika\/ 112\/ (2), asaf022
2025
-
[2]
Briol, A
Barp, A., F.-X. Briol, A. Duncan, M. Girolami, and L. Mackey (2019). Minimum stein discrepancy estimators. Advances in Neural Information Processing Systems\/ 32
2019
-
[3]
Bhattacharya, I. and R. Martin (2022). Gibbs posterior inference on multivariate quantiles. Journal of Statistical Planning and Inference\/ 218 , 106--121
2022
-
[4]
Bissiri, P. G., C. C. Holmes, and S. G. Walker (2016). A general framework for updating belief distributions. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 78\/ (5), 1103--1130
2016
- [5]
-
[6]
Chernozhukov, V. and H. Hong (2003). An mcmc approach to classical estimation. Journal of econometrics\/ 115\/ (2), 293--346
2003
- [7]
-
[8]
Ghosh, A. and A. Basu (2016). Robust bayes estimation using the density power divergence. Annals of the Institute of Statistical Mathematics\/ 68\/ (2), 413--437
2016
Show all 27 references
-
[9]
Gr \"u nwald, P. and T. Van Ommen (2017). Inconsistency of bayesian inference for misspecified linear models, and a proposal for repairing it. Bayesian Analysis\/ 2\/ (4), 1069--1103
2017
-
[10]
Kirichenko, P
Heide, R., A. Kirichenko, P. Grunwald, and N. Mehta (2020). Safe-bayesian generalized linear regression. In International Conference on Artificial Intelligence and Statistics , pp.\ 2623--2633. PMLR
2020
-
[11]
Holmes, C. C. and S. G. Walker (2017). Assigning a value to a power likelihood in a general bayesian model. Biometrika\/ 104\/ (2), 497--503
2017
-
[12]
Huber, P. J. et al. (1967). The behavior of maximum likelihood estimates under nonstandard conditions. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability , Volume 1, pp.\ 221--233. Berkeley, CA: University of California Press
1967
-
[13]
Hyv \"a rinen, A. (2005). Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research\/ 6\/ (4)
2005
-
[14]
Jiang, W. and M. A. Tanner (2008). Gibbs posterior for variable selection in high-dimensional classification and data mining . The Annals of Statistics\/ 36\/ (5), 2207 -- 2231
2008
-
[15]
Long, J. S. (1990). The origins of sex differences in science. Social forces\/ 68\/ (4), 1297--1316
1990
-
[16]
Long, J. S. and J. Freese (2001). Predicted probabilities for count models. The Stata Journal\/ 1\/ (1), 51--57
2001
-
[17]
Lyddon, S. P., C. C. Holmes, and S. G. Walker (2019). General bayesian updating and the loss-likelihood bootstrap. Biometrika\/ 106\/ (2), 465--478
2019
-
[18]
Knoblauch, F.-X
Matsubara, T., J. Knoblauch, F.-X. Briol, and C. J. Oates (2022). Robust generalised bayesian inference for intractable likelihoods. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 84\/ (3), 997--1022
2022
-
[19]
Knoblauch, F.-X
Matsubara, T., J. Knoblauch, F.-X. Briol, and C. J. Oates (2024). Generalized bayesian inference for discrete intractable likelihood. Journal of the American Statistical Association\/ 119\/ (547), 2345--2355
2024
-
[20]
Miller, J. W. (2021). Asymptotic normality, concentration, and coverage of generalized posteriors. Journal of Machine Learning Research\/ 22\/ (168), 1--53
2021
-
[21]
Nakagawa, T. and S. Hashimoto (2020). Robust bayesian inference via -divergence. Communications in Statistics-Theory and Methods\/ 49\/ (2), 343--360
2020
-
[22]
Onizuka, T. and S. Hashimoto (2025). Robust bayesian graphical modeling using -divergence. Journal of Multivariate Analysis\/ 209 , 105461
2025
-
[23]
Hashimoto, and S
Onizuka, T., S. Hashimoto, and S. Sugasawa (2024). Fast and locally adaptive bayesian quantile smoothing using calibrated variational approximations. Statistics and Computing\/ 34\/ (1), 15
2024
-
[24]
Rigon, T., A. H. Herring, and D. B. Dunson (2023). A generalized bayes framework for probabilistic clustering. Biometrika\/ 110\/ (3), 559--578
2023
-
[25]
Syring, N. and R. Martin (2019). Calibrating general posterior credible regions. Biometrika\/ 106\/ (2), 479--486
2019
-
[26]
Van der Vaart, A. W. (2000). Asymptotic statistics , Volume 3. Cambridge university press
2000
-
[27]
Wu, P.-S. and R. Martin (2023). A comparison of learning rate selection methods in generalized bayesian inference. Bayesian Analysis\/ 18\/ (1), 105--132
2023
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.