REVIEW 2 minor 12 references
When can a posterior predictive check identify the learning rate? Exact degeneracy in Gaussian models and implications for Generalised Bayesian Inerence
T0 review · 0 major / 2 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read In Gaussian linear models the PPC selector for the learning rate is invariant to η or fixed before seeing data.
desk verdict The paper gives exact finite-sample derivations showing the PPC selector for the learning rate is data-independent or η-invariant in the Gaussian linear model under flat and reference priors. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The pivotality property of the Gaussian scale-location family under the flat or reference prior, which renders the log-likelihood PPC p-value independent of η or of the data.
What would settle it
Compute the PPC p-value across a grid of η values on data generated from a non-Gaussian model or under an informative prior and check whether the p-value changes with η or with the realised sample.
Extended reading notes
Core claim
With known variance and a flat prior the PPC p-value equals P(χ²_n > RSS/σ₀²) for every η, so the selector is η-invariant; under variance misspecification it is two-sided non-identifying. With unknown variance and the reference prior the p-value depends only on (n,d,η) and not on the realised data or the data-generating process, so the selector output is fixed before any data are seen and typically collapses to the smallest grid value.
Load-bearing premise
The data-generating process belongs to the Gaussian linear model family together with either a flat prior or the reference prior.
Editorial extensions
If this is right
- The selector cannot identify η inside the Gaussian linear model with known variance.
- Under variance misspecification the selector is two-sided non-identifying.
- With unknown variance the selector output is predetermined by n, d and the grid and collapses to the smallest value.
- The chosen η over-tempers relative to held-out selection and inflates predictive intervals.
- The degeneracy disappears once an informative prior replaces the reference prior.
Reading between the lines
- The same pivotality may appear in other location-scale families that admit exact marginal likelihoods under conjugate priors.
- A data-free pre-screening step can be run on any new model class before committing to the PPC selector.
- Alternative selection criteria that avoid the log-likelihood PPC may remain identifiable even when this particular PPC fails.
- The results motivate checking whether the p-value expression factors through sufficient statistics alone in other exponential-family models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper provides an exact finite-sample analysis of the PPC-based selector for the learning rate η proposed by Zafar and Nicholls (2024) when applied to the Gaussian linear model. With known variance and a flat prior, the log-likelihood PPC p-value equals P(χ²_n > RSS/σ₀²) for every η, rendering the selector η-invariant (and two-sided non-identifying under variance misspecification). With unknown variance and the reference prior, the p-value depends only on (n,d,η) and is independent of the data and data-generating process, so the selector output is fixed in advance and typically selects the smallest grid value. The results are presented as a pivotality property specific to the Gaussian scale-location family under these priors, vanishing under informative priors, and the paper motivates a data-free pre-screening diagnostic.
Significance. If the algebraic derivations hold, the work is significant for precisely delineating the scope of the PPC selector by exhibiting a canonical model class on which it cannot identify η. Strengths include the exact finite-sample character of the results (no asymptotics or approximations), explicit statement of scope limits, and the data-free diagnostic. These findings are useful for practitioners and for understanding when tempering via η can or cannot be calibrated by PPC.
minor comments (2)
- [Abstract] The abstract states that the selector 'typically collapsing to the smallest grid value'; a brief numerical illustration with concrete (n,d) values and grid would make this concrete without lengthening the paper.
- [Abstract] Notation for the reference prior and the exact form of the PPC statistic could be cross-referenced to the main text in the abstract for readers who encounter the claim first.
Simulated Author's Rebuttal
We thank the referee for their accurate summary of the manuscript, for highlighting its strengths, and for recommending acceptance. No major comments were raised in the report.
Circularity Check
No significant circularity; results are direct algebraic derivations
full rationale
The paper derives exact finite-sample invariance of the PPC p-value to η (known variance, flat prior) and sole dependence on (n,d,η) (unknown variance, reference prior) via properties of the Gaussian linear model and chi-squared pivots. These follow immediately from the model assumptions and prior choice without any fitted parameters renamed as predictions, self-definitional loops, or load-bearing self-citations. The scope is explicitly limited to this family and priors, with the result vanishing under informative priors. No reduction of the central claim to its own inputs occurs.
Assumptions & free parameters
assumptions (1)
- domain assumption Data follow a Gaussian linear model with either known variance and flat prior or unknown variance and reference prior.
Cite this review
Pith. "Pith review of When can a posterior predictive check identify the learning rate? Exact degeneracy in Gaussian models and implications for Generalised Bayesian Inerence." pith.science (2026). https://pith.science/paper/GUVD2JL7
@misc{pith2026260607169,
author = {Pith},
title = {Pith review of: When can a posterior predictive check identify the learning rate? Exact degeneracy in Gaussian models and implications for Generalised Bayesian Inerence},
year = {2026},
howpublished = {\url{https://pith.science/paper/GUVD2JL7}},
note = {Machine review of arXiv:2606.07169}
}
abstract
Generalised Bayesian inference tempers the likelihood by a learning rate $\eta$ to mitigate model misspecification, and the choice of $\eta$ is consequential. Zafar and Nicholls (2024) proposed selecting $\eta$ by a posterior predictive check (PPC): one chooses the smallest $\eta$ at which a log-likelihood PPC $p$-value is not rejected. An exact, finite-sample analysis of this selector on the Gaussian linear model is given. With known variance and a flat prior, the PPC $p$-value equals $P(\chi^2_n > \mathrm{RSS}/\sigma_0^2)$ for every $\eta$, so the selector is $\eta$-invariant; under variance misspecification it is two-sided non-identifying. With unknown variance and the reference prior, the $p$-value depends only on $(n,d,\eta)$ and not on the realised data or the data-generating process. Consequently the selector's output is fixed before any data are seen, typically collapsing to the smallest grid value, which over-tempers and inflates predictive intervals relative to held-out selection. The phenomenon is a pivotality property specific to the Gaussian scale--location family and the reference prior; it disappears under informative priors. These results delineate the selector's scope, identify a canonical class on which it cannot identify the learning rate, and motivate a cheap, data-free pre-screening diagnostic.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
G., Holmes, C
Bissiri, P. G., Holmes, C. C., Walker, S. G. (2016). A general framework for updating belief distributions.J. R. Stat. Soc. B78(5), 1103–1130
2016
-
[2]
Gelman, A., Meng, X.-L., Stern, H. (1996). Posterior predictive assessment of model fitness via realized discrepancies. Statist. Sinica6, 733–807. Grünwald, P., van Ommen, T. (2017). Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it.Bayesian Anal.12(4)
1996
-
[3]
C., Walker, S
Holmes, C. C., Walker, S. G. (2017). Assigning a value to a power likelihood in a general Bayesian model.Biometrika 104(2), 497–503
2017
-
[4]
Bayesian inference for the learning rate in Generalised Bayesian inference
Lee, J. E., Liu, S., Nicholls, G. K. (2025). Bayesian inference for the learning rate in generalised Bayesian inference. arXiv:2506.12532
work page Pith review arXiv 2025
-
[5]
P., Holmes, C
Lyddon, S. P., Holmes, C. C., Walker, S. G. (2019). General Bayesian updating and the loss-likelihood bootstrap. Biometrika106(2), 465–478
2019
-
[6]
McLatchie, Y., Fong, E., Frazier, D. T., Knoblauch, J. (2024). Predictive performance of power posteriors. arXiv:2408.08806
-
[7]
Meng, X.-L. (1994). Posterior predictive𝑝-values.Ann. Statist.22(3), 1142–1160
1994
-
[8]
M., van der Vaart, A., Ventura, V
Robins, J. M., van der Vaart, A., Ventura, V. (2000). Asymptotic distribution of𝑝-values in composite null models.J. Amer. Statist. Assoc.95(452), 1143–1156
2000
Show all 12 references
-
[9]
Rubin-Delanchy, P., Lawson, D. J. (2015). Posterior predictive𝑝-values and the convex order.arXiv:1412.3442
2015 arXiv
-
[10]
Syring, N., Martin, R. (2019). Calibrating general posterior credible regions.Biometrika106(2), 479–486
2019
-
[11]
Wu, P.-S., Martin, R. (2023). A comparison of learning rate selection methods in generalized Bayesian inference. Bayesian Anal.18(1), 105–132
2023
-
[12]
Zafar, S., Nicholls, G. K. (2024). Exploring learning rate selection in generalised Bayesian inference using posterior predictive checks.arXiv:2410.01475. 6
2024
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.