Pith. sign in

REVIEW 2 minor 12 references

When can a posterior predictive check identify the learning rate? Exact degeneracy in Gaussian models and implications for Generalised Bayesian Inerence

T0 review · 0 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read In Gaussian linear models the PPC selector for the learning rate is invariant to η or fixed before seeing data.

desk verdict The paper gives exact finite-sample derivations showing the PPC selector for the learning rate is data-independent or η-invariant in the Gaussian linear model under flat and reference priors. read the letter →

arxiv 2606.07169 v1 pith:GUVD2JL7 submitted 2026-06-05 stat.ME math.STstat.TH

classification stat.MEmath.STstat.TH
keywords posteriorpredictivechecklearningrateselectiongeneralisedBayesianinferenceGaussianlinearmodelpivotalitymisspecificationreferenceprior
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper shows that selecting the learning rate η via a posterior predictive check p-value on the log-likelihood fails to work in the Gaussian linear model. With known variance and a flat prior the p-value equals P(χ²_n > RSS/σ₀²) for every η, so every candidate produces the same result. With unknown variance and the reference prior the p-value depends only on n, d and η, making the selector output the same no matter what data arrive. The degeneracy is traced to a pivotality property that holds only for the Gaussian scale-location family under these specific priors and vanishes with informative priors. The analysis supplies an exact finite-sample diagnosis and a cheap pre-screening test to detect when the selector cannot function.

What carries the argument

The pivotality property of the Gaussian scale-location family under the flat or reference prior, which renders the log-likelihood PPC p-value independent of η or of the data.

What would settle it

Compute the PPC p-value across a grid of η values on data generated from a non-Gaussian model or under an informative prior and check whether the p-value changes with η or with the realised sample.

Watch

Extended reading notes

Core claim

With known variance and a flat prior the PPC p-value equals P(χ²_n > RSS/σ₀²) for every η, so the selector is η-invariant; under variance misspecification it is two-sided non-identifying. With unknown variance and the reference prior the p-value depends only on (n,d,η) and not on the realised data or the data-generating process, so the selector output is fixed before any data are seen and typically collapses to the smallest grid value.

Load-bearing premise

The data-generating process belongs to the Gaussian linear model family together with either a flat prior or the reference prior.

Editorial extensions

If this is right

  • The selector cannot identify η inside the Gaussian linear model with known variance.
  • Under variance misspecification the selector is two-sided non-identifying.
  • With unknown variance the selector output is predetermined by n, d and the grid and collapses to the smallest value.
  • The chosen η over-tempers relative to held-out selection and inflates predictive intervals.
  • The degeneracy disappears once an informative prior replaces the reference prior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same pivotality may appear in other location-scale families that admit exact marginal likelihoods under conjugate priors.
  • A data-free pre-screening step can be run on any new model class before committing to the PPC selector.
  • Alternative selection criteria that avoid the log-likelihood PPC may remain identifiable even when this particular PPC fails.
  • The results motivate checking whether the p-value expression factors through sufficient statistics alone in other exponential-family models.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The paper provides an exact finite-sample analysis of the PPC-based selector for the learning rate η proposed by Zafar and Nicholls (2024) when applied to the Gaussian linear model. With known variance and a flat prior, the log-likelihood PPC p-value equals P(χ²_n > RSS/σ₀²) for every η, rendering the selector η-invariant (and two-sided non-identifying under variance misspecification). With unknown variance and the reference prior, the p-value depends only on (n,d,η) and is independent of the data and data-generating process, so the selector output is fixed in advance and typically selects the smallest grid value. The results are presented as a pivotality property specific to the Gaussian scale-location family under these priors, vanishing under informative priors, and the paper motivates a data-free pre-screening diagnostic.

Significance. If the algebraic derivations hold, the work is significant for precisely delineating the scope of the PPC selector by exhibiting a canonical model class on which it cannot identify η. Strengths include the exact finite-sample character of the results (no asymptotics or approximations), explicit statement of scope limits, and the data-free diagnostic. These findings are useful for practitioners and for understanding when tempering via η can or cannot be calibrated by PPC.

minor comments (2)
  1. [Abstract] The abstract states that the selector 'typically collapsing to the smallest grid value'; a brief numerical illustration with concrete (n,d) values and grid would make this concrete without lengthening the paper.
  2. [Abstract] Notation for the reference prior and the exact form of the PPC statistic could be cross-referenced to the main text in the abstract for readers who encounter the claim first.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their accurate summary of the manuscript, for highlighting its strengths, and for recommending acceptance. No major comments were raised in the report.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; results are direct algebraic derivations

full rationale

The paper derives exact finite-sample invariance of the PPC p-value to η (known variance, flat prior) and sole dependence on (n,d,η) (unknown variance, reference prior) via properties of the Gaussian linear model and chi-squared pivots. These follow immediately from the model assumptions and prior choice without any fitted parameters renamed as predictions, self-definitional loops, or load-bearing self-citations. The scope is explicitly limited to this family and priors, with the result vanishing under informative priors. No reduction of the central claim to its own inputs occurs.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claim rests on the Gaussian linear model assumption together with the flat or reference prior; these are standard domain assumptions that produce the pivotality property. No free parameters or invented entities are introduced.

assumptions (1)
  • domain assumption Data follow a Gaussian linear model with either known variance and flat prior or unknown variance and reference prior.
    The abstract derives the p-value expressions and independence properties under precisely these model and prior specifications.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When can a posterior predictive check identify the learning rate? Exact degeneracy in Gaussian models and implications for Generalised Bayesian Inerence." pith.science (2026). https://pith.science/paper/GUVD2JL7

@misc{pith2026260607169,
  author       = {Pith},
  title        = {Pith review of: When can a posterior predictive check identify the learning rate? Exact degeneracy in Gaussian models and implications for Generalised Bayesian Inerence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GUVD2JL7}},
  note         = {Machine review of arXiv:2606.07169}
}
abstract

Generalised Bayesian inference tempers the likelihood by a learning rate $\eta$ to mitigate model misspecification, and the choice of $\eta$ is consequential. Zafar and Nicholls (2024) proposed selecting $\eta$ by a posterior predictive check (PPC): one chooses the smallest $\eta$ at which a log-likelihood PPC $p$-value is not rejected. An exact, finite-sample analysis of this selector on the Gaussian linear model is given. With known variance and a flat prior, the PPC $p$-value equals $P(\chi^2_n > \mathrm{RSS}/\sigma_0^2)$ for every $\eta$, so the selector is $\eta$-invariant; under variance misspecification it is two-sided non-identifying. With unknown variance and the reference prior, the $p$-value depends only on $(n,d,\eta)$ and not on the realised data or the data-generating process. Consequently the selector's output is fixed before any data are seen, typically collapsing to the smallest grid value, which over-tempers and inflates predictive intervals relative to held-out selection. The phenomenon is a pivotality property specific to the Gaussian scale--location family and the reference prior; it disappears under informative priors. These results delineate the selector's scope, identify a canonical class on which it cannot identify the learning rate, and motivate a cheap, data-free pre-screening diagnostic.

Figures

Figures reproduced from arXiv: 2606.07169 by the authors.

Figure 1
Figure 1. Known variance. (A) The PPC 𝑝-value is exactly flat in 𝜂 (Theorem 1); each line is a different variance ratio 𝑟. (B) Two-sided non-identifiability: 𝑝 → 1 for 𝑟 < 1 and 𝑝 → 0 for 𝑟 > 1 (Corollary 1). 0.5 1.0 1.5 2.0 2.5 3.0 learning rate η 0.0 0.2 0.4 0.6 0.8 1.0 PPC p-value Unknown variance: pn, d(η) is data-free (Theorem 2) MC points (any DGP) land on the exact curve exact, n=30, d=1 exact, n=30, d=5 exact, n=100, … view at source ↗
Figure 2
Figure 2. Unknown variance. Solid: the exact data-free curve [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. (A) The PPC-selected 𝜂 is data-free and equals the grid floor across most of the (𝑛, 𝑑) plane. (B) min𝜂∈𝒢+ 𝑝𝑛,𝑑(𝜂) stays above the level 𝛼 = 0.10, forcing grid-floor collapse. Gaussian Student-t3 contam 0.800 0.825 0.850 0.875 0.900 0.925 0.950 0.975 1.000 90% predictive coverage (A) Coverage by selection rule nominal 0.90 PPC (η = 0.10) held-out η = 1 Gaussian Student-t3 contam 0 2 4 6 8 10 12 mean predictive inter… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Calibration cost (𝑛 = 80, 𝑑 = 3, 90% intervals). The data-free PPC choice 𝜂 = 0.1 over-covers and inflates width in every DGP, whereas held-out selection and 𝜂 = 1 track the nominal level. 6 A practical diagnostic, and discussion Pre-screening diagnostic. Theorem 2 sug…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

12 extracted references · 4 canonical work pages

  1. [1]

    G., Holmes, C

    Bissiri, P. G., Holmes, C. C., Walker, S. G. (2016). A general framework for updating belief distributions.J. R. Stat. Soc. B78(5), 1103–1130

  2. [2]

    Gelman, A., Meng, X.-L., Stern, H. (1996). Posterior predictive assessment of model fitness via realized discrepancies. Statist. Sinica6, 733–807. Grünwald, P., van Ommen, T. (2017). Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it.Bayesian Anal.12(4)

  3. [3]

    C., Walker, S

    Holmes, C. C., Walker, S. G. (2017). Assigning a value to a power likelihood in a general Bayesian model.Biometrika 104(2), 497–503

  4. [4]

    Bayesian inference for the learning rate in Generalised Bayesian inference

    Lee, J. E., Liu, S., Nicholls, G. K. (2025). Bayesian inference for the learning rate in generalised Bayesian inference. arXiv:2506.12532

  5. [5]

    P., Holmes, C

    Lyddon, S. P., Holmes, C. C., Walker, S. G. (2019). General Bayesian updating and the loss-likelihood bootstrap. Biometrika106(2), 465–478

  6. [6]

    T., Knoblauch, J

    McLatchie, Y., Fong, E., Frazier, D. T., Knoblauch, J. (2024). Predictive performance of power posteriors. arXiv:2408.08806

  7. [7]

    Meng, X.-L. (1994). Posterior predictive𝑝-values.Ann. Statist.22(3), 1142–1160

  8. [8]

    M., van der Vaart, A., Ventura, V

    Robins, J. M., van der Vaart, A., Ventura, V. (2000). Asymptotic distribution of𝑝-values in composite null models.J. Amer. Statist. Assoc.95(452), 1143–1156

Show all 12 references
  1. [9]

    Rubin-Delanchy, P., Lawson, D. J. (2015). Posterior predictive𝑝-values and the convex order.arXiv:1412.3442

  2. [10]

    Syring, N., Martin, R. (2019). Calibrating general posterior credible regions.Biometrika106(2), 479–486

  3. [11]

    Wu, P.-S., Martin, R. (2023). A comparison of learning rate selection methods in generalized Bayesian inference. Bayesian Anal.18(1), 105–132

  4. [12]

    Zafar, S., Nicholls, G. K. (2024). Exploring learning rate selection in generalised Bayesian inference using posterior predictive checks.arXiv:2410.01475. 6

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.