REVIEW 2 major objections 4 minor 18 references
Asymptotic exchangeability — the shifted tail of a data sequence converging in distribution to an exchangeable sequence — is far too weak to license treating the tail as approximately exchangeable or as inferentially equivalent to its excha
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 13:07 UTC pith:POPN32NF
load-bearing objection Asymptotic exchangeability is indeed too weak for inference — the warning is correct, but Example 7's unproved 'it can be shown' claims need proofs or references before this becomes a citable caution. the 2 major comments →
Some cautionary tales about Bayesian predictive inference
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that asymptotic exchangeability, defined as convergence in distribution of (X_{n+1}, X_{n+2}, ...) to an exchangeable sequence Z, is far too weak to justify common Bayesian interpretations: that for large n the tail is approximately i.i.d. given some random measure, or that all sequences with the same exchangeable limit are inferentially equivalent. The main demonstration is a construction with X_n = n^{-1/2} (U_1 + ... + U_n) + V_n, where U and V are independent standard-normal i.i.d. sequences. Its shifted tail converges in distribution to Z = (U_1+V_1, U_1+V_2, ...), an exchangeable sequence, yet the paper argues that the empirical mean of X fails to converge
What carries the argument
The relation X∼Z, in which two data sequences share the same exchangeable limit sequence in distribution, is the object whose inferential weight is interrogated. The counterexample engine is the predictive construction X_n = n^{-1/2} sum_{i=1}^n U_i + V_n with independent standard-normal U and V sequences; its tail has the same distributional limit as Z = (U_1+V_1, U_1+V_2, ...), but its empirical mean, probability law, and predictive distributions behave differently from Z's. A supporting mechanism compares two predictive sequences that are asymptotically close in a probability metric and shows that they can still induce very different data laws.
Load-bearing premise
The argument's main demonstration rests on three assertions in Example 7 that the paper states with 'it can be shown' and does not prove: the empirical mean of the constructed X fails to converge in probability, X and the exchangeable limit Z have mutually singular distributions, and X's predictives fail to converge weakly; if any of these assertions is false, the claim that asymptotic exchangeability does not even imply the law of large numbers loses its key support.
What would settle it
Check the unproved part of Example 7 directly: compute or simulate whether Xbar_n = (1/n) sum_{i=1}^n (i^{-1/2} sum_{k=1}^i U_k + V_i) converges in probability. If it does converge, the paper's main demonstration fails; if it provably does not, the cautionary conclusion stands. The covariance calculation alone cannot settle the convergence question.
If this is right
- Verifying that the data tail converges in distribution to an exchangeable sequence is not enough to justify Bayesian updates that assume near-exchangeability of the tail; a stronger mode of predictive convergence is needed.
- The law of large numbers for the original data does not follow from asymptotic exchangeability: the paper's construction has an exchangeable limit while its empirical mean fails to converge in probability.
- Asymptotic credible intervals computed from the exchangeable limit sequence can have wrong coverage for the original data, because the studentized CLT that holds for exchangeable sequences can fail along X.
- Two predictive sequences that are asymptotically close in a metric can still generate qualitatively different data sequences, so closeness of predictives is not a license to substitute one model for another.
- Stronger conditions, such as total-variation convergence of the predictives or of the shifted sequence law, would restore the approximate-exchangeability interpretation.
Where Pith is reading between the lines
- The covariance calculation suggests a practical diagnostic: before applying an exchangeable-limit approximation, estimate Cov(X_{n+i}, X_{n+j}) for several lags and check whether it stabilizes as n grows; the pattern of covariances drifting toward both 0 and 1 would warn against the approximation.
- A testable extension is a simulation study that draws data from the paper's construction, fits the exchangeable limit model, and measures the actual coverage of nominal credible intervals, which would quantify the finite-sample gap.
- The paper's strengthenings point to a program: characterize, for each inferential target (law of large numbers, central limit theorem, credible intervals), the minimal mode of predictive convergence required; total variation is sufficient but likely not necessary.
- The two views of the data-generating mechanism imply a terminology fix for applied work: reserve 'predictive distribution' for the actual conditional law of the next observation given the past, and call estimators of an unknown sampling measure something else.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a short cautionary note on two misunderstandings in Bayesian predictive inference. The first (M1) is the conflation of two views of the data sequence: (i) P is constructed from a chosen sequence of predictive distributions via the Ionescu-Tulcea theorem, and (ii) P is an unknown structured law, with the chosen β_n serving as estimators of an unknown random measure. The second (M2) is the overinterpretation of asymptotic exchangeability: the belief that X∼Z (both shifted sequences converge to the same exchangeable limit) makes X and Z inferentially equivalent. The paper illustrates M1 with bootstrap, Hill's predictive distributions, Bayesian nonparametrics, and predictive resampling, and M2 with examples showing differences in SLLN behavior, CLT behavior, predictive convergence, and the effect of asymptotically equivalent predictives. It closes by proposing two possible strengthenings of asymptotic exchangeability.
Significance. If the mathematical claims are correct, the note fills a useful didactic gap: it crisply separates two meanings of 'predictive distribution' and gives a sharp warning that asymptotic exchangeability is too weak to license de Finetti-type conditional i.i.d. interpretations or asymptotic inferential equivalence. The examples are simple and readable; the covariance calculation in Example 7 is a transparent and correct demonstration. The paper makes no fitted-parameter claims and is honest about relying on published results, notably [2] and [16]. However, several headline claims in Example 7 are asserted without proof, and the final characterization of a total-variation strengthening is incorrect as stated. These issues are local and fixable, but they prevent the note from being citable in its current form.
major comments (2)
- [§2, Example 7] The three load-bearing facts in Example 7 are each introduced by 'it can be shown' with no proof and only a vague reference to Example 4.6 of [16]: (i) X̄_n fails to converge in probability, (ii) X and Z have singular probability distributions, and (iii) the predictives of Z converge weakly while those of X do not. Fact (i) is the paper's demonstration that asymptotic exchangeability does not even imply the SLLN; fact (ii) is the sharpest evidence of inferential non-equivalence. These claims are not immediate from the displayed covariance calculation, which addresses tail exchangeability only. I have verified independently that the claims are mathematically correct, but a citable note should either prove them in an appendix or supply an exact theorem/example reference for each. As written, the central M2 conclusion rests on unverifiable assertions.
- [§2, final remark (condition (ii))] The asserted equivalence for condition (ii) is false in general. Let Z be i.i.d. N(0,1) and set X_n = Z_n + a_n with a_n = 1/√n. Then X is asymptotically exchangeable with limit Z, and σ(X_{n+1},...) = σ(Z_{n+1},...) for every n, so the two laws agree on the trivial tail σ-field. Yet for each n the law of (X_{n+1}, X_{n+2}, ...) is an infinite product of N(a_{n+k},1) measures and is mutually singular with the law of Z by Kakutani's theorem, since Σ_k a_{n+k}^2 = ∞. Hence ∥ν_n - P_Z∥ = 1 for every n, not 0. The 'only if' direction is correct, but the converse needs additional hypotheses or should be withdrawn.
minor comments (4)
- [§2, Example 9] The displayed formula for α_n^* is garbled; it should presumably read (c α_0 + n μ_n)/(n+c). Please fix.
- [§2, Example 6] The abbreviation 'c.i.d.' is used without definition. Define it as 'conditionally identically distributed' and give the precise definition or a reference to [2].
- [§2, Example 7] The covariance formula Cov(X_{n+i}, X_{n+j}) = √((n+i)/(n+j)) is correct, but a one-line derivation from the independence of U and V and the variance of partial sums would help the reader follow the point.
- [§1.2] The definition of the relation 'Y∼X' uses the same symbol as the example 'X∼Z'; this is perhaps unavoidable, but consider an explicit phrase to avoid confusion with 'approximately distributed'.
Circularity Check
No circular derivation: the paper's examples are explicit constructions backed by published, parameter-free prior results; Example 7's unproved 'it can be shown' claims are a verifiability gap, not a circular reduction.
full rationale
No circularity is present. The paper contains no fitted parameters, no data-fitting step, and no quantity that is fitted to a subset of data and then renamed as a prediction. The M2 thesis—that asymptotic exchangeability is too weak to license treating the tail as approximately exchangeable—is supported by explicit counterexamples, not by a definitional identification. The load-bearing citations, Theorem 2.2 and Example 4.6 of [16] and Theorem 3.1/Example 3.2 of [2], are published, peer-reviewed, parameter-free mathematical results with stated assumptions that do not include the present conclusion; under the review rules these are independent evidence and do not raise the circularity score. The main caveat is in Example 7, where the assertions 'Firstly, it can be shown that \(\bar X_n = \frac1n \sum_{i=1}^n X_i\) fails to converge in probability', 'X and Z have singular probability distributions', and 'the predictive distributions of Z converge (weakly a.s.) while those of X do not' are introduced by 'it can be shown' with no proof or precise reference. These are load-bearing for the M2 demonstration, and the manuscript should supply proofs or references before it is cited as self-contained. But this is an omitted-proof/verifiability gap, not circularity: the claims are direct properties of the constructed Gaussian sequence \(X_n = \sum_{i=1}^n U_i/\sqrt n + V_n\), not restatements of the definition of asymptotic exchangeability, and no equation in the paper reduces the conclusion to its inputs. No other circularity pattern (self-definition, fitted input called prediction, uniqueness imported from authors, ansatz smuggled via citation, renaming of a known result) occurs.
Axiom & Free-Parameter Ledger
free parameters (2)
- c (Example 9 concentration constant) =
arbitrary positive constant
- q (Example 6 weighting constant) =
arbitrary in (0,1)
axioms (6)
- standard math Ionescu-Tulcea theorem: a law on R^∞ is uniquely determined by arbitrary predictives (α_n)
- standard math de Finetti's representation theorem for exchangeable sequences
- standard math Aldous Lemma (8.2): α_n → μ weakly a.s. implies asymptotic exchangeability
- standard math Published self-cited theorems: Theorem 2.2 of [16] (weak convergence in probability suffices), Theorem 3.1 of [2] (stable CLT for exchangeable Z), Example 3.2 of [2] (CLT failure for c.i.d. X)
- domain assumption A.s. existence of θ = lim_n (1/n)Σ_{i≤n} X_i
- domain assumption A common probability space (Ω, A, P) carrying all sequences
read the original abstract
Two misunderstandings, frequently arising in Bayesian predictive inference, are discussed. The first deals with the data generating mechanism, while the second consists in overestimating the role played by asymptotic exchangeability. Some consequences of such misunderstandings are highlighted through examples.
Reference graph
Works this paper leans on
-
[1]
(1985) Exchangeability and related topics,Ecole d’et´ e de Probabilit´ es de Saint- Flour XIII, Lecture Notes in Math
Aldous D.J. (1985) Exchangeability and related topics,Ecole d’et´ e de Probabilit´ es de Saint- Flour XIII, Lecture Notes in Math. 1117, Springer
1985
-
[2]
(2004) Limit theorems for a class of identically distributed random variables,Ann
Berti P., Pratelli L., Rigo P. (2004) Limit theorems for a class of identically distributed random variables,Ann. Probab., 32, 2029-2052
2004
-
[3]
(2013) Exchangeable sequences driven by an absolutely contin- uous random measure,Ann
Berti P., Pratelli L., Rigo P. (2013) Exchangeable sequences driven by an absolutely contin- uous random measure,Ann. Probab., 41, 2090-2102
2013
-
[4]
(2025) A probabilistic view on predictive constructions for Bayesian learning,Statist
Berti P., Dreassi E., Leisen F., Pratelli L., Rigo P. (2025) A probabilistic view on predictive constructions for Bayesian learning,Statist. Sci., 40, 25-39
2025
-
[5]
Bissiri P.G., Holmes C., Walker S.G. (2026) Posterior Inference via Hill’s Prediction Model, arXiv:2603.20071v1 8 EMANUELA DREASSI, F ABRIZIO LEISEN, LUCA PRATELLI, AND PIETRO RIGO
arXiv 2026
-
[6]
(2026) Posterior uncertainty for kernel density estimates, arXiv:2607.03927v1
Christensen D., Svardal T., Rønneberg L., Stoltenberg E.A. (2026) Posterior uncertainty for kernel density estimates, arXiv:2607.03927v1
Pith/arXiv arXiv 2026
-
[7]
(1999) Prequential probability: principles and properties,Bernoulli, 5, 125-162
Dawid A.P., Vovk V.G. (1999) Prequential probability: principles and properties,Bernoulli, 5, 125-162
1999
-
[8]
(2023) Neural diffusion processes,Proc
Dutordoir V., Saul A., Ghahramani Z., Simpson F. (2023) Neural diffusion processes,Proc. 40th Intern. Conf. on Machine Learning, PMLR 202, 8990-9012
2023
-
[9]
(2023) Martingale posterior distributions (with discussion), J
Fong E., Holmes C., Walker S.G. (2023) Martingale posterior distributions (with discussion), J. Royal Stat. Soc. B, 85, 1357-1391
2023
-
[10]
(2000) Exchangeability, predictive distributions and para- metric models,Sankhya A, 62, 86–109
Fortini S., Ladelli L., Regazzini E. (2000) Exchangeability, predictive distributions and para- metric models,Sankhya A, 62, 86–109
2000
-
[11]
(2012) Predictive construction of priors in Bayesian nonparametrics, Brazilian J
Fortini S., Petrone S. (2012) Predictive construction of priors in Bayesian nonparametrics, Brazilian J. Prob. Stat., 26, 423-449
2012
-
[12]
(2020) Quasi-Bayes properties of a procedure for sequential learning in mixture models,J
Fortini S., Petrone S. (2020) Quasi-Bayes properties of a procedure for sequential learning in mixture models,J. Royal Stat. Soc. B, 82, 1087-1114
2020
-
[13]
(2025) Exchangeability, prediction and predictive modeling in Bayesian statistics,Statist
Fortini S., Petrone S. (2025) Exchangeability, prediction and predictive modeling in Bayesian statistics,Statist. Sci., 40, 40–67
2025
-
[14]
(2018) On recursive Bayesian predictive distributions, J.A.S.A., 113, 1085-1093
Hahn P.R., Martin R., Walker S.G. (2018) On recursive Bayesian predictive distributions, J.A.S.A., 113, 1085-1093
2018
-
[15]
(1968) Posterior distribution of percentiles: Bayes’ theorem for sampling from a population,J.A.S.A., 63, 677–691
Hill B.M. (1968) Posterior distribution of percentiles: Bayes’ theorem for sampling from a population,J.A.S.A., 63, 677–691
1968
-
[16]
(2026) Weak convergence in probability of predictive distribu- tions,Electronic Commun
Leisen F., Pratelli L., Rigo P. (2026) Weak convergence in probability of predictive distribu- tions,Electronic Commun. Probab., 31, 1-15
2026
-
[17]
(2023) Bayesian causal inference: A critical review,Phil
Li F., Ding P., Mealli F. (2023) Bayesian causal inference: A critical review,Phil. Trans. R. Soc.A, 381
2023
-
[18]
G. Par- enti
Pitman J. (1996) Some developments of the Blackwell-MacQueen urn scheme,Statistics, Probability and Game Theory, IMS Lect. Notes Mon. Series, 30, 245-267. Emanuela Dreassi, Dipartimento di Statistica, Informatica, Applicazioni “G. Par- enti”, Universit`a di Firenze, viale Morgagni 59, 50134 Firenze, Italy Email address:emanuela.dreassi@unifi.it F abrizio ...
1996
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.