REVIEW 2 major objections 5 minor 2 references
This paper shows that any Bayesian decision problem generates four interchangeable coherent functions—uncertainty, scoring rule, discrepancy, and dependence—and that Bayesian predictive experimental design can be based on any of them with i
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A decision problem induces equivalent coherent measures of uncertainty, scoring, discrepancy, and dependence; design criteria based on any of them coincide.
T0 review reviewed 2026-08-02 challenge →
load-bearing objection A 1998 unification of scoring rules, concave uncertainty, discrepancies, and expected value of sample information—strong on characterizations, conditional on a presentability assumption that the paper openly flags. the 2 major comments →
Coherent Measures of Discrepancy, Uncertainty and Dependence, with Applications to Bayesian Predictive Experimental Design
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that coherence is a single decision-theoretic property wearing four hats. Starting from a loss L, define H(P) = inf_a E_P L(Z,a); set S(z,Q) = L(z,a(Q)) with a(Q) a Bayes act; set D(P,Q) = S(P,Q) - S(P,P); and set C as the expected reduction in Bayes loss from observing an auxiliary variable. The paper proves that each of these is coherent, that H is concave, that D(P,Q) - D(P,Q0) is affine in P, and—under the presentability assumption—conversely any concave H and any discrepancy with affine differences can be realized from some decision problem. It follows that design criteria based on any of the four coincide: choose the experiment minimizing expected posterior uncerta
What carries the argument
The central mechanism is the chain of identities linking a terminal decision problem to its derived functions. The key identity is D(P,Q) = S(P,Q) - S(P,P), which converts any proper scoring rule S into a discrepancy, and conversely any coherent discrepancy arises this way from the scoring rule built from Bayes actions. For uncertainty, the load-bearing condition is the supporting-hyperplane identity H(P) ≤ E_P[f_Q(Z)] with equality at P=Q and f_Q presentable; this turns a concave H into a proper scoring rule S(z,Q) = f_Q(z). For discrepancies, the characterizing identity is the affineness in P of D(P,Q) - D(P,Q0), which allows construction of S via presentability.
Load-bearing premise
The load-bearing premise is that every affine function on a family of distributions can be written as the expectation of some function of the outcome (presentability); the paper notes this can fail in general, and without it concavity alone need not imply coherence.
What would settle it
Find a concave function H on a convex family of distributions that has no presentable supporting hyperplane—equivalently, no proper scoring rule S with H(P)=S(P,P). If such an H exists, the claim that coherence is essentially concavity is false without extra conditions; the paper itself points to such a counterexample, so the meaningful falsifier is a concrete non-presentable concave H or a discrepancy D with affine differences but no scoring-rule realization.
If this is right
- Any Bayesian predictive design problem can be solved equivalently by minimizing expected posterior uncertainty, minimizing expected discrepancy between sampling and predictive distributions, or maximizing dependence between predictand and observation; the same experiment is optimal under all three.
- Every concave uncertainty function (with presentability) yields a proper scoring rule and hence a coherent design criterion, so concavity is the test for whether a proposed uncertainty criterion is Bayesian-coherent.
- A discrepancy is coherent exactly when its differences in the second argument are affine in the first; quadratic and characteristic-function discrepancies pass this test, so coherent discrepancies are plentiful, not unique.
- The expected value of sample information is a coherent dependence function, and its maximization is equivalent to uncertainty-minimizing design.
- The class of A-coherent discrepancies is broader than the Kullback-Leibler distance; the conjecture of uniqueness is false.
Where Pith is reading between the lines
- This equivalence suggests a practical recipe: to check whether any proposed design criterion is Bayesian, verify concavity (for uncertainty) or the affine-difference condition (for discrepancy) rather than constructing a full loss function.
- Because the characterization depends on presentability, in infinite-dimensional or non-parametric settings one should expect counterexamples; a robust version might require additional topological hypotheses.
- The concluding discussion points toward a differential-geometric reading: coherent discrepancies induce metrics on distribution space; if only coherent discrepancies are considered, the Fisher information metric may be characterizable as the unique coherent choice, giving Bayesian foundations for information geometry.
- For applied design, this means that criteria like expected entropy or expected quadratic loss are not competing philosophies but re-expressions of the same underlying expected-loss minimization, so empirical comparisons among them compare numerical approximations, not different objectives.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a decision-theoretic framework in which a terminal decision problem with loss L induces a coherent uncertainty function H(P)=inf_a E_P L(Z,a), a proper scoring rule S(z,Q)=L(z,a(Q)), a discrepancy D(P,Q)=S(P,Q)-H(P), and a dependence function C(P^{Z,U})=H(P^Z)-E H(P^{Z|U}). It shows that all of these yield equivalent formulations of Bayesian predictive experimental design. It then characterizes coherent uncertainty functions as concave functions satisfying a supporting-hyperplane/presentability condition, and coherent discrepancy functions as those for which D(P,Q)-D(P,Q0) is affine and presentable in P. It proves that coherent discrepancies are A-coherent, refuting Aitchison's conjecture that only the Kullback-Leibler distance has that property. Eight worked examples link the theory to variance, entropy, characteristic-function scores, quadratic (Brier) scores, exponential families, and D-optimality.
Significance. If correct, the paper gives a clean, unifying decision-theoretic account of many existing criteria, showing that proper scoring rules, entropy/divergences, mutual information, and Bayesian design criteria are facets of one structure. The characterization theorems, while building on prior work (Dawid 1986, 1994; Dawid & Sebastiani 1996; Hendrickson & Buehler 1971), are proved from the definitions and are illustrated by explicit verification in every example. The refutation of Aitchison's conjecture and the link to D-optimality are valuable. The manuscript is notably candid about the presentability condition and about the lack of a full dependence-function characterization, which is a strength.
major comments (2)
- [§10 (pp. 27-28), esp. (10.1)-(10.7)] The statement that 'coherence is essentially equivalent to concavity' is not made precise. The converse proof uses the supporting-hyperplane property (10.1) and presentability of the affine function φ_Q; the paper cites Hendrickson & Buehler (1971), Theorem 4.1, and Johnson (1991), but does not state the 'suitable additional technical conditions' promised in the Introduction. Since the abstract and summary treat this equivalence as a central result, please state those conditions explicitly, or clearly label the general equivalence as heuristic and rely on the direct verifications in the examples. This is a clarity issue rather than a mathematical error, because each example is checked directly.
- [§9 and §11, esp. (11.2)-(11.3)] The presentability assumption that 'every affine function on the domain P is presentable' is load-bearing for the converse characterization of coherent discrepancy functions and for constructing a scoring rule from a concave H in (10.4)-(10.7). The paper explicitly concedes that this assumption can fail (Hendrickson & Buehler 1971, Example 4.1). Consequently the Summary's claim that 'each function essentially determines the others' is stronger than the theorems as stated. I recommend adding an explicit caveat to the Summary/Abstract, and noting Lemma 9.1 as a sufficient condition. The design-coincidence results for the examples survive because presentability is either immediate (point masses) or checked directly.
minor comments (5)
- [Summary and throughout] There are several typographical errors: 'each functions' in the Summary, 'funtions' in §12 heading, and 'Kullback-Liebler' used throughout instead of 'Kullback-Leibler'.
- [§4 and elsewhere] The notation 'P~~Y' appears garbled in the text (e.g., in the description of the predictive distribution). Please use a consistent notation such as P_{Z|Y} or P^Y_Z.
- [References] Several key references (Dawid 1994, Dawid & Sebastiani 1996, Johnson 1991) are research reports or theses. If this is intended as a journal submission, please update them to published versions where available.
- [Figures] The influence diagrams (Figures 1-4) are referenced but not fully reproduced in the text shown. If the final version includes them, ensure the labels are legible and the notation matches the text.
- [§12] The dependence-function section is explicitly incomplete, which is acknowledged. However, the Summary's 'each function essentially determines the others' claim should be qualified to exclude dependence functions until a full characterization is available.
Circularity Check
No significant circularity: the uncertainty/scoring/discrepancy/dependence constructions are explicit and self-contained; the nonzero score reflects only the paper's acknowledged presentability restriction and reliance on some prior self-citations, neither of which is load-bearing.
full rationale
The paper's derivation chain is not circular. Forward constructions are explicit: from any loss L, H(P)=inf_a E_P L(Z,a), S(z,Q)=L(z,a(Q)), D(P,Q)=S(P,Q)-S(P,P), and C(P^{Z,U})=E[H(P^Z|U)]-H(P^Z). The reverse characterizations are proven by explicit construction rather than by assuming the conclusion. In Section 10, concavity plus the supporting-hyperplane property gives an affine φ_Q(P), and presentability is used to define S(z,Q)=f_Q(z), after which H(P)=S(P,P) follows from equality at P=Q. In Section 11, D(P,Q)-D(P,Q0) affine in P is used, via presentability, to define S(z;Q)=f(z;Q,Q0), yielding S(P,Q)-S(P,P)=D(P,Q) and hence a proper scoring rule whose discrepancy is D. These are mathematical equivalences, not fitted inputs or renamed conclusions. Self-citations to Dawid (1986), Dawid (1994), Dawid and Sebastiani (1996), and Sebastiani (1995) support standard or secondary facts, while the load-bearing characterizations are derived in this text. The paper explicitly acknowledges that presentability can fail (Hendrickson and Buehler 1971, Example 4.1), and the characterizations are conditional on that assumption; this is a scope limitation, not circularity. No empirical prediction is made from fitted parameters, and no target result is smuggled into the assumptions. Accordingly, no circular step is identified; the score of 2 reflects the conditional status of the general characterizations and the presence of minor self-citations, not actual circularity.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Bayesian decision theory: rational preferences imply a loss function L and a distribution P, and acts are ranked by expected loss L(P,a)=E_P[L(Z,a)].
- domain assumption The infimum defining H(P) in (3.3) is attained for every P in the domain, so a Bayes act a(P) exists.
- ad hoc to paper Every affine function on the domain P is presentable: F(P)=E_P[f(Z)] for some f.
- domain assumption In the design problem, Z is conditionally independent of (ξ, Y, a) given Θ, and the prior Π on Θ does not depend on ξ.
Cite this review
Pith. "Pith review of Coherent Measures of Discrepancy, Uncertainty and Dependence, with Applications to Bayesian Predictive Experimental Design." pith.science (2026). https://pith.science/paper/EZC6MU6U
@misc{pith2026260726077,
author = {Pith},
title = {Pith review of: Coherent Measures of Discrepancy, Uncertainty and Dependence, with Applications to Bayesian Predictive Experimental Design},
year = {2026},
howpublished = {\url{https://pith.science/paper/EZC6MU6U}},
note = {Machine review of arXiv:2607.26077}
}
read the original abstract
We show how, associated with any decision problem, we may derive related functions measuring uncertainty, discrepancy and dependence of distributions. Such "coherent" functions have special properties, which we characterise, and each function essentially determines the others. The theory is applied to the Bayesian formulation of the problem of choosing an experiment in order to make a subsequent prediction. It is shown that coherent choice criteria may be based on any of the coherent functions, with related functions yielding identical solutions.
Figures
Reference graph
Works this paper leans on
-
[1]
Aitchison, J. (1990). On coherence in parametric density estimation. Biometrika 77, 905-908. Amari, S. (1985). Differential-Geometrical Methods in Statistics. Springer Lec ture Notes in Statistics 28, Springer. Barndorff-Nielsen,
1990
-
[2]
Chichester: Wiley Brooks, R
(1978) Information and Exponential Families in Statistical Theory. Chichester: Wiley Brooks, R. J. (1972). A decision theory approach to optimal regression designs. Biometrika 59, 563-571. Dawid, A. P. (1986). Probability forecasting. Encyclopedia of Statistical Sciences, vol. 7, edited by S. Kotz, N. I. Johnson and C. B. Read. Wiley-Interscience, 210-218...
1978
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.