{"id":"d2962a2f-098a-428c-a22c-bbaf8d796d0c","arxiv_id":"2607.26077","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A decision problem induces equivalent coherent measures of uncertainty, scoring, discrepancy, and dependence; design criteria based on any of them coincide.","lead":"This paper shows that any standard decision problem -- choosing an action to minimize expected loss -- automatically yields a coherent measure of uncertainty, a scoring rule for probability forecasts, a distance between distributions, and a measure of dependence. Because all of these arise from the same loss function, experimental designs chosen using any of them are identical.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Presentability assumption leaves the general characterizations conditional: §10's concavity-implies-coherence and §11's discrepancy iff require every affine function on the domain to be E_P[f(Z)]; the paper concedes this can fail, so the advertised 'each function essentially determines the others' i","rationale":"The core of the paper—deriving S, H, D, and C from a terminal loss and showing the associated Bayesian design criteria coincide—is sound and is verified in the examples. The soft spot is the converse characterizations in §§10-11, which require both the supporting-hyperplane property and presentability of every affine function. The reader identified presentability as the weakest assumption, and I agree. This is a genuine limitation, but it is explicitly flagged by the paper and does not undermine the central applications: each worked example either satisfies Lemma 9.1 because point masses are in the domain, or verifies the required presentable scoring rule directly. The abstract's phrase 'each function essentially determines the others' overstates the general theorems, especially for dependence functions, but this is a scope caveat rather than an internal inconsistency. Hence I keep the ACCEPT verdict unchanged while emphasizing that the general characterizations are conditional on unstated regularity conditions.","tokens_in":20511,"tokens_out":13739,"duration_ms":161341,"concrete_test":"Take the Hendrickson-Buehler (1971) Example 4.1 counterexample: let P be the convex family of distributions for which they exhibit a concave H that is not coherent. For each Q in P, check explicitly whether there exists a function f_Q: Z -> R such that H(P) <= E_P[f_Q(Z)] for all P in P, with equality at P=Q. If no such presentable supporting affine function exists at some Q, the §10 converse fails exactly at the presentability assumption. Then state the precise topological/continuity conditions under which every concave H has a presentable supporting hyperplane; if the conditions exclude the Hendrickson-Buehler example, the theorem must be amended to include those conditions before the general characterization is claimed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the converse direction of the characterizations. Section 9 defines presentable affine functions and then 'simply assume[s]' every affine function on the domain is presentable. Section 10 uses that assumption to turn a supporting hyperplane of a concave H into a pointwise scoring rule S(z,Q) via (10.4)-(10.7), concluding that concavity essentially implies coherence. Section 11 similarly uses presentability to construct S(z,Q) from the affine difference D(P,Q)-D(P,Q0) in (11.2)-(11.3), giving the iff characterization of coherent discrepancy functions. If presentability fails, concavity alone does not imply coherence: the supporting hyperplane may exist only as a non-presentable affine functional, so there is no pointwise scoring rule. The paper itself cites Hendrickson & Buehler (1971), Example 4.1, as a case where this converse fails, and it never states precisely the 'suitable additional technical conditions' promised in §10. For the worked examples (1-8), presentability is either immediate because the domain contains point masses (Lemma 9.1) or is checked directly, so the design-coincidence results survive. But the abstract/summary claim that each function essentially determines the others is stronger than the theorems as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a decision-theoretic framework in which a terminal decision problem with loss L induces a coherent uncertainty function H(P)=inf_a E_P L(Z,a), a proper scoring rule S(z,Q)=L(z,a(Q)), a discrepancy D(P,Q)=S(P,Q)-H(P), and a dependence function C(P^{Z,U})=H(P^Z)-E H(P^{Z|U}). It shows that all of these yield equivalent formulations of Bayesian predictive experimental design. It then characterizes coherent uncertainty functions as concave functions satisfying a supporting-hyperplane/presentability condition, and coherent discrepancy functions as those for which D(P,Q)-D(P,Q0) is affine and presentable in P. It proves that coherent discrepancies are A-coherent, refuting Aitchison's conjecture that only the Kullback-Leibler distance has that property. Eight worked examples link the theory to variance, entropy, characteristic-function scores, quadratic (Brier) scores, exponential families, and D-optimality.","tokens_in":20816,"tokens_out":4224,"duration_ms":43849,"significance":"If correct, the paper gives a clean, unifying decision-theoretic account of many existing criteria, showing that proper scoring rules, entropy/divergences, mutual information, and Bayesian design criteria are facets of one structure. The characterization theorems, while building on prior work (Dawid 1986, 1994; Dawid & Sebastiani 1996; Hendrickson & Buehler 1971), are proved from the definitions and are illustrated by explicit verification in every example. The refutation of Aitchison's conjecture and the link to D-optimality are valuable. The manuscript is notably candid about the presentability condition and about the lack of a full dependence-function characterization, which is a strength.","major_comments":[{"comment":"The statement that 'coherence is essentially equivalent to concavity' is not made precise. The converse proof uses the supporting-hyperplane property (10.1) and presentability of the affine function φ_Q; the paper cites Hendrickson & Buehler (1971), Theorem 4.1, and Johnson (1991), but does not state the 'suitable additional technical conditions' promised in the Introduction. Since the abstract and summary treat this equivalence as a central result, please state those conditions explicitly, or clearly label the general equivalence as heuristic and rely on the direct verifications in the examples. This is a clarity issue rather than a mathematical error, because each example is checked directly.","section":"§10 (pp. 27-28), esp. (10.1)-(10.7)"},{"comment":"The presentability assumption that 'every affine function on the domain P is presentable' is load-bearing for the converse characterization of coherent discrepancy functions and for constructing a scoring rule from a concave H in (10.4)-(10.7). The paper explicitly concedes that this assumption can fail (Hendrickson & Buehler 1971, Example 4.1). Consequently the Summary's claim that 'each function essentially determines the others' is stronger than the theorems as stated. I recommend adding an explicit caveat to the Summary/Abstract, and noting Lemma 9.1 as a sufficient condition. The design-coincidence results for the examples survive because presentability is either immediate (point masses) or checked directly.","section":"§9 and §11, esp. (11.2)-(11.3)"}],"minor_comments":[{"comment":"There are several typographical errors: 'each functions' in the Summary, 'funtions' in §12 heading, and 'Kullback-Liebler' used throughout instead of 'Kullback-Leibler'.","section":"Summary and throughout"},{"comment":"The notation 'P~~Y' appears garbled in the text (e.g., in the description of the predictive distribution). Please use a consistent notation such as P_{Z|Y} or P^Y_Z.","section":"§4 and elsewhere"},{"comment":"Several key references (Dawid 1994, Dawid & Sebastiani 1996, Johnson 1991) are research reports or theses. If this is intended as a journal submission, please update them to published versions where available.","section":"References"},{"comment":"The influence diagrams (Figures 1-4) are referenced but not fully reproduced in the text shown. If the final version includes them, ensure the labels are legible and the notation matches the text.","section":"Figures"},{"comment":"The dependence-function section is explicitly incomplete, which is acknowledged. However, the Summary's 'each function essentially determines the others' claim should be qualified to exclude dependence functions until a full characterization is available.","section":"§12"}],"recommendation":"minor_revision","confidential_remarks":"This is a strong theoretical paper, clearly within the scope of a mathematical statistics journal. I would be comfortable with acceptance after a short revision. The main reason I am not recommending immediate accept is that the abstract/summary overstate the generality of the characterisation results without explicitly flagging the presentability condition and the incompleteness of the dependence-function characterization. The author's caveats are present in the body but should be reflected in the summary so that readers do not take the claims as unconditional."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a real paper from 1998, and it deserves serious attention. The main contribution is a unification: starting from a single terminal decision problem, you get a proper scoring rule, a concave uncertainty function, a discrepancy function, and a dependence function, and the Bayesian design criteria based on any of these coincide. The new characterization theorems are the core: coherent discrepancy iff the affine-difference condition (11.1), coherent uncertainty essentially iff concavity plus a supporting-hyperplane/presentability condition, and coherent implies A-coherent, which knocks down Aitchison's conjecture that only KL has that property. Example 8, showing that log-det D-optimality is coherent, is a nice worked payoff.\n\nThe paper is honest about its main limitation. Section 9 introduces presentability and then just assumes every affine function on the domain is presentable; §§10 and 11 use that to turn supporting hyperplanes into pointwise scoring rules. The paper itself cites Hendrickson & Buehler (1971, Example 4.1) where concavity alone fails, so the characterizations are conditional, not unconditional. The stress-test note is right about this: the 'suitable additional technical conditions' promised in §10 are never pinned down. But the paper flags the assumption explicitly, and the worked examples (1–8) satisfy presentability either via point masses or by direct construction. So the design-coincidence results survive; only the abstract-level claim that each function essentially determines the others is a bit stronger than the theorems as stated. The dependence characterization is also explicitly incomplete, which again the abstract glosses over. Still, these are proportionality issues, not fatal holes.\n\nThe citation pattern is fine. The paper relies on Dawid's earlier work for some scaffolding, but the main theorems are proved from definitions, and the examples are verified explicitly. There is no fitting to data here; this is pure theory. I would not dock it for self-citation when the results are internally derived.\n\nWho should read it: anyone working on proper scoring rules, Bayesian experimental design, or coherence. It is valuable to see the old theory in one place. Send it to a serious referee; it deserves careful review. My own verdict: accept, with minor revision to tone down the abstract and to state the presentability condition as a hypothesis rather than an assumption.","headline":"A 1998 unification of scoring rules, concave uncertainty, discrepancies, and expected value of sample information—strong on characterizations, conditional on a presentability assumption that the paper openly flags.","tokens_in":21298,"tokens_out":2829,"would_cite":true,"duration_ms":27678,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62C10","62K05","62B10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that any Bayesian decision problem generates four interchangeable coherent functions—uncertainty, scoring rule, discrepancy, and dependence—and that Bayesian predictive experimental design can be based on any of them with i","keywords":["coherent uncertainty","proper scoring rule","Bayesian experimental design","discrepancy function","dependence function","expected value of sample information","Kullback-Leibler distance","concavity"],"falsifier":"Find a concave function H on a convex family of distributions that has no presentable supporting hyperplane—equivalently, no proper scoring rule S with H(P)=S(P,P). If such an H exists, the claim that coherence is essentially concavity is false without extra conditions; the paper itself points to such a counterexample, so the meaningful falsifier is a concrete non-presentable concave H or a discrepancy D with affine differences but no scoring-rule realization.","tokens_in":20357,"feed_emoji":"🎯","tokens_out":5268,"duration_ms":47326,"temperature":0.7,"pith_summary":"This paper shows that any Bayesian terminal decision problem—a loss function, an action space, and a distribution for an uncertain outcome—generates a matched set of four 'coherent' functions: an uncertainty function (the Bayes loss), a proper scoring rule, a discrepancy function, and a dependence function. Each can be used to express the same design problem, and they all yield identical choices of experiment. The paper then characterises when a function given independently is coherent: an uncertainty function is coherent precisely when it is concave (under a technical presentability condition), and a discrepancy is coherent exactly when D(P,Q) - D(P,Q0) is affine in P for fixed Q and Q0. This refutes the earlier conjecture that only the Kullback-Leibler distance can serve as a coherent discrepancy. A sympathetic reader should care because the result unifies many classical criteria—entropy, mutual information, D-optimality, quadratic scoring—under one decision-theoretic umbrella.","feed_headline":"Four coherent measures yield the same experiment","feed_subtitle":"Bayesian design can minimize uncertainty, discrepancy, or maximize dependence—each choice is equivalent.","key_machinery":"The central mechanism is the chain of identities linking a terminal decision problem to its derived functions. The key identity is D(P,Q) = S(P,Q) - S(P,P), which converts any proper scoring rule S into a discrepancy, and conversely any coherent discrepancy arises this way from the scoring rule built from Bayes actions. For uncertainty, the load-bearing condition is the supporting-hyperplane identity H(P) ≤ E_P[f_Q(Z)] with equality at P=Q and f_Q presentable; this turns a concave H into a proper scoring rule S(z,Q) = f_Q(z). For discrepancies, the characterizing identity is the affineness in P of D(P,Q) - D(P,Q0), which allows construction of S via presentability.","core_discovery":"The central claim is that coherence is a single decision-theoretic property wearing four hats. Starting from a loss L, define H(P) = inf_a E_P L(Z,a); set S(z,Q) = L(z,a(Q)) with a(Q) a Bayes act; set D(P,Q) = S(P,Q) - S(P,P); and set C as the expected reduction in Bayes loss from observing an auxiliary variable. The paper proves that each of these is coherent, that H is concave, that D(P,Q) - D(P,Q0) is affine in P, and—under the presentability assumption—conversely any concave H and any discrepancy with affine differences can be realized from some decision problem. It follows that design criteria based on any of the four coincide: choose the experiment minimizing expected posterior uncerta","pith_inferences":["This equivalence suggests a practical recipe: to check whether any proposed design criterion is Bayesian, verify concavity (for uncertainty) or the affine-difference condition (for discrepancy) rather than constructing a full loss function.","Because the characterization depends on presentability, in infinite-dimensional or non-parametric settings one should expect counterexamples; a robust version might require additional topological hypotheses.","The concluding discussion points toward a differential-geometric reading: coherent discrepancies induce metrics on distribution space; if only coherent discrepancies are considered, the Fisher information metric may be characterizable as the unique coherent choice, giving Bayesian foundations for information geometry.","For applied design, this means that criteria like expected entropy or expected quadratic loss are not competing philosophies but re-expressions of the same underlying expected-loss minimization, so empirical comparisons among them compare numerical approximations, not different objectives."],"forward_implications":["Any Bayesian predictive design problem can be solved equivalently by minimizing expected posterior uncertainty, minimizing expected discrepancy between sampling and predictive distributions, or maximizing dependence between predictand and observation; the same experiment is optimal under all three.","Every concave uncertainty function (with presentability) yields a proper scoring rule and hence a coherent design criterion, so concavity is the test for whether a proposed uncertainty criterion is Bayesian-coherent.","A discrepancy is coherent exactly when its differences in the second argument are affine in the first; quadratic and characteristic-function discrepancies pass this test, so coherent discrepancies are plentiful, not unique.","The expected value of sample information is a coherent dependence function, and its maximization is equivalent to uncertainty-minimizing design.","The class of A-coherent discrepancies is broader than the Kullback-Leibler distance; the conjecture of uniqueness is false."],"fun_headline_variants":["Four coherent measures, one optimal experiment","Uncertainty, discrepancy, dependence: same design choice","Coherent functions unify Bayesian design criteria","One decision problem, four equivalent measures","Bayesian design: any coherent measure works"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that every affine function on a family of distributions can be written as the expectation of some function of the outcome (presentability); the paper notes this can fail in general, and without it concavity alone need not imply coherence.","fun_headline_variants_meta":{"raw":{"variants":["Four coherent measures, one optimal experiment","Uncertainty, discrepancy, dependence: same design choice","Coherent functions unify Bayesian design criteria","One decision problem, four equivalent measures","Bayesian design: any coherent measure works"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000542,"raw_usage":{"total_tokens":2369,"prompt_tokens":616,"completion_tokens":1753,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":360,"completion_tokens_details":{"reasoning_tokens":1686}},"tokens_in":360,"tokens_out":1753,"duration_ms":11756,"temperature":1.0,"reasoning_tokens":1686,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:27:18.376432+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a concave function H on a convex family of distributions that has no presentable supporting hyperplane—equivalently, no proper scoring rule S with H(P)=S(P,P). If such an H exists, the claim that coherence is essentially concavity is false without extra conditions; the paper itself points to such a counterexample, so the meaningful falsifier is a concrete non-presentable concave H or a discrepancy D with affine differences but no scoring-rule realization.","supporting_citations":[],"review_version":1}