Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

Convex Cost of Information via Statistical Divergence

T0 review · 2 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A cost of information is Blackwell monotone, mixture convex, sub-additive, and identity additive exactly when it is the maximum over extended Rényi divergences between state-dependent signal distributions.

desk verdict Genuine extension of PST (2023) with clean axioms; the Eq. (9) complaint is a misreading, and the paper deserves refereeing despite some technical density. read the letter →

arxiv 2509.00229 v1 pith:P4X34JTN submitted 2025-08-29 econ.TH

classification econ.TH MSC 91B0691B4462B1094A17
keywords informationcostRényidivergenceBlackwellmonotonicitymixtureconvexitysub-additivityposteriorseparabilityrationalinattentionexperiments
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper pins down what 'convex' means for the cost of acquiring information. It shows that four natural axioms—monotonicity with respect to Blackwell's informativeness order, convexity under probabilistic mixtures, sub-additivity under independent bundling, and exact additivity for identical repeated draws—are satisfied by exactly one functional family: the maximum, over a menu of measures, of extended Rényi divergences between the signal distributions induced by different states. This generalizes the standard KL cost of information, which is the special case where the menu is a single measure. The extension matters because additive posterior-separable costs predict that an optimal decision maker never uses more actions than there are states, whereas laboratory subjects repeatedly use more actions; the paper shows that a simple Rényi cost reproduces that behavior.

What carries the argument

The load-bearing object is the extended Rényi divergence D_{α,β}(µ), a multi-state generalization of Rényi's divergence: for a probability weighting α over states (normalized, not concentrated on one state), it is a log of the integral over signals of the product of state likelihoods raised to α; the boundary α = e_i is the weighted sum of Kullback–Leibler divergences from state i. These divergences are additive under independent bundling and monotone with respect to Blackwell dominance. The cost function takes the maximum, over a compact set of probability measures m on parameter pairs (α,β), of the expected divergence; the max operator is what produces strict convexity and strict sub-addit

What would settle it

Find finite experiments µ and ν such that D_{α,β}(µ) ≥ D_{α,β}(ν) for every divergence parameter, but some cost function satisfying Blackwell monotonicity, mixture convexity, sub-additivity, and identity additivity assigns C(µ) < C(ν). Lemma B.1 says no such pair can exist, so exhibiting one would falsify the representation.

Watch

Extended reading notes

Core claim

The central result (Theorem 1) is a characterization: a cost function defined on bounded experiments is Blackwell monotone, mixture convex, sub-additive, and identity additive if and only if there is a compact set M of measures on divergence parameters such that for every experiment µ, C(µ) = max_{m∈M} ∫ D_{α,β}(µ) dm(α,β). Here D_{α,β} collects the multi-state Rényi divergence D_α(µ) = (1/(α_max−1)) log ∫ ∏_{i∈Θ} (dµ_i/dλ)^{α_i} dλ for non-unit weight vectors α in the simplex, together with the weighted KL divergences reached as limits. The theorem says the four axioms are exactly the structure 'take the most expensive divergence from a menu': strict mixture convexity arises either because

Load-bearing premise

The proof leans on an external large-sample theorem: if one experiment's entire profile of extended Rényi divergences strictly exceeds another's, then some tensor power of the first Blackwell-dominates the same power of the second; if that dominance step fails, the link between divergence rankings and cost rankings collapses, and with it the representation.

Editorial extensions

If this is right

  • KL costs are exactly the additive cases: the menu M collapses to a singleton, so bundling different experiments gives no discount (Proposition 1).
  • Mixture linearity plus the four axioms also forces the KL form, so any genuinely convex cost must use either several divergences in the menu or non-KL Rényi parameters.
  • The two tractable special cases are characterized by weaker axioms: Max-KL costs equal the class satisfying dilution linearity, and Rényi costs equal the class satisfying independence; both are disciplined ways to depart from posterior separability.
  • In the two-state, three-action decision problem, a symmetric Rényi cost makes using all three actions optimal on an interval of payoff values, while posterior-separable and symmetric Max-KL costs make it optimal at a single knife-edge value at most—matching the frequency of three-action policies in the lab data used by the paper.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I read the menu M as a flexible, nonparametric object: any convex cost family can be fit by expanding M, so the characterization turns the qualitative debate over convexity into a quantitative question of estimating M from stochastic-choice data.
  • The Rényi special case is differentiable in signal probabilities while max-form costs are not, which likely makes it the more convenient member of this family for structural estimation and for models that need interior solutions; the paper notes the differentiability advantage but does not pursue estimation.
  • The appendix's observation that Chernoff information and worst-case privacy leakage are instances of the max form suggests the same four axioms may underlie information measures used in learning and privacy, a connection the paper records but leaves for later work.
  • A direct laboratory test would vary the payoff parameters in the three-action problem and measure the empirical frequency of three-action policies: the Rényi cost predicts a positive-measure region, Max-KL and posterior-separable costs predict a knife-edge.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper characterizes convex information costs on bounded Blackwell experiments. The main result (Theorem 1) states that a cost function is Blackwell monotone, mixture convex, sub-additive, and identity additive if and only if it is a Max-Rényi cost: C(µ) = max_{m∈M} ∫ D_{α,β}(µ) dm(α,β), where D_{α,β} are extended Rényi divergences and M is a compact set of measures. Theorem 2 characterizes two special cases: Max-KL costs (via dilution linearity) and Rényi costs (via an independence axiom). The proof strategy is to first obtain a more general max-of-integrals representation over generalized divergences (Theorem B.1), using the large-sample Blackwell dominance theorem of Farooq, Fritz, Haapasalo, and Tomamichel (2024), and then to impose mixture convexity to restrict the support of the representing measures. An application to a three-action decision problem shows that symmetric Rényi costs can rationalize the use of all three actions for an interval of parameters, in line with experiments by Dean and Neligh (2023).

Significance. If the proof is repaired, the paper makes a substantial contribution: it provides the first axiomatic characterization of convex information costs as maxima of Rényi divergences, strictly generalizing the additive KL costs of Pomatto, Strack, and Tamuz (2023). The representation is derived from stated axioms rather than assumed, and no parameters are fitted to data in the application; the connection to posterior-separable costs and the falsifiable prediction about support size are valuable. The appendix contains genuine proof machinery, including a monotone homogeneous subadditive functional representation and a clean use of the Aliprantis-Border theorem. However, one load-bearing formula in the proof of Theorem 1 is incorrect as printed, and the manuscript's reliance on the external matrix-majorization theorem should be made fully explicit before the result can be considered established.

major comments (2)
  1. [Appendix B.2, Eq. (9)] The formula for D_{γ,ψ}(ν_k) is incorrect. For ν_k := (1/k)μ^{⊗k}, the Hellinger integral is H_α(ν_k) = (k−1)/k + (1/k)H_α(μ)^k, so D_{γ,ψ}(ν_k) = 1/(γ−1) log((k−1)/k + (1/k)H_α(μ)^k). The printed equation places the exponent k on the entire bracket, i.e. log(((k−1)/k + H/k)^k). This changes the limits used later: for γ>1 with H>1, the printed expression tends to (H−1)/(γ−1), not ∞, so the argument excluding measures supported on γ>1 in the proof of Theorem 1 fails as written; for γ<1 with H<1, the printed expression tends to (H−1)/(γ−1)>0, not 0, so the dilution-linearity argument in Appendix B.3.3 also fails. With the corrected formula the desired ∞ and 0 limits are restored, so the error is local, but it is load-bearing and must be fixed.
  2. [Appendix B.1.1, Lemma B.1] The proof invokes Theorem 19 of Farooq, Fritz, Haapasalo, and Tomamichel (2024) for the full divergence domain [1/|Θ|,∞]×Ψ, including γ>1 and γ=∞. This large-sample Blackwell-dominance result is the only bridge from divergence dominance to C-monotonicity, so the hypotheses and domain of the published theorem must be stated precisely and checked. If the theorem is restricted to γ∈[1/2,1] or to finite α, then Lemma B.1 and Theorem B.1 collapse. Please quote the theorem and confirm explicitly that it applies to the entire domain used in the intermediate representation.
minor comments (4)
  1. [Lemma B.2] Property (ii) is misprinted: it says F(αx)=αF(αx), which is vacuous; it should read F(αx)=αF(x).
  2. [Appendix B.4.1] The phrase 'Dilation linearity' should be 'Dilution linearity' to match the terminology used elsewhere.
  3. [Section 2] The parameter t is introduced for Rényi divergence in [1/2,1), but later subsections use γ for the same role; the notation is manageable but could be unified for readability.
  4. [Proof of Theorem 2, second part] The envelope-theorem argument with the second derivative in equation (12) is terse; adding a short justification of why m_ε can be selected so that the derivative formulas hold would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Max-Renyi representation is derived from explicitly stated axioms, with no fitted parameters or self-citational load-bearing premise.

full rationale

The paper's central Theorem 1 characterizes Max-Renyi cost functions from four stated axioms—Blackwell monotonicity, mixture convexity, sub-additivity, and identity additivity—using a representation theorem (Theorem B.1) and an external matrix-majorization result. The functional form is not assumed in the primitives: the proof first establishes a more general divergence representation and then uses the axioms to restrict the support of the representing measures. No parameter is fitted to data; in the application, parameters are chosen only to exhibit an open region where a three-action optimum exists, and the target behavior is not used to set the model. The citations to Pomatto, Strack, and Tamuz (2023), Mu, Pomatto, Strack, and Tamuz (2021), and Farooq, Fritz, Haapasalo, and Tomamichel (2024) are prior independent mathematical results or benchmark characterizations, not self-citations by the present authors; the only related self-citation (Frick, Iijima, and Ishii) appears in a non-load-bearing literature discussion. The skeptic's objection concerning Eq. (9) is a potential internal computational error in the proof of the gamma>1 exclusion step, but it is an issue of correctness, not circularity: the conclusion is not equivalent to the premise by construction, nor does the proof rely on assuming the representation. The paper is therefore self-contained against external benchmarks and receives score 0.

Assumptions & free parameters 4 free parameters · 8 assumptions · 0 invented entities

The central theorems are pure iff characterizations and contain no fitted numbers: the axioms themselves are the input. The paper's own axioms (mixture convexity, sub-additivity, identity additivity, Blackwell monotonicity) are domain assumptions about the primitive cost function, not hidden assumptions stacked to force the conclusion; the proof of Theorem 1 is substantive. The technical load is carried by published external theorems from non-overlapping authors (matrix majorization, monotone functional representation, Banach-Alaoglu). Hand-chosen parameters (lambda, t, v, w) appear only in the illustrative application and do not enter the central claim. No invented entities of the graviton type: the extended Renyi divergence and the Max-Renyi class are mathematical constructions derived from axioms, and their behavioral consequences are derived rather than assumed.

free parameters (4)
  • lambda (Renyi cost scale)
    In Claim 1(1), C(mu) = lambda(R_t(mu0||mu1) + R_t(mu1||mu0)) with lambda > 0. Chosen by hand to exhibit a parameter region with three-action optimal policies; not estimated from the Dean-Neligh data.
  • t (Renyi divergence order in Claim 1(1)) = any t in (0,1); t = 1/2 for the paper's symmetric Renyi cost
    The proof of the three-action result works for every t in (0,1), so the value is not tuned to data; the paper's symmetric Renyi cost corresponds to t = 1/2.
  • v, w (matching and safe-action payoffs)
    Claim 1(1) needs v above a threshold and w in an open interval W(v); these are constructed payoff values for a possibility statement, not point estimates.
  • compact set M of measures (Theorem 1 representation)
    The representation is quantified over any compact set of Borel measures on divergence parameters; this is the theorem's structural flexibility (a set-valued free object), not a constant fitted to data.
assumptions (8)
  • domain assumption Mixture convexity: C(alpha mu + (1-alpha) nu) <= alpha C(mu) + (1-alpha) C(nu) for all experiments and alpha in (0,1)
    Main postulate capturing 'balanced outputs are cheaper than extreme ones'; its equality case (mixture linearity) yields the KL/posterior-separable benchmark (Proposition 1). Section 3.1.
  • domain assumption Sub-additivity: C(mu tensor nu) <= C(mu) + C(nu) for bundled independent experiments
    Bundling yields 'balanced' information and is never costlier than separate acquisition; strict cases of this axiom are captured by the maximum operator (|M| > 1) in the representation. Section 3.1.
  • domain assumption Identity additivity: C(mu^(tensor k)) = k C(mu) for every k
    Repeating the identical experiment gives no balancing benefit; this pins down the homogeneity of the functional representation and rules out entropy-type costs. Section 3.1.
  • domain assumption Blackwell monotonicity: C(mu) >= C(nu) whenever mu Blackwell-dominates nu
    Used to pass from divergence dominance to cost dominance via large-sample Blackwell dominance (Lemma B.1). Section 3.2.
  • domain assumption Domain restriction to bounded experiments (log likelihood ratios bounded a.s.)
    Technical restriction in Section 2; Lemma A.1 uses it to approximate any bounded experiment by finite-signal ones uniformly in the divergence parameters. The paper notes Max-Renyi costs extend uniquely to unbounded experiments (footnote 16).
  • standard math Theorem 19 of Farooq, Fritz, Haapasalo, Tomamichel (2024): profile-wise divergence dominance implies Blackwell dominance of some tensor power
    Load-bearing external result used in Lemma B.1 to get C(mu) >= C(nu) from D(mu) > D(nu); published in IEEE Transactions on Information Theory.
  • standard math Representation of monotone, positively homogeneous, sub-additive functionals as suprema of integrals (Aliprantis-Border Thm 7.51) and Banach-Alaoglu compactness
    Converts the constructed functional F on the cone of divergences into the max-over-measures representation in Theorem B.1.
  • standard math Theorem 1 of Mu, Pomatto, Strack, Tamuz (2021) on binary-state Blackwell dominance in large samples
    Used in the proof sketch of Proposition 2 for the pairwise-Blackwell version, replacing the many-state majorization theorem.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Convex Cost of Information via Statistical Divergence." pith.science (2026). https://pith.science/paper/P4X34JTN

@misc{pith2026250900229,
  author       = {Pith},
  title        = {Pith review of: Convex Cost of Information via Statistical Divergence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P4X34JTN}},
  note         = {Machine review of arXiv:2509.00229}
}
read the original abstract

This paper characterizes convex information costs using an axiomatic approach. We employ mixture convexity and sub-additivity, which capture the idea that producing "balanced" outputs is less costly than producing ``extreme'' ones. Our analysis leads to a novel class of cost functions that can be expressed in terms of R\'enyi divergences between signal distributions across states. This representation allows for deviations from the standard posterior-separable cost, thereby accommodating recent experimental evidence. We also characterize two simpler special cases, which can be written as either the maximum or a convex transformation of posterior-separable costs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A mathematical study of the excess growth rate

    cs.IT 2025-10 conditional novelty 7.0 of 10

    The excess growth rate is the unique functional, up to a constant, satisfying each of three axiom systems; its deterministic maximizer invests only in the best- and worst-performing assets.

Reference graph

Works this paper leans on

4 extracted references · 4 canonical work pages · cited by 1 Pith paper

  1. [1]

    Economies and Diseconomies of Scale in the Cost of Information,

    Aliprantis, C. D., and K. C. Border (2006): Infinite Dimensional Analysis A Hitchhiker’s Guide . Springer. Baker, C. (2023): “Economies and Diseconomies of Scale in the Cost of Information,” working paper. Bloedel, A. W., T. Denti, and L. Pomatto (2025): “Understanding information acquisition through f-informativity and duality,,” . Bloedel, A. W., and W....

  2. [36]

    Some properties of Matusita’s measure of affinity of several distributions,

    Cambridge University Press. Toussaint, G. T. (1974): “Some properties of Matusita’s measure of affinity of several distributions,” Annals of the Institute of Statistical Mathematics , 26(1), 389–394. Tsallis, C. (1988): “Possible generalization of Boltzmann-Gibbs statistics,” Journal of statistical physics, 52, 479–487. V an Erven, T., and P. Harremo ¨es ...

  3. [690]

    Rational inattention dynamics: Inertia and delay in decision-making,

    Steiner, J., C. Stewart, and F. Mat ˇejka (2017): “Rational inattention dynamics: Inertia and delay in decision-making,” Econometrica, 85(2), 521–553. Torgersen, E. (1991): Comparison of statistical experiments , vol

  4. [3259]

    Experimental cost of information,

    Denti, T., M. Marinacci, and A. Rustichini (2022): “Experimental cost of information,” American Economic Review, 112(9), 3106–3123. Dwork, C. (2006): “Differential privacy,” in International colloquium on automata, languages, and programming, pp. 1–12. Springer. Ellis, A. (2018): “Foundations for optimal inattention,” Journal of Economic Theory , 173, 56–...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.