REVIEW 2 major objections 4 minor 1 cited by
Convex Cost of Information via Statistical Divergence
T0 review · 2 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A cost of information is Blackwell monotone, mixture convex, sub-additive, and identity additive exactly when it is the maximum over extended Rényi divergences between state-dependent signal distributions.
desk verdict Genuine extension of PST (2023) with clean axioms; the Eq. (9) complaint is a misreading, and the paper deserves refereeing despite some technical density. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the extended Rényi divergence D_{α,β}(µ), a multi-state generalization of Rényi's divergence: for a probability weighting α over states (normalized, not concentrated on one state), it is a log of the integral over signals of the product of state likelihoods raised to α; the boundary α = e_i is the weighted sum of Kullback–Leibler divergences from state i. These divergences are additive under independent bundling and monotone with respect to Blackwell dominance. The cost function takes the maximum, over a compact set of probability measures m on parameter pairs (α,β), of the expected divergence; the max operator is what produces strict convexity and strict sub-addit
What would settle it
Find finite experiments µ and ν such that D_{α,β}(µ) ≥ D_{α,β}(ν) for every divergence parameter, but some cost function satisfying Blackwell monotonicity, mixture convexity, sub-additivity, and identity additivity assigns C(µ) < C(ν). Lemma B.1 says no such pair can exist, so exhibiting one would falsify the representation.
Extended reading notes
Core claim
The central result (Theorem 1) is a characterization: a cost function defined on bounded experiments is Blackwell monotone, mixture convex, sub-additive, and identity additive if and only if there is a compact set M of measures on divergence parameters such that for every experiment µ, C(µ) = max_{m∈M} ∫ D_{α,β}(µ) dm(α,β). Here D_{α,β} collects the multi-state Rényi divergence D_α(µ) = (1/(α_max−1)) log ∫ ∏_{i∈Θ} (dµ_i/dλ)^{α_i} dλ for non-unit weight vectors α in the simplex, together with the weighted KL divergences reached as limits. The theorem says the four axioms are exactly the structure 'take the most expensive divergence from a menu': strict mixture convexity arises either because
Load-bearing premise
The proof leans on an external large-sample theorem: if one experiment's entire profile of extended Rényi divergences strictly exceeds another's, then some tensor power of the first Blackwell-dominates the same power of the second; if that dominance step fails, the link between divergence rankings and cost rankings collapses, and with it the representation.
Editorial extensions
If this is right
- KL costs are exactly the additive cases: the menu M collapses to a singleton, so bundling different experiments gives no discount (Proposition 1).
- Mixture linearity plus the four axioms also forces the KL form, so any genuinely convex cost must use either several divergences in the menu or non-KL Rényi parameters.
- The two tractable special cases are characterized by weaker axioms: Max-KL costs equal the class satisfying dilution linearity, and Rényi costs equal the class satisfying independence; both are disciplined ways to depart from posterior separability.
- In the two-state, three-action decision problem, a symmetric Rényi cost makes using all three actions optimal on an interval of payoff values, while posterior-separable and symmetric Max-KL costs make it optimal at a single knife-edge value at most—matching the frequency of three-action policies in the lab data used by the paper.
Reading between the lines
- I read the menu M as a flexible, nonparametric object: any convex cost family can be fit by expanding M, so the characterization turns the qualitative debate over convexity into a quantitative question of estimating M from stochastic-choice data.
- The Rényi special case is differentiable in signal probabilities while max-form costs are not, which likely makes it the more convenient member of this family for structural estimation and for models that need interior solutions; the paper notes the differentiability advantage but does not pursue estimation.
- The appendix's observation that Chernoff information and worst-case privacy leakage are instances of the max form suggests the same four axioms may underlie information measures used in learning and privacy, a connection the paper records but leaves for later work.
- A direct laboratory test would vary the payoff parameters in the three-action problem and measure the empirical frequency of three-action policies: the Rényi cost predicts a positive-measure region, Max-KL and posterior-separable costs predict a knife-edge.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper characterizes convex information costs on bounded Blackwell experiments. The main result (Theorem 1) states that a cost function is Blackwell monotone, mixture convex, sub-additive, and identity additive if and only if it is a Max-Rényi cost: C(µ) = max_{m∈M} ∫ D_{α,β}(µ) dm(α,β), where D_{α,β} are extended Rényi divergences and M is a compact set of measures. Theorem 2 characterizes two special cases: Max-KL costs (via dilution linearity) and Rényi costs (via an independence axiom). The proof strategy is to first obtain a more general max-of-integrals representation over generalized divergences (Theorem B.1), using the large-sample Blackwell dominance theorem of Farooq, Fritz, Haapasalo, and Tomamichel (2024), and then to impose mixture convexity to restrict the support of the representing measures. An application to a three-action decision problem shows that symmetric Rényi costs can rationalize the use of all three actions for an interval of parameters, in line with experiments by Dean and Neligh (2023).
Significance. If the proof is repaired, the paper makes a substantial contribution: it provides the first axiomatic characterization of convex information costs as maxima of Rényi divergences, strictly generalizing the additive KL costs of Pomatto, Strack, and Tamuz (2023). The representation is derived from stated axioms rather than assumed, and no parameters are fitted to data in the application; the connection to posterior-separable costs and the falsifiable prediction about support size are valuable. The appendix contains genuine proof machinery, including a monotone homogeneous subadditive functional representation and a clean use of the Aliprantis-Border theorem. However, one load-bearing formula in the proof of Theorem 1 is incorrect as printed, and the manuscript's reliance on the external matrix-majorization theorem should be made fully explicit before the result can be considered established.
major comments (2)
- [Appendix B.2, Eq. (9)] The formula for D_{γ,ψ}(ν_k) is incorrect. For ν_k := (1/k)μ^{⊗k}, the Hellinger integral is H_α(ν_k) = (k−1)/k + (1/k)H_α(μ)^k, so D_{γ,ψ}(ν_k) = 1/(γ−1) log((k−1)/k + (1/k)H_α(μ)^k). The printed equation places the exponent k on the entire bracket, i.e. log(((k−1)/k + H/k)^k). This changes the limits used later: for γ>1 with H>1, the printed expression tends to (H−1)/(γ−1), not ∞, so the argument excluding measures supported on γ>1 in the proof of Theorem 1 fails as written; for γ<1 with H<1, the printed expression tends to (H−1)/(γ−1)>0, not 0, so the dilution-linearity argument in Appendix B.3.3 also fails. With the corrected formula the desired ∞ and 0 limits are restored, so the error is local, but it is load-bearing and must be fixed.
- [Appendix B.1.1, Lemma B.1] The proof invokes Theorem 19 of Farooq, Fritz, Haapasalo, and Tomamichel (2024) for the full divergence domain [1/|Θ|,∞]×Ψ, including γ>1 and γ=∞. This large-sample Blackwell-dominance result is the only bridge from divergence dominance to C-monotonicity, so the hypotheses and domain of the published theorem must be stated precisely and checked. If the theorem is restricted to γ∈[1/2,1] or to finite α, then Lemma B.1 and Theorem B.1 collapse. Please quote the theorem and confirm explicitly that it applies to the entire domain used in the intermediate representation.
minor comments (4)
- [Lemma B.2] Property (ii) is misprinted: it says F(αx)=αF(αx), which is vacuous; it should read F(αx)=αF(x).
- [Appendix B.4.1] The phrase 'Dilation linearity' should be 'Dilution linearity' to match the terminology used elsewhere.
- [Section 2] The parameter t is introduced for Rényi divergence in [1/2,1), but later subsections use γ for the same role; the notation is manageable but could be unified for readability.
- [Proof of Theorem 2, second part] The envelope-theorem argument with the second derivative in equation (12) is terse; adding a short justification of why m_ε can be selected so that the derivative formulas hold would improve readability.
Circularity Check
No significant circularity: the Max-Renyi representation is derived from explicitly stated axioms, with no fitted parameters or self-citational load-bearing premise.
full rationale
The paper's central Theorem 1 characterizes Max-Renyi cost functions from four stated axioms—Blackwell monotonicity, mixture convexity, sub-additivity, and identity additivity—using a representation theorem (Theorem B.1) and an external matrix-majorization result. The functional form is not assumed in the primitives: the proof first establishes a more general divergence representation and then uses the axioms to restrict the support of the representing measures. No parameter is fitted to data; in the application, parameters are chosen only to exhibit an open region where a three-action optimum exists, and the target behavior is not used to set the model. The citations to Pomatto, Strack, and Tamuz (2023), Mu, Pomatto, Strack, and Tamuz (2021), and Farooq, Fritz, Haapasalo, and Tomamichel (2024) are prior independent mathematical results or benchmark characterizations, not self-citations by the present authors; the only related self-citation (Frick, Iijima, and Ishii) appears in a non-load-bearing literature discussion. The skeptic's objection concerning Eq. (9) is a potential internal computational error in the proof of the gamma>1 exclusion step, but it is an issue of correctness, not circularity: the conclusion is not equivalent to the premise by construction, nor does the proof rely on assuming the representation. The paper is therefore self-contained against external benchmarks and receives score 0.
Assumptions & free parameters
free parameters (4)
- lambda (Renyi cost scale)
- t (Renyi divergence order in Claim 1(1)) =
any t in (0,1); t = 1/2 for the paper's symmetric Renyi cost
- v, w (matching and safe-action payoffs)
- compact set M of measures (Theorem 1 representation)
assumptions (8)
- domain assumption Mixture convexity: C(alpha mu + (1-alpha) nu) <= alpha C(mu) + (1-alpha) C(nu) for all experiments and alpha in (0,1)
- domain assumption Sub-additivity: C(mu tensor nu) <= C(mu) + C(nu) for bundled independent experiments
- domain assumption Identity additivity: C(mu^(tensor k)) = k C(mu) for every k
- domain assumption Blackwell monotonicity: C(mu) >= C(nu) whenever mu Blackwell-dominates nu
- domain assumption Domain restriction to bounded experiments (log likelihood ratios bounded a.s.)
- standard math Theorem 19 of Farooq, Fritz, Haapasalo, Tomamichel (2024): profile-wise divergence dominance implies Blackwell dominance of some tensor power
- standard math Representation of monotone, positively homogeneous, sub-additive functionals as suprema of integrals (Aliprantis-Border Thm 7.51) and Banach-Alaoglu compactness
- standard math Theorem 1 of Mu, Pomatto, Strack, Tamuz (2021) on binary-state Blackwell dominance in large samples
Cite this review
Pith. "Pith review of Convex Cost of Information via Statistical Divergence." pith.science (2026). https://pith.science/paper/P4X34JTN
@misc{pith2026250900229,
author = {Pith},
title = {Pith review of: Convex Cost of Information via Statistical Divergence},
year = {2026},
howpublished = {\url{https://pith.science/paper/P4X34JTN}},
note = {Machine review of arXiv:2509.00229}
}
read the original abstract
This paper characterizes convex information costs using an axiomatic approach. We employ mixture convexity and sub-additivity, which capture the idea that producing "balanced" outputs is less costly than producing ``extreme'' ones. Our analysis leads to a novel class of cost functions that can be expressed in terms of R\'enyi divergences between signal distributions across states. This representation allows for deviations from the standard posterior-separable cost, thereby accommodating recent experimental evidence. We also characterize two simpler special cases, which can be written as either the maximum or a convex transformation of posterior-separable costs.
Forward citations
Cited by 1 Pith paper
-
A mathematical study of the excess growth rate
The excess growth rate is the unique functional, up to a constant, satisfying each of three axiom systems; its deterministic maximizer invests only in the best- and worst-performing assets.
Reference graph
Works this paper leans on
-
[1]
Economies and Diseconomies of Scale in the Cost of Information,
Aliprantis, C. D., and K. C. Border (2006): Infinite Dimensional Analysis A Hitchhiker’s Guide . Springer. Baker, C. (2023): “Economies and Diseconomies of Scale in the Cost of Information,” working paper. Bloedel, A. W., T. Denti, and L. Pomatto (2025): “Understanding information acquisition through f-informativity and duality,,” . Bloedel, A. W., and W....
work page 2006
-
[36]
Some properties of Matusita’s measure of affinity of several distributions,
Cambridge University Press. Toussaint, G. T. (1974): “Some properties of Matusita’s measure of affinity of several distributions,” Annals of the Institute of Statistical Mathematics , 26(1), 389–394. Tsallis, C. (1988): “Possible generalization of Boltzmann-Gibbs statistics,” Journal of statistical physics, 52, 479–487. V an Erven, T., and P. Harremo ¨es ...
work page 1974
-
[690]
Rational inattention dynamics: Inertia and delay in decision-making,
Steiner, J., C. Stewart, and F. Mat ˇejka (2017): “Rational inattention dynamics: Inertia and delay in decision-making,” Econometrica, 85(2), 521–553. Torgersen, E. (1991): Comparison of statistical experiments , vol
work page 2017
-
[3259]
Experimental cost of information,
Denti, T., M. Marinacci, and A. Rustichini (2022): “Experimental cost of information,” American Economic Review, 112(9), 3106–3123. Dwork, C. (2006): “Differential privacy,” in International colloquium on automata, languages, and programming, pp. 1–12. Springer. Ellis, A. (2018): “Foundations for optimal inattention,” Journal of Economic Theory , 173, 56–...
work page 2022
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.