{"id":"1a8ab51d-ab74-4d2a-a044-2914925fe80c","arxiv_id":"2501.10099","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Every common variant of alpha-mutual information can be rewritten as the multiplicative increase in the maximal generalized mean of an adversary's gain, for suitably chosen scoring rules.","lead":"This paper derives new mathematical formulas that express five types of alpha-mutual information as privacy leakage measures based on the gains of a guessing adversary. The results unify several existing information leakage metrics under one framework.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's new conditional entropies use unnormalized powers p_X^{1/α} and p_X^{α/(2α−1)}; on a noiseless channel this doubles I^S_α and I^LP_α, so the central representation as printed is false.","rationale":"The reader's weakest_assumption concerns the operational meaning of the gain functions in the privacy interpretation. My check found a more basic, correctness-level problem: Theorem 1, as written, is false because the conditional entropies in Eqs. (23) and (26) omit the normalization of the tilted distributions. The noiseless-channel example is decisive: Eq. (21) yields twice the true Sibson MI, and Eq. (25) similarly fails for Lapidoth–Pfister MI. This is not a matter of interpretation; it is an algebraic identity that fails as stated. The likely fix is to replace p_X^{1/α} by p_{X_{1/α}} and p_X^{α/(2α−1)} by p_{X_{α/(2α−1)}}, which is consistent with Lemma 2 and with the proof of Theorem 1 in Appendix A. I verified that Theorem 2's privacy ratios, which use the normalized tilted distributions, reproduce the correct values on the noiseless example, so those statements appear salvageable. The reader's verdict of CONDITIONAL is still appropriate, but the condition should include correcting Theorem 1 and Proposition 5, not just softening the novelty claims.","tokens_in":14022,"tokens_out":35656,"duration_ms":301306,"concrete_test":"Take binary X with p_X(0)=0.9, p_X(1)=0.1, α=2, and Y=X. Compute Eq. (21) as printed: H_{1/2}=2 log(√0.9+√0.1)≈0.470, and H^S_2 (Eq. 23) = -2 log(√0.9+√0.1)≈-0.470, so the RHS is ≈0.940. The definition of Sibson MI gives I^S_2(X;X)=min_q log(0.9/q_0+0.1/q_1)=2 log(√0.9+√0.1)≈0.470. The identity fails. Repeat with p_{X_{1/2}}(x)=p_X(x)^{1/2}/Σ√p in Eq. (23); then H^S_2=0 and Eq. (21) holds. Similarly check Eq. (25) with p_{X_{2/3}} instead of the unnormalized power.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1's definitions of the new conditional entropies are written with unnormalized powers: Eq. (23) uses p_X^{1/α}(x) and Eq. (26) uses p_X^{α/(2α−1)}(x), rather than the normalized tilted distributions p_{X_{1/α}} and p_{X_{α/(2α−1)}}. As printed, this makes the central representation false. For the noiseless channel p_{Y|X}=δ_{y=x} and α=2, Eq. (21) gives I^S_2(X;X)=H_{1/2}(X)-H^S_2(X|X)=2 log Σ√p - (-2 log Σ√p)=4 log Σ√p, while Proposition 2 and the definition of Sibson MI give I^S_2(X;X)=H_{1/2}(X)=2 log Σ√p; the printed RHS is exactly double. The same factor-of-two error occurs in Eq. (25) for I^LP. The missing normalization constant is Z=Σ p^{1/α}, so the printed RHS exceeds the true MI by H_{1/α}(X) (respectively H_{α/(2α−1)}(X)). Proposition 5's tilde entropies inherit this issue. The privacy ratios in Theorem 2 appear to use the normalized tilted distributions and are not affected, but Theorem 1—one of the paper's main answers to Q2—must be corrected. If p_X^{1/α} was intended as shorthand for the tilted distribution, the notation is nonstandard and the proof must state this explicitly.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies five variants of α-mutual information (Sibson, Arimoto, Augustin–Csiszár, Hayashi, and Lapidoth–Pfister) and addresses three questions: (Q1) whether Arimoto and Hayashi MI can be expressed via Rényi divergence; (Q2) whether Sibson, Augustin–Csiszár, and Lapidoth–Pfister MI can be expressed as Rényi entropy minus a conditional Rényi entropy; and (Q3) whether all five can be interpreted as privacy leakage measures based on a guessing adversary's gain functions. The main results are Proposition 4 for Q1, Theorem 1 for Q2, Proposition 5 for CRE and DPI of the newly defined conditional Rényi entropies, and Theorem 2 for Q3. The paper also introduces power-mean and generalized-geometric-mean interpretations of these leakage measures.","tokens_in":14380,"tokens_out":12296,"duration_ms":101917,"significance":"If Theorem 1 and Theorem 2 are correct, the paper provides a useful unification: all five α-MI measures are represented through a single reverse-channel variational form, and each is interpreted as the multiplicative increase of a maximal generalized mean of a guessing adversary's gain. This extends the α-leakage framework of Liao et al. [7] and the generalized-average interpretation of Sibson MI by Zarrabian and Sadeghi [38]. The proposed conditional Rényi entropies with CRE and DPI, though largely definitional, are a convenient byproduct. The mathematical style is mostly clear and the proofs are sketched with appropriate references, including reliance on the authors' earlier work [31] for the Augustin–Csiszár and Lapidoth–Pfister variational characterizations. However, the normalization error in Theorem 1 is load-bearing and must be corrected before the results can be accepted.","major_comments":[{"comment":"The definitions of H^S_α(X|Y) and H^LP_α(X|Y) use the unnormalized powers p_X^{1/α}(x) and p_X^{α/(2α−1)}(x), respectively, instead of the normalized tilted distributions p_{X_{1/α}} and p_{X_{α/(2α−1)}}. As printed, Theorem 1 is false. For a noiseless channel p_{Y|X}=δ_{y=x} and α=2, Eq. (21) yields I^S_2(X;X)=H_{1/2}(X)−H^S_2(X|X)=2 log Σ√p − (−2 log Σ√p)=4 log Σ√p, whereas Proposition 2 gives I^S_2(X;X)=H_{1/2}(X)=2 log Σ√p. The same factor-of-two error appears in Eq. (25) for I^LP_α. Notably, the proof of the LP case in Appendix A divides by the normalization constant in Eq. (81), which indicates that the displayed formulas are typographical rather than intended; nevertheless, the theorem as stated must be corrected to use the tilted distributions throughout.","section":"Section III-A, Eqs. (23) and (26)"},{"comment":"The proof of Proposition 5 states that the inequalities follow from nonnegativity and DPI of the corresponding mutual informations. That is correct, but it also reveals that the quantities ~H^S_α and ~H^LP_α are defined as H_α(X)−I^S_{1/α}(X;Y) and H_α(X)−I^LP_{1/(2α−1)}? (up to the index conventions), so their CRE and DPI properties hold by construction rather than being intrinsic properties of genuinely new conditional entropy definitions. The authors should acknowledge this definitional triviality explicitly and present Proposition 5 as a corollary of Theorem 1, not as an independent contribution. In addition, the statements in Proposition 5 inherit the normalization errors of Theorem 1 and need to be rechecked after the correction.","section":"Section III-B, Proposition 5"},{"comment":"The proof of the Lapidoth–Pfister privacy interpretation says that Eq. (58) follows from Theorem 1. Since Eq. (25) in Theorem 1 is misprinted, the derivation of Eqs. (58)–(59) is currently invalid, even though the final privacy ratios may be correct when the normalized tilted distribution is used. The proof needs to be reworked and cross-checked with the corrected Theorem 1. I verified that the other ratios in Theorem 2 that use p_{X_{1/α}} and p_{X_{α/(2α−1)}} are consistent with normalized tilts and are not affected by the typo, but the LP case specifically depends on the flawed Eq. (26).","section":"Section III-C, Theorem 2, Eqs. (58)–(59)"}],"minor_comments":[{"comment":"There is a typo in the sentence introducing p_Xα: it reads 'where and p_Xα := p^{(α)}_X'; the word 'and' should be removed.","section":"Proposition 1"},{"comment":"The notation \\|\\|r_{X|Y}(·|y)\\|\\|_α^α should be written as \\|r_{X|Y}(·|y)\\|_α^α (single bars around the function and α as subscript), to avoid confusion with a double-bar norm.","section":"Remark 2, Eq. (28)"},{"comment":"The notation for tilted distributions is inconsistent: Theorem 1 uses p_X^{1/α}(x) and p_X^{α/(2α−1)}(x), while Theorem 2 and Remark 5 use p_{X_{1/α}} and p_{X_{α/(2α−1)}}. Please standardize and explicitly define the tilted distribution at each point of use.","section":"Throughout"},{"comment":"In the denominator of Eq. (48), 'max_{r_\\hat X} G^{p_X}_{1/α}[g_∞]' is missing the argument; it should read 'max_{r_\\hat X} G^{p_X}_{1/α}[g_∞(X,r_\\hat X)]'.","section":"Section III-C, Eq. (48)"},{"comment":"The proof of Lemma 3 is long but relies on footnotes 3 and 4 to justify the minimax theorem; it would be helpful to state explicitly the concavity-convexity conditions in the main text or in a table, rather than only in footnotes.","section":"Appendix A, Lemma 3"},{"comment":"In Definition 3, the case α = 1 is written as log r(x) − 1, but in the later equations the α-score is used with α ∈ (0,1)∪(1,∞). Clarify that g_1 is used only for the limiting case and is not needed for α ∈ (0,1)∪(1,∞).","section":"Section II, Eq. (15)"}],"recommendation":"major_revision","confidential_remarks":"The normalization error in Theorem 1 appears to be a typographical mistake that is already corrected in the appendix for the LP case, so a revision is clearly feasible. The paper's reliance on the authors' prior work [31] for the Augustin–Csiszár and Lapidoth–Pfister variational characterizations is acceptable, but the novelty of Proposition 5 should be toned down to avoid overclaiming: the new conditional entropies are, by definition, H(X)−MI, so their CRE and DPI are immediate. The privacy interpretations in Theorem 2 are the strongest part of the paper and should be highlighted once the proof is repaired. I recommend that the editor request a revision that fixes the normalization, re-verifies Theorem 2's LP case, and clarifies the status of Proposition 5."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The two things to know: Theorem 1 as printed is false, and Theorem 2 is the part worth keeping. The privacy-leakage unification of all five α-MI variants is a genuinely useful piece of work, but it sits on top of a normalization error that needs to be fixed before anyone should build on it.\n\nThe issue is in Eqs. (23) and (26). They define the new conditional entropies using p_X^{1/α}(x) and p_X^{α/(2α−1)}(x), unnormalized. The proof of Theorem 1 in Appendix A goes through the tilted distribution p_{X_{1/α}} (and its LP analog), so the printed definitions lack the normalization constant Z. As written, plug a noiseless channel into Eq. (21) with α=2 and you get twice the Sibson MI; the same factor-two problem appears in Eq. (25). The fix is straightforward: replace those powers with the normalized tilted distributions defined in Eq. (1). The authors should have been careful here because the notation p_X^{1/α} is not defined anywhere and conflicts with their own p_{X_{1/α}} later.\n\nWhat is actually new and good: Proposition 4 gives clean divergence-form identities for Arimoto and Hayashi MI, though they follow quickly from the Gallager exponent forms. Theorem 2 is the real contribution, extending Liao et al.'s α-leakage to all five α-MI measures using generalized means and proper scoring rules. That is a useful way to match a privacy metric to a threat model, and the proof seems to go through once Theorem 1 is corrected.\n\nThe soft spots beyond the bug: the 'novel conditional entropies' in Section III-B are defined as H − I for nonnegative DPI measures, so their CRE and DPI properties are true by construction. The paper concedes as much in the proof of Proposition 5, but should say this explicitly in the body rather than presenting them as a discovery. The proofs in Appendix A are also sketchy: Lemma 3 for α>1 is just 'replace min with max,' and the Csiszár and Lapidoth–Pfister cases lean on [31]. That reliance is fine if [31] is solid, but a referee should verify.\n\nBottom line: this deserves peer review, but only after the normalization error is acknowledged. Send it, and ask the referee to check the definitions in Theorem 1 and re-derive the noiseless case. A corrected version would be a reasonable contribution to the α-MI toolkit.\n\nWho it's for: information theorists working on Rényi measures and privacy leakage. The paper is readable and the main idea is sound despite the misstep.","headline":"Fix the normalization bug in Theorem 1 and this is a useful unification of α-MI as privacy leakage; as printed, a headline result is off by a factor of two.","tokens_in":14898,"tokens_out":8017,"would_cite":false,"duration_ms":65022,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A17","94A15"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that all five standard α-mutual-information variants equal the log-ratio of an adversary's maximal generalized-mean gain after and before observing released data, with each variant selecting a different gain function and…","keywords":["α-mutual information","Rényi divergence","conditional Rényi entropy","privacy leakage","proper scoring rules","generalized mean","data processing inequality"],"falsifier":"Take a finite alphabet, e.g., binary $X$ with $P(X=1)=p$ and a binary channel, and compute the left- and right-hand sides of Theorem 2 for Arimoto and Sibson MI by direct numerical optimization over all randomized decision rules $r_{X|Y}$ and over $q_Y$ in the defining minimization of each $\\alpha$-MI. If any of the claimed equalities differs by more than numerical precision for some $p$ and $\\alpha$, the representation fails.","tokens_in":13810,"feed_emoji":"🔐","tokens_out":7802,"duration_ms":72732,"temperature":0.7,"pith_summary":"α-mutual information (α-MI) is a family of generalizations of Shannon mutual information controlled by a parameter α, and it appears in coding, hypothesis testing, and privacy. This paper tries to show that five standard α-MI variants—Sibson, Arimoto, Augustin–Csiszár, Hayashi, and Lapidoth–Pfister—are not just algebraically related but are all instances of one operational story: how much a guessing adversary's best expected gain increases after seeing released data Y. The increase is measured as a log-ratio of maximal generalized means of a gain function, with the particular gain function and mean prescribing which α-MI variant appears. Along the way, the paper derives new conditional Rényi entropies that keep two properties Shannon-style entropies have: conditioning reduces entropy, and data-processing cannot increase information. A sympathetic reader would care because this turns a zoo of definitions into a single adversary model, making the choice of metric a choice about the threat model.","feed_headline":"Five mutual-information variants collapse into one leakage ratio","feed_subtitle":"A single adversary-gain ratio now describes Sibson, Arimoto, Augustin–Csiszár, Hayashi, and Lapidoth–Pfister metrics","key_machinery":"The reverse-channel variational representation of mutual information and entropy is the engine: any conditional distribution $r_{X|Y}$ acts as a randomized decision rule for an adversary, and entropy terms are re-expressed as maxima over $r_{X|Y}$ of expected scores. The $\\alpha$-tilted (escort) distribution $p^{(\\alpha)}_X(x) = p_X(x)^\\alpha / \\sum_{x'} p_X(x')^\\alpha$ moves the input distribution between the Arimoto and Sibson expressions. The generalized (power) mean $M^p_t[\\cdot]$, its $q$-generalized geometric-mean relative $G^p_q[\\cdot]$, and Lemma 1's identity $M^p_{1-q} = G^p_q$ tie the score maxima to the log-ratio form. The named gain functions—$\\alpha$-score, pseudospherical score, power score—are all proper scoring rules, so the maximizing decisions are posterior or tilted-posterior estimators.","core_discovery":"The paper's central claim is Theorem 2: for every $\\alpha \\in (0,1) \\cup (1,\\infty)$, Arimoto MI, Sibson MI, Augustin–Csiszár MI, and Hayashi MI (and, for $\\alpha \\in (1/2,1) \\cup (1,\\infty)$, Lapidoth–Pfister MI) can each be written as a constant multiple of the logarithm of the ratio between the maximal generalized mean of an adversary's gain after observing $Y$ and the maximal generalized mean before observing $Y$. For Arimoto MI the gain is the $\\alpha$-score or pseudospherical score with ordinary expectation; for Sibson MI the same score is evaluated under an $\\alpha$-tilted distribution of the input; for Augustin–Csiszár MI the outer mean is geometric; for Hayashi MI the gain is the power score; for Lapidoth–Pfister MI the mean is a power mean of order $\\alpha/(2\\alpha-1)$. The paper also proves differential representations: Sibson MI is $H_{1/\\alpha}(X) - H^S_\\alpha(X|Y)$, Augustin–Csiszár MI is $H(X) - H^C_\\alpha(X|Y)$, and Lapidoth–Pfister MI is $H_{\\alpha/(2\\alpha-1)}(X) - H^{LP}_\\alpha(X|Y)$, with the new conditional entropies defined by reverse-channel minimization.","pith_inferences":["Beyond the paper: because the adversary in the Sibson case uses the tilted distribution $p_X^{1/\\alpha}$, the formalism suggests a design principle—choose $\\alpha$ to encode how much the threat model overweights rare inputs, and tune a privacy mechanism against that tilted adversary.","Beyond the paper: the generalized-mean viewpoint may extend to continuous alphabets or to other proper scoring rules; if a new score satisfies the same variational entropy identity, it would automatically produce a new member of the $\\alpha$-MI family.","Beyond the paper: the restriction of Lapidoth–Pfister MI to $\\alpha > 1/2$ in Theorem 2 (inherited from the definition's range) leaves open whether a suitable score or mean can cover $\\alpha \\le 1/2$; testing this on a concrete finite-alphabet example would clarify whether the restriction is fundamental.","Beyond the paper: the ratio form could be used as a numerical estimator of $\\alpha$-MI from samples—replace expectations by empirical means and optimize over reverse channels—yielding a plug-in leakage estimator with a clear adversary interpretation."],"forward_implications":["Each $\\alpha$-MI variant now has an explicit adversary: a randomized guesser maximizing an expected proper score, so privacy analyses can pick the variant matching the threat model.","The conditional Rényi entropies $H^S_\\alpha$, $H^C_\\alpha$, and $H^{LP}_\\alpha$ satisfy conditioning-reduces-entropy and data-processing inequalities, the properties needed in information-theoretic security proofs.","Because every representation is variational (optimization over reverse channels), the same numerical machinery that computes one $\\alpha$-MI can be reused for the others.","The representations answer the paper's three questions directly: Arimoto and Hayashi MI become Rényi divergences, Sibson, Augustin–Csiszár, and Lapidoth–Pfister MI become Rényi-entropy differences, and all five become leakage measures."],"supporting_citations":[{"why":"proves the Arimoto MI $\\alpha$-leakage ratio (Proposition 3 here) that Theorem 2 generalizes to the other four variants.","marker":"[7]"},{"why":"provides the Gallager $E_0$ representations of Sibson and Arimoto MI used to derive the Rényi-divergence and conditional-entropy formulas.","marker":"[2]"},{"why":"defines Lapidoth–Pfister MI and supplies the divergence characterization used in Lemma 3 and Theorem 1.","marker":"[17]"},{"why":"gives the variational representation of Augustin–Csiszár MI in terms of reverse channels that underlies Lemma 3.","marker":"[4]"},{"why":"establishes the nonnegativity, data-processing, and noiseless-channel properties of Sibson MI invoked in Propositions 2 and 5.","marker":"[41]"},{"why":"is the source of the earlier conditional Rényi entropies that satisfy conditioning-reduces-entropy and data-processing, against which the paper's new conditional entropies are compared.","marker":"[16]"}],"fun_headline_variants":["All α-MI forms collapse into one gain ratio","Single ratio now describes five MI variants","Leakage measure unified: one ratio for all α-MI","All five MI metrics become one adversarial gain ratio","α-MI unified: every variant is a single gain ratio"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The privacy interpretation rests on the assumption that an adversary's utility is faithfully described by the paper's gain functions ($\\alpha$-score, pseudospherical score, power score) and by the multiplicative increase of maximal expected gain as the measure of leakage.","fun_headline_variants_meta":{"raw":{"variants":["All α-MI forms collapse into one gain ratio","Single ratio now describes five MI variants","Leakage measure unified: one ratio for all α-MI","All five MI metrics become one adversarial gain ratio","α-MI unified: every variant is a single gain ratio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000813,"raw_usage":{"total_tokens":3570,"prompt_tokens":960,"completion_tokens":2610,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":2533}},"tokens_in":576,"tokens_out":2610,"duration_ms":19258,"temperature":1.0,"reasoning_tokens":2533,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:24:04.104483+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a finite alphabet, e.g., binary $X$ with $P(X=1)=p$ and a binary channel, and compute the left- and right-hand sides of Theorem 2 for Arimoto and Sibson MI by direct numerical optimization over all randomized decision rules $r_{X|Y}$ and over $q_Y$ in the defining minimization of each $\\alpha$-MI. If any of the claimed equalities differs by more than numerical precision for some $p$ and $\\alpha$, the representation fails.","supporting_citations":[{"cited_title":"Tunabl e measures for information leakage and applications to privacy-utili ty tradeoffs,","cited_arxiv_id":null,"evidence_quote":"proves the Arimoto MI $\\alpha$-leakage ratio (Proposition 3 here) that Theorem 2 generalizes to the other four variants."},{"cited_title":"Generalized cutoff rates and R´ enyi’s inf ormation measures,","cited_arxiv_id":null,"evidence_quote":"provides the Gallager $E_0$ representations of Sibson and Arimoto MI used to derive the Rényi-divergence and conditional-entropy formulas."},{"cited_title":"Two measures of depen- dence,","cited_arxiv_id":null,"evidence_quote":"defines Lapidoth–Pfister MI and supplies the divergence characterization used in Lemma 3 and Theorem 1."},{"cited_title":"On R´ enyi measures and hypothesis testin g,","cited_arxiv_id":null,"evidence_quote":"gives the variational representation of Augustin–Csiszár MI in terms of reverse channels that underlies Lemma 3."},{"cited_title":"Revisiting conditional R´ e nyi entropies and generalizing shannon’s bounds in information theoreti cally secure encryption,","cited_arxiv_id":null,"evidence_quote":"is the source of the earlier conditional Rényi entropies that satisfy conditioning-reduces-entropy and data-processing, against which the paper's new conditional entropies are compared."}],"review_version":1}