{"id":"690463a5-3973-4f0a-8b1d-a665e05ee5ec","arxiv_id":"2504.15779","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The average degree of redundancy equals the normalized sum of single-source mutual informations, and the average degree of vulnerability equals the normalized sum of leave-one-out conditional mutual informations.","lead":"This paper introduces two simple averages, called Shannon invariants, that capture how information about a target is spread across many sources. Because they depend only on Shannon entropy, they are computable for large systems where full information decomposition is intractable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The invariant identities are sound; the load-bearing gap is that their redundancy/synergy and 'average' interpretation assumes nonnegative PID atoms, which is unverified for the neural data and fails for some PID measures.","rationale":"The paper's central mathematical contribution, that rbar and vbar are computable from Shannon entropies alone and are independent of the choice of PID, follows directly from the consistency equation A7 and is correct. The stress-test therefore does not attack the invariant identities themselves. The load-bearing weakness is that the paper uses these invariants to make atom-level claims about redundancy, synergy, robustness, and vulnerability. Those claims are only entailed when the PID atoms are nonnegative, a property the paper acknowledges is satisfied by some but not all PID approaches. Since the empirical sections never verify nonnegativity for the neural activation distributions, the practical conclusions and the advertised resolution of PID ambiguities are conditional. The reader identified this same assumption as the weakest point, and the Appendix B Corollary 5 summation-index error further supports the view that the atom-level interpretation needs more care. A targeted check on the actual neural distributions with a nonnegative and a possibly-negative PID measure would settle whether the concern lands in practice. The theoretical core remains intact, so the existing CONDITIONAL verdict is appropriate and no verdict change is needed.","tokens_in":18854,"tokens_out":13209,"duration_ms":130920,"concrete_test":"On the trained MNIST network (n=5 neurons per hidden layer) and the autoencoder bottleneck (n=12, 16, 20), take 10 random source subsets of size n=3-4 and compute the exact PID atoms under two standard measures, e.g. IBROJA and I_ccs, using the same quantized activation distributions as Section IV. For each subset, compare I_r^(0)/I with 1-rbar and I_v^(0)/I with 1-vbar. If any negative atom or violation of Proposition 5's inequality appears, the paper's threshold interpretation is not warranted for these empirical data; if all tested PIDs have nonnegative atoms and the inequalities hold, the concern fails for these systems.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Propositions 1 and 2 (Eqs. 6 and 10) are algebraic consequences of the PID consistency equation (A7) and do not require atom nonnegativity; as pure identities they are solid. The central interpretive claim, that rbar and vbar are average degrees of redundancy and vulnerability, and that thresholds such as rbar<1 imply source-level synergy, does require nonnegative atoms. The paper explicitly assumes this in Section III and Appendix B (Prop. 5, Cor. 3, Prop. 6, Cor. 4), and notes that some PID measures violate it, but it never checks the activation distributions in Section IV. If negative atoms occur there, the bounds I_r^(0)/I >= 1-rbar and I_v^(0)/I >= 1-vbar can fail, so Figures 2 and 3 do not strictly support statements like 'redundancy increases with depth' as claims about actual PID atoms. Additionally, Corollary 5 in Appendix B sums k=1 to n over I_r^(k), where the derivation requires k>=2 and the displayed formula corresponds to proper redundancy; this summation-index error is a symptom that the atom-level interpretation is less settled than the invariant part. This does not invalidate the invariant identities, but it makes the 'resolving ambiguities' claim conditional on a nonnegativity property that is neither guaranteed nor empirically verified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces two aggregate quantities, the average degree of redundancy rbar and the average degree of vulnerability vbar, defined respectively as the ratio of the sum of marginal mutual informations to the joint mutual information (Eq. 6) and the ratio of the sum of conditional mutual informations to the joint mutual information (Eq. 10). The main theoretical results, Propositions 1 and 2, show that these quantities can be rewritten as weighted averages of PID atom degrees (Eqs. 5 and 9), and are therefore independent of the choice of PID measure. The paper further shows that the redundancy-synergy index satisfies RSI = (rbar-1)I(X;Y) (Cor. 1), and introduces a dual index DRSI = (1-vbar)I(X;Y) (Prop. 4, Cor. 2). It then applies these invariants to quantized deep neural networks, reporting layer- and training-dependent trends for an MNIST classifier and a face autoencoder. The core derivations are elementary counting arguments based on the PID consistency equation (A7), and the paper explicitly notes that the interpretational claims require nonnegative PID atoms.","tokens_in":19044,"tokens_out":11254,"duration_ms":103146,"significance":"The theoretical core is sound and potentially useful: if the nonnegativity assumption holds, rbar and vbar provide scalable summaries of high-order information structure that circumvent the super-exponential cost of full PID, and they yield a rigorous interpretation of the widely used RSI as well as a new DRSI. The derivations are parameter-free, transparent, and correct as algebraic identities under the standard PID consistency equation. The paper also frames the empirical analysis carefully as an exact computation on the training-set distribution, avoiding the injectivity pitfalls of continuous neural activations. The main limitations are that the interpretative conclusions depend on an unverified nonnegativity property of PID atoms in the neural data, and the empirical trends are presented without statistical tests or null models. With those caveats addressed, the framework would be a valuable contribution to multivariate information theory and its applications.","major_comments":[{"comment":"The interpretation of rbar and vbar as average degrees of redundancy and vulnerability, and the threshold statements such as rbar<1 implying source-level synergy, presuppose nonnegative PID atoms. The manuscript explicitly assumes nonnegativity in Section III (after Prop. 3) and in Appendix B (Prop. 5, Prop. 6, Cor. 3, Cor. 4), but it never verifies this property for the activation distributions analyzed in Section IV. Since several PID measures admit negative atoms for n>2, the values in Figures 2 and 3 are strictly statements about the Shannon ratios, not about nonnegative PID atoms; consequently, statements such as \"redundancy increases with depth\" in Section IV.B are not fully supported as claims about actual redundant information atoms. The authors should either verify nonnegativity for the specific PID they intend (or for the data at hand), or explicitly weaken the interpretational language throughout.","section":"Section III and Appendix B, Eqs. (21), (B2)–(B9)"},{"comment":"The empirical claims rest on medians and min-max ranges over only 10 runs, with no null models, confidence intervals, or statistical tests. For example, the asserted monotonic increase of redundancy with depth and the decrease of vulnerability in the MNIST classifier (Fig. 2c,d), and the ordering by bottleneck size in the autoencoder (Fig. 3d,e), could plausibly fall within run-to-run variability. The manuscript should either add appropriate statistical analyses (e.g., permutation tests against null models, or at least confidence intervals) or explicitly describe these results as qualitative, exploratory observations rather than established signatures.","section":"Section IV, Figures 2 and 3"}],"minor_comments":[{"comment":"The summation in Corollary 5 runs over k=1 to n for I_r^(k) and I_v^(k), but the derivation and the displayed threshold correspond to proper redundancy/vulnerability, i.e., k>=2. As written, the equivalence is false; for example, with n=2 and λ=0.5, taking I_r^(1)/I=0.6 and I_r^(0)/I=0.4 gives rbar=0.6, so the left-hand side of (B10) holds but the claimed rbar>1.5 does not. The index should be corrected to k=2,...,n.","section":"Appendix B, Corollary 5, Eqs. (B10)–(B11)"},{"comment":"The text says the architecture has three hidden layers, while the Figure 2 caption and the layer labels L3–L5 in the text indicate five hidden layers. Please reconcile this inconsistency.","section":"Section IV.B and Figure 2 caption"},{"comment":"The phrase \"The inset in E shows\" should refer to panel (e) or \"the inset in e\"; also, Figure 2's caption contains the typo \"wile\" for \"while\".","section":"Figure 3 caption"},{"comment":"Definition 1 defines Shannon-invariance for linear combinations of atoms, but the two main quantities rbar and vbar are ratios of such combinations. Consider extending the definition to include normalized invariants, or add a remark that all results apply to ratios of Shannon-invariant linear combinations.","section":"Definition 1"},{"comment":"The statements that rbar and vbar lie in [0,n] are asserted before the Appendix B proofs; a one-line justification using I(X_i;Y) ≤ I(X;Y) and I(X_j;Y|X_-j) ≤ I(X;Y) would improve readability.","section":"Section II.B and II.C"},{"comment":"References [42] and [49] are the same Chechik et al. conference paper and should be merged or cross-referenced.","section":"References"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the two invariants are correct and worth having; the interpretation and the empirical sections need work before I'd repeat the claims in public.\n\nThe actual math is clean. Propositions 1 and 2 give rbar and vbar as sums of marginal or conditional MIs normalized by joint MI, and they follow immediately from the PID consistency equation. That's a genuinely useful observation: you can compute a redundancy/vulnerability profile without choosing a PID measure, and the cost is linear in the number of sources. The RSI gets a nice interpretation as (rbar−1)·I, and the DRSI is the obvious dual. I believe the authors when they say the average redundancy overlaps with earlier constructions (Ehrlich et al., Varley & Hoel); they cite them, which is the right move.\n\nThe soft spots are real but localized. The inequalities in Appendix B and the “rbar<1 means synergy” statements assume non-negative PID atoms. That assumption is standard in some definitions but not all, and the authors never check the actual activation distributions in Section IV. So Figure 2 can only support those conclusions under an extra condition that is plausible but not established. The empirical sections otherwise are thin: medians and min–max over 10 runs, no null models, no significance tests, and no reproducible code or data shipped. The authors mention null models as future work, which is honest but shows the current results are anecdotal. Also, Corollary 5 in Appendix B sums k from 1 to n over I_r^(k), while the derivation requires k≥2; that's a typo but it's the kind of typo that makes you pause when the atom-level claims are load-bearing.\n\nThe abstract overreaches. This does not “resolve long-standing ambiguities” about redundancy measures; it shows that certain averages are independent of the choice of PID. That is a real step, but the conceptual ambiguity about what the atoms mean remains, and the resolution is conditional on a non-negativity property. I'd suggest the authors soften that framing.\n\nBottom line: this deserves a proper peer review. The invariant identities are valuable for anyone doing information decomposition at scale, and the RSI interpretation is a nice payoff. But I'd want the authors to (1) fix Corollary 5, (2) either verify non-negativity for the analyzed distributions or explicitly qualify the claims, (3) add some statistical discipline to the neural experiments, and (4) revise the abstract to match what is actually proven. If they do that, I'd be happy to see it published.","headline":"The invariant identities are correct and worth publishing, but the interpretive claims and empirical sections need substantial tightening.","tokens_in":19645,"tokens_out":2148,"would_cite":true,"duration_ms":20281,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A17","94A15"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that the average degree of redundancy and the average degree of vulnerability are Shannon invariants, so they can be computed from Shannon entropies without choosing a PID measure.","keywords":["partial information decomposition","Shannon invariants","redundancy-synergy index","dual redundancy-synergy index","mutual information","higher-order information","deep neural networks","information decomposition"],"falsifier":"For a three-source distribution, compute $\\bar r$ and $\\bar v$ from their entropy formulas, then compute the atom-level sums under any decomposition satisfying Eq. (A7) that permits a negative atom. If the inequality $I_r^{(0)}/I(X;Y) \\ge 1-\\bar r$ is violated, or if $\\bar r<1$ occurs while every source-level-synergy atom $I_r^{(0)}$ is zero, the interpretive reading of the invariants fails even though the entropy formulas still hold.","tokens_in":18624,"feed_emoji":"🧮","tokens_out":8834,"duration_ms":76981,"temperature":0.7,"pith_summary":"Partial information decomposition splits the information a set of sources carries about a target into many fine-grained atoms, but the atoms cannot be uniquely defined and their number explodes with the number of sources. This paper shows that two aggregate quantities of those atoms, the average degree of redundancy and the average degree of vulnerability, are Shannon invariants: their values are fixed by ordinary Shannon entropies alone and are therefore the same under every decomposition. The redundancy average equals the ratio of the sum of individual mutual informations to the joint mutual information, and the vulnerability average equals the corresponding ratio for conditional mutual informations. Because these ratios are easy to compute, the framework brings information-decomposition reasoning to systems with many variables, such as the layers of deep neural networks, and gives the long-used redundancy-synergy index a precise meaning.","feed_headline":"Redundancy and vulnerability averages need no decomposition","feed_subtitle":"These mutual-information ratios give the redundancy-synergy index a clear meaning and scale to deep networks.","key_machinery":"The load-bearing object is the PID atom lattice: each information atom $\\Pi(\\alpha)$ is indexed by an antichain $\\alpha$ of source subsets, and the consistency equation (A7) fixes how atoms compose every mutual information term. On this lattice the paper puts two integer counting functions: $r(\\alpha)$, the number of singleton source sets $\\{i\\}\\in\\alpha$, and $v(\\alpha)$, the number of indices that appear in every subset of $\\alpha$. The key identity is a double-counting argument: summing the marginal mutual informations over $i$ counts each atom once per accessible singleton source, giving $\\sum_i I(X_i;Y)=\\sum_\\alpha r(\\alpha)\\Pi(\\alpha)$, and summing the conditional mutual informations counts each atom once per critically necessary source, giving $\\sum_j I(X_j;Y|X_{-j})=\\sum_\\alpha v(\\alpha)\\Pi(\\alpha)$. Normalizing by $I(X;Y)$ converts these identities into the two Shannon-invariant averages.","core_discovery":"The paper's central claim is that although individual PID atoms are not Shannon quantities, certain averages over them are. For $n$ sources $X_1,\\dots,X_n$ and a target $Y$, it defines the average degree of redundancy as $\\bar r = \\sum_{i=1}^n I(X_i;Y) / I(X;Y)$ and proves in Proposition 1 that this equals the atom-weighted average of $r(\\alpha)$, the number of singleton sources through which an atom is accessible. It defines the average degree of vulnerability as $\\bar v = \\sum_{j=1}^n I(X_j;Y|X_{-j}) / I(X;Y)$ and proves in Proposition 2 that this equals the atom-weighted average of $v(\\alpha)$, the number of sources on which an atom critically depends. The paper then reads the redundancy-synergy index through the first average, $\\mathrm{RSI} = (\\bar r - 1)I(X;Y)$, and introduces a dual index $\\mathrm{DRSI} = (1-\\bar v)I(X;Y)$, showing both are Shannon invariants. If the claims are right, these averaged quantities are properties of the joint distribution itself, not of any chosen decomposition.","pith_inferences":["If these invariants are as fundamental as claimed, they can serve as consistency checks: any proposed PID measure must reproduce the same $\\bar r$ and $\\bar v$ from its atoms, so deviations signal a violation of the counting identities.","The same double-counting logic likely extends to other coefficient choices: any linear combination of atoms whose coefficients are functions of $r(\\alpha)$ and $v(\\alpha)$ may be Shannon-invariant whenever the coefficient pattern collapses by the same counting argument, suggesting a larger family of invariants than the two introduced here.","The vulnerability average could be deployed directly as a robustness diagnostic for machine-learning representations, for example comparing models of different widths or after pruning, because it needs only conditional mutual informations that can be estimated from finite samples.","The paper's network analyses treat the training set as the full population; applying the same invariants on held-out data would test whether the redundancy and vulnerability trends reflect generalization rather than memorization of the training set."],"forward_implications":["The redundancy-synergy index is no longer a heuristic: $\\mathrm{RSI}=(\\bar r-1)I(X;Y)$, so its sign and magnitude are set by how the average degree of redundancy deviates from 1.","The new dual index satisfies $\\mathrm{DRSI}=(1-\\bar v)I(X;Y)$, giving a second, vulnerability-based balance measure that is also computable from Shannon entropies.","Under non-negative atoms, $\\bar r<1$ forces some source-level synergy and $\\bar v<1$ forces some robust information, with quantitative bounds in Appendix B for $\\bar r>1$ and $\\bar v>1$.","Shannon invariants require only entropies of small subsets, so the computational cost scales linearly with the number of sources, unlike a full PID.","In the trained feed-forward classifier studied here, redundancy increases across layers and training while vulnerability decreases; in the autoencoder, larger bottlenecks show higher redundancy and lower vulnerability."],"supporting_citations":[{"why":"Supplies the partial information decomposition framework and the lattice of information atoms over which the invariants are defined.","marker":"[24]"},{"why":"Provides the atom-focused formulation of PID and the consistency equation (A7) used in the proofs.","marker":"[39]"},{"why":"Gives the Shannon entropy and mutual information definitions underlying every Shannon-invariant quantity.","marker":"[45]"},{"why":"Supplies the interaction information identity that motivates the redundancy-minus-synergy difference for two sources.","marker":"[47]"},{"why":"Introduces the redundancy-synergy index that the paper reinterprets via the average degree of redundancy.","marker":"[49]"},{"why":"Provides the stochastic quantization approach used to make neural-network activations amenable to information-theoretic analysis.","marker":"[27]"},{"why":"Establishes why deterministic injective networks make standard information quantities ill-defined, motivating the quantization setup.","marker":"[58]"},{"why":"Provides the MNIST benchmark on which the feed-forward classification experiment is run.","marker":"[60]"},{"why":"Provides the Labeled Faces in the Wild dataset on which the autoencoder experiment is run.","marker":"[61]"}],"fun_headline_variants":["Averaging redundant and vulnerable atoms yields Shannon invariants","Shannon invariants: averages that skip PID decomposition","Redundancy and vulnerability: scaling info decomposition","New invariants resolve PID ambiguity and scale to deep nets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The interpretation that $\\bar r<1$ means source-level synergy and $\\bar v<1$ means robustness, and the bounds in Appendix B, all assume that no information atom is negative; several PID measures allow negative atoms when there are more than two sources.","fun_headline_variants_meta":{"raw":{"variants":["Averaging redundant and vulnerable atoms yields Shannon invariants","Shannon invariants: averages that skip PID decomposition","Redundancy and vulnerability: scaling info decomposition","New invariants resolve PID ambiguity and scale to deep nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000241,"raw_usage":{"total_tokens":1529,"prompt_tokens":957,"completion_tokens":572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":508}},"tokens_in":573,"tokens_out":572,"duration_ms":4858,"temperature":1.0,"reasoning_tokens":508,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:18:53.412688+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a three-source distribution, compute $\\bar r$ and $\\bar v$ from their entropy formulas, then compute the atom-level sums under any decomposition satisfying Eq. (A7) that permits a negative atom. If the inequality $I_r^{(0)}/I(X;Y) \\ge 1-\\bar r$ is violated, or if $\\bar r<1$ occurs while every source-level-synergy atom $I_r^{(0)}$ is zero, the interpretive reading of the invariants fails even though the entropy formulas still hold.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the partial information decomposition framework and the lattice of information atoms over which the invariants are defined."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the Shannon entropy and mutual information definitions underlying every Shannon-invariant quantity."},{"cited_title":"The Fast M\\\"obius Transform: An algebraic approach to information decomposition","cited_arxiv_id":"2410.06224","evidence_quote":"Supplies the interaction information identity that motivates the redundancy-minus-synergy difference for two sources."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the stochastic quantization approach used to make neural-network activations amenable to information-theoretic analysis."},{"cited_title":"Chechik, A","cited_arxiv_id":null,"evidence_quote":"Establishes why deterministic injective networks make standard information quantities ill-defined, motivating the quantization setup."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the MNIST benchmark on which the feed-forward classification experiment is run."},{"cited_title":"Griffith and C","cited_arxiv_id":null,"evidence_quote":"Provides the Labeled Faces in the Wild dataset on which the autoencoder experiment is run."}],"review_version":1}