Pith. sign in

REVIEW 3 major objections 6 minor 23 references

An Informational Parsimony Perspective on Probabilistic Symmetries

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper proves that a single compression objective, the Divergence Information Bottleneck, exactly recovers channel invariances, channel equivariances, and distribution invariances in the discrete full-support case, and that relaxing…

desk verdict A genuine generalization of IB with a clean zero-distortion characterization, but the experiments sit outside the full-support regime the theory actually proves. read the letter →

arxiv 2412.08954 v2 pith:KMRC55I6 submitted 2024-12-12 cs.IT math.IT

classification cs.ITmath.IT MSC 94A1562B10
keywords DivergenceInformationBottleneckprobabilisticsymmetrieschannelequivariancedistributioninvariancesoftexponentialfamilygeometriccomplexity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that probabilistic symmetries — exact invariances and equivariances of channels and distributions — are not an extra layer imposed on data but can be derived from a single principle of informational parsimony. It introduces the Divergence Information Bottleneck (DIB), a rate-distortion-style problem that compresses a variable while preserving its divergence from a chosen hierarchical (exponential-family) model of non-structure. The central result states that, for discrete distributions with full support and unconstrained encoders, the fully divergence-preserving solutions of DIB coincide, up to lossless post-processing, with the clustering of the alphabet defined by the ratio identity $p(a)\tilde p(a') = p(a')\tilde p(a)$. By picking the hierarchical model appropriately, this one theorem yields exact characterisations of channel invariances, channel equivariances, and distribution invariances as those transformations that are quotiented out by the bottleneck. Once divergence preservation is only partial, the same construction defines nested soft symmetries whose coarseness is set by the compression trade-off, which the authors demonstrate on a synthetic grid-world channel.

What carries the argument

The load-bearing object is the Divergence Information Bottleneck functional, $\mathrm{DIB}(\lambda) = \arg\min_{\kappa \in C} I_\kappa(A; T)$ subject to $D(\kappa \cdot p \,||\, \kappa \cdot E) \geq \lambda$, where $D(p||E) = \inf_{r \in \mathrm{cl}\,E} D(p||r)$. The proof machinery is the ratio identity $p(a)\tilde p(a') = p(a')\tilde p(a)$: $\tilde p$ is the unique distribution in the closure of $E$ achieving the divergence, and the equality case of the log-sum inequality shows that full divergence preservation is possible only when the encoder's probabilistic pre-images are contained in the partition elements defined by that identity. A Blahut-Arimoto-style fixed-point algorithm derived from the Lagrangian relaxation provides the computational route to approximate solutions used in the experiments.

What would settle it

Take the grid-world channel of Section 4, restore full support by replacing the zero-probability position-direction pairs with a small uniform mass, and compare the $\mathrm{DIB}_{ce}(\Lambda)$ clustering to the partition defined by $p(x,y)p(x') = p(x',y')p(x)$; if any fully divergence-preserving encoder fails to factor through that partition, the characterisation in Theorem 4(i) is false.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is Theorem 3: if $C = C(A,T)$ and $\mathrm{supp}(p(A)) = A$, then $\mathrm{DIB}(\Lambda) = \{\gamma \circ \pi : \gamma \in C_{\mathrm{cong}}(\bar A, T)\}$ — every fully divergence-preserving encoder is, up to a congruent channel (a lossless post-processing), the projection $\pi$ onto the equivalence classes of the relation $p(a)\tilde p(a') = p(a')\tilde p(a)$, where $\tilde p$ is the projection of $p$ onto the exponential family $E$. Applied to the family $E_{ce} = \{r(X)U(Y)\}$, the theorem says a pair $(\sigma,\tau)$ is an equivariance of the channel $p(Y|X)$ exactly when the maximal bottleneck erases it: $\kappa \circ (\sigma \otimes \tau) = \kappa$ for $\kappa \in \mathrm{DIB}_{ce}(\Lambda)$. Applied to $E_{di} = \{U(A)\}$, it says the same for distribution invariances. The paper also shows the quotienting property holds at every trade-off level $\lambda$, even though full equivalence is proven only at $\lambda = \Lambda$, and it notes that the projection from the DIB clustering need not coincide with the true orbit projection in the equivariance case.

Load-bearing premise

The full-support condition $\mathrm{supp}(p(A)) = A$: the proof needs the projection $\tilde p$ of $p$ onto the chosen exponential family to lie in the interior of the probability simplex, and it is the load-bearing assumption for Theorems 3 and 4; the paper's own numerical experiment uses a distribution without full support, so the exact characterisation is proven outside that numerical regime.

Editorial extensions

If this is right

  • For any full-support discrete system, exact channel invariances, channel equivariances, and distribution invariances can be read off from the solutions of a single optimisation problem at $\lambda = \Lambda$, without enumerating group elements.
  • Lowering $\lambda$ turns exact symmetries into a nested family of $\lambda$-equivariances; the smaller the perturbation that breaks a symmetry, the less compression is needed to make it reappear, as shown in the grid-world experiment.
  • The classic Information Bottleneck is a special case of DIB with $E = E_{ce}$ and shape-restricted encoders, so the invariance-extraction behaviour of IB is a special case of the general characterisation.
  • The effective cardinality of DIB solutions increases monotonically with $\beta$, and changes of effective cardinality coincide with slope discontinuities of the information curve, giving a bifurcation picture of how soft symmetries emerge.
  • Symmetries in data afford drastic compression: in the experiments, most of the divergence from the hierarchical model is retained even after a large reduction of mutual information.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the right-hand side of the ratio identity depends only on $p$ and $\tilde p$ and not on any pre-specified group action, the framework suggests a data-driven notion of structure: any partition of the alphabet into constant-$p/\tilde p$ cells is a candidate symmetry that the data itself defines, even when no group is known in advance.
  • The full-support restriction is not purely technical: when zero-probability states exist, the DIB must send those states to a dummy symbol, so the effective cardinality gains an extra bin; this suggests support boundaries will show up as spurious structure in soft-symmetry recovery unless handled explicitly.
  • A natural testable extension is to continuous alphabets: if $p$ and $\tilde p$ are densities, the relation $p/\tilde p = \text{constant}$ defines level sets that could play the role of orbits, and the DIB fixed-point equations could be checked numerically for Gaussian or other exponential-family hierarchical models.
  • The bifurcation structure of effective cardinality suggests a principled way to choose the number of symmetry classes: stop compression where the preserved divergence $D(\kappa\cdot p||\kappa\cdot\tilde p)$ starts to fall sharply, in analogy with an elbow rule.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces the Divergence Information Bottleneck (DIB), a rate-distortion-like framework that trades off compression against preservation of the divergence from a hierarchical model, and uses it to characterize channel invariance, channel equivariance, and distribution invariance in terms of the solution sets of the corresponding bottleneck problems. The main theoretical results are Theorem 2 (classic IB solutions characterize channel invariances), Theorem 3 (for full-support p(A) and unconstrained encoders, the fully divergence-preserving DIB solutions are exactly the clusterings defined by the relation p(a)p~(a')=p(a')p~(a), up to congruent post-processing), Theorem 4 (equivariances are characterized by κ∘(σ⊗τ)=κ for κ at maximum trade-off parameter), and Theorem 5 (analogous statement for distribution invariances). The paper also proposes a notion of soft λ-equivariances, gives a Blahut-Arimoto-style algorithm for the DIB Lagrangian, and reports synthetic experiments where symmetry recovery occurs at bifurcations of the trade-off parameter.

Significance. If the results are made fully rigorous, the paper offers a genuinely useful bridge between information-theoretic parsimony and group-theoretic symmetry: it recovers the known invariance-extraction property of the IB in a group-theoretic language, extends it to equivariances without the uniformity assumption needed in prior work, and gives a principled way to soften exact symmetries. The appendices contain substantial proof detail, the statements are honest about the failure of the equivariance clustering to coincide with the orbit projection in general (Theorem 4(iii)), and the synthetic experiments provide a concrete demonstration of the proposed soft-symmetry extraction mechanism. The main theoretical gaps—the full-support assumption and the compactness of the optimization domain—do not invalidate the overall research direction but must be addressed before the central theorems can be relied upon as stated.

major comments (3)
  1. [Section 4.1, footnote 5] The experiments use a distribution p(X,Y) with supp(p(X,Y)) strictly contained in X×Y, while Theorem 3 and Theorem 4(i) are proved only under the assumption supp(p(A))=A. The paper's central observation that Dp(κβ||C_Gce) vanishes at bifurcations and that nested soft equivariances are recovered in increasing order of compression is therefore outside the proven regime. The authors acknowledge this and defer the non-full-support case to future work, but as written the numerical demonstration does not test the theorem. Please either extend Theorem 3 and Theorem 4 to the non-full-support case under suitable conditions (for example, the conditions used in Appendix D.4), or explicitly present the experiments as heuristic evidence outside the scope of the theorems.
  2. [Appendix C.2 (and Appendix B.2)] The proof of Theorem 3 asserts that the DIB problem is 'the minimisation of a continuous function on a compact domain' and concludes that a solution q* exists. However, the paper sets T:=N in Section 3.1, and the space C(A,T) of channels from a finite set A to a countably infinite alphabet T is not compact in the usual topologies on the simplex. The same compactness assertion is used in the proof of Theorem 2 in Appendix B.2. This is load-bearing for the converse inclusions E⊆DIB(Λ) and E⊆IB(Λ), since those arguments compare arbitrary elements of E with an existing minimizer. Please either restrict the theorems to finite bottleneck alphabets and justify the passage to T=N by a limiting argument, or prove existence of minimizers in a well-defined topology on C(A,T) for infinite T.
  3. [Section 3.1 / Appendix C.2, Lemma 15] The equality case of the log-sum inequality in Lemma 15 is used to conclude that p(a)/p~(a) is constant on each probabilistic pre-image A_q^t. This step depends on the fact that p~(a)>0 for every a, which follows from supp(p)=A only if the projection p~ lies in the interior of the simplex. The authors do prove in Appendix C.1 that full support of p implies full support of p~, so the logic is internally consistent. However, the proof of Lemma 15 contains a notational slip: the sum over a∈S is written with an undefined S, which should presumably be a∈A or a∈A_q^t. Please correct this and clarify the summation domain, since the equality condition is central to Theorem 3.
minor comments (6)
  1. [Appendix C.2, Lemma 14] In the proof of Lemma 14(i), the displayed equation for (26) states 'p(a)p~(a') = p(a)p~(a')' but the intended relation from equation (6) is 'p(a)p~(a') = p(a')p~(a)'. Please fix this typo.
  2. [Appendix C.6] In the proof of Theorem 5, the sentence 'This yields point (iii) of Theorem 3' should refer to point (ii) of Theorem 5, since the preceding implication establishes that elements of Gdi are quotiented out for all λ. The subsequent paragraph then proves point (iii).
  3. [Section 3.4 / Appendix D.2] The main text states that all alphabets are finite except T:=N, but the Blahut-Arimoto algorithm in Appendix D.2 explicitly assumes T is finite. Please state clearly in Section 3.4 or Appendix D that the algorithmic results apply to finite bottleneck alphabets, and explain how the infinite-alphabet theoretical results relate to the finite-alphabet algorithm used in the experiments.
  4. [Section 4.1 / Figure 1] The description of Figure 1 refers to colors representing a clustering of supp(p(X,Y)), but the figure itself is not reproduced in color in the text. Please ensure that the color coding is clear in the final version or add textual labels to the clusters.
  5. [Appendix D.3] The choice of threshold 10^{-3} for rounding |q(t|a)–q(t|a')| to zero when computing effective cardinality is stated but not justified. Please comment on the sensitivity of the reported bifurcation locations to this threshold, or justify the choice.
  6. [Section 6] There is a typo in the conclusion: 'exponetial families' should be 'exponential families'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the DIB characterization is derived self-containedly; the full-support gap is an acknowledged limitation, not a circular step.

full rationale

The paper's central result, Theorem 3, is a genuine mathematical derivation rather than a restatement of its inputs: the DIB problem is defined by optimizing I_kappa(A;T) subject to D(kappa.p || kappa.E) >= Lambda, with no mention of the partition pi or the relation p(a)p~(a') = p(a')p~(a). The proof derives pi from the equality condition of the log-sum inequality in Appendix C.2, so the characterization is not assumed by construction. The choice of E_ce is deliberately made so that the emergent relation becomes p(y|x) = p(y'|x'), but the paper proves this equivalence through equation (6) and the specific form of p~ = p(X)U(Y), rather than simply defining the DIB solution to be the equivariance clustering. The cited prior work, Charvin et al. (2023), is used only for elementary group-theoretic facts, such as the equivalence between the equivariance diagram and equality of conditional distributions, and for Proposition 17 about the classic IB; these are independently checkable and are not load-bearing for the main new characterization. The full-support caveat is explicitly acknowledged in footnote 5 and Section 5: the experiments use a non-full-support p(X,Y), so they fall outside the proven regime. That is a limitation or correctness risk, not circularity, because the experiments are presented as sanity checks rather than as predictions forced by the theorem. No fitted parameter is renamed as a prediction, and no uniqueness claim is imported from the authors' own earlier work to forbid alternatives. Therefore the derivation chain is self-contained and warrants a circularity score of 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central theorems rely on standard information-theoretic inequalities, the projection existence/uniqueness theorem for exponential families, and the full-support assumption. The numerical section introduces implementation parameters (rounding threshold, perturbation amplitudes) that affect the reported bifurcation curves but are not load-bearing for the theoretical results. No new physical or mathematical entities are postulated.

free parameters (2)
  • effective cardinality rounding threshold = 0.001
    Used in Appendix D.3 to decide when two symbols are considered equal in the computation of effective cardinality; this threshold affects the reported bifurcation points in Figure 2.
  • perturbation amplitudes in synthetic experiments = not specified
    The two perturbations applied to p(Y|X) in Section 4.1 have 'larger' and 'smaller' amplitudes, but numerical values are not given. These are hand-chosen to produce the nested equivariance structure and directly determine the observed recovery order.
assumptions (4)
  • standard math Log-sum inequality and its equality conditions (Csiszár and Körner, 2011)
    Used throughout the proofs of Lemmas 6, 11, 14, 15, and in Appendix C.1 to establish the projection property of p~.
  • domain assumption Existence and uniqueness of the projection of a distribution onto the closure of an exponential family (Ay et al., 2017)
    Invoked in Section 3.1 and Appendix C.1 to define p~ and to assert it simultaneously minimizes the latent-space divergence.
  • standard math Compactness of the space of channels C(A,T) with T = N under the relevant topology
    Used implicitly in the proofs of Theorems 2 and 3 to assert existence of a minimizer. With T countable, compactness holds in the product topology, but the paper does not justify this technical point.
  • domain assumption Full support of p(A) in the central theorems
    Theorem 3 and its applications require supp(p(A)) = A to ensure p~ is in the interior of the simplex, enabling the equality-case analysis in the log-sum inequality.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Informational Parsimony Perspective on Probabilistic Symmetries." pith.science (2026). https://pith.science/paper/KMRC55I6

@misc{pith2026241208954,
  author       = {Pith},
  title        = {Pith review of: An Informational Parsimony Perspective on Probabilistic Symmetries},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KMRC55I6}},
  note         = {Machine review of arXiv:2412.08954}
}
read the original abstract

Extraction of structure, in particular of group symmetries, is increasingly crucial to understanding and building intelligent models. In particular, some information-theoretic models of parsimonious learning have been argued to induce invariance extraction. Here, we formalise these arguments from a group-theoretic perspective. We then extend them to the study of more general probabilistic symmetries, through compressions preserving geometric measures of complexity. More precisely, our framework implements a trade-off between compression and preservation of the divergence from a given hierarchical model, yielding a novel generalisation of the Information Bottleneck framework. Through appropriate choices of hierarchical models, we fully characterise (in the discrete and full support case) channel invariance, channel equivariance and distribution invariance under permutation. Allowing imperfect divergence preservation then leads to principled definitions of "soft symmetries", where the "coarseness" corresponds to the degree of compression of the system. In simple synthetic experiments, we demonstrate that our method successively recovers, at increasingly compressed "resolutions", nested but increasingly perturbed equivariances, where new equivariances emerge at bifurcation points of the trade-off parameter. Our framework suggests a new path for the extraction of generalised probabilistic symmetries.

Figures

Figures reproduced from arXiv: 2412.08954 by the authors.

Figure 1
Figure 1. Left: representation of p(Y |X), where X is the position on the grid, Y the gradient direction, and probabilities are proportional to arrow lengths. Thus equivariances are here pairs (σ, τ ) that send each arrow on an arrow of equal length. Right: same figure with colors representing a clustering of supp(p(X, Y )) — which defines a clustering of X × Y if we add the cluster supp(p(X, Y ))c , made of position￾orientat… view at source ↗
Figure 2
Figure 2. Dβ := D(κβ · p||κβ · p˜) as a function of Iβ := Iκβ (X, Y ; T). Bottom left: Effective cardinality k(κ) as a function of Iβ. Right: Divergence of compression channels κβ as a function of Iβ, for the groups Gce and Grota. The vertical dashed lines represent specific bifurcations of the parameter β at which Dp (κβ||CGrota ), resp. Dp (κβ||CGce ), approximately vanishes (in decreasing order of Iβ). equivariant, but whe… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 15 canonical work pages

  1. [1]

    Emergence of Invariance and Disentanglement in Deep Representations

    Alessandro Achille and Stefano Soatto. Emergence of Invariance and Disentanglement in Deep Representations . In 2018 Information Theory and Applications Workshop ( ITA ) , pages 1--9, February 2018. doi:10.1109/ITA.2018.8503149

  2. [2]

    Critical Slowing Down Near Topological Transitions in Rate-Distortion Problems

    Shlomi Agmon, Etam Benger, Or Ordentlich, and Naftali Tishby. Critical Slowing Down Near Topological Transitions in Rate-Distortion Problems . In 2021 IEEE International Symposium on Information Theory ( ISIT ) , pages 2625--2630, July 2021. doi:10.1109/ISIT45174.2021.9517956

  3. [3]

    Alemi, Ian Fischer, Joshua V

    Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, and Kevin Murphy. Deep Variational Information Bottleneck . In International Conference on Learning Representations , February 2017

  4. [4]

    Matthew Ashman, Cristiana Diaconu, Adrian Weller, Wessel Bruinsma, and Richard E. Turner. Approximately Equivariant Neural Processes . Advances in Neural Information Processing Systems, 37: 0 97088--97123, December 2024

  5. [5]

    Information Geometry on Complexity and Stochastic Interaction

    Nihat Ay. Information Geometry on Complexity and Stochastic Interaction . Entropy, 17 0 (4): 0 2432--2458, April 2015. ISSN 1099-4300. doi:10.3390/e17042432

  6. [6]

    A geometric approach to complexity

    Nihat Ay, Eckehard Olbrich, Nils Bertschinger, and J \"u rgen Jost. A geometric approach to complexity. Chaos: An Interdisciplinary Journal of Nonlinear Science, 21 0 (3): 0 037103, September 2011. ISSN 1054-1500. doi:10.1063/1.3638446

  7. [7]

    u rgen Jost, H \^o ng V \^a n L \^e , and Lorenz Schwachh \

    Nihat Ay, J \"u rgen Jost, H \^o ng V \^a n L \^e , and Lorenz Schwachh \"o fer. Information Geometry , volume 64 of Ergebnisse Der Mathematik Und Ihrer Grenzgebiete 34 . Springer International Publishing, Cham, 2017. ISBN 978-3-319-56477-7 978-3-319-56478-4. doi:10.1007/978-3-319-56478-4

  8. [8]

    Towards Information Theory-Based Discovery of Equivariances

    Hippolyte Charvin, Nicola Catenacci Volpi, and Daniel Polani. Towards Information Theory-Based Discovery of Equivariances . In NeurIPS 2023 Workshop on Symmetry and Geometry in Neural Representations , November 2023

Show all 23 references
  1. [9]

    Information Theory : Coding Theorems for Discrete Memoryless Systems

    Imre Csisz \'a r and J \'a nos K \"o rner. Information Theory : Coding Theorems for Discrete Memoryless Systems . Cambridge University Press, Cambridge, 2 edition, 2011. ISBN 978-0-521-19681-9. doi:10.1017/CBO9780511921889

  2. [10]

    Gerken, Jimmy Aronsson, Oscar Carlsson, Hampus Linander, Fredrik Ohlsson, Christoffer Petersson, and Daniel Persson

    Jan E. Gerken, Jimmy Aronsson, Oscar Carlsson, Hampus Linander, Fredrik Ohlsson, Christoffer Petersson, and Daniel Persson. Geometric deep learning and equivariant neural networks. Artificial Intelligence Review, 56 0 (12): 0 14605--14662, December 2023. ISSN 1573-7462. doi:10...

  3. [11]

    An Information Theoretic Tradeoff between Complexity and Accuracy

    Ran Gilad-Bachrach , Amir Navot, and Naftali Tishby. An Information Theoretic Tradeoff between Complexity and Accuracy . In Gerhard Goos, Juris Hartmanis, Jan Van Leeuwen, Bernhard Sch \"o lkopf, and Manfred K. Warmuth, editors, Learning Theory and Kernel Machines , volume 277...

  4. [12]

    A Formal Account of Structuring Motor Actions With Sensory Prediction for a Naive Agent

    Jean-Merwan Godon, Sylvain Argentieri, and Bruno Gas. A Formal Account of Structuring Motor Actions With Sensory Prediction for a Naive Agent . Frontiers in Robotics and AI, 7, 2020. ISSN 2296-9144. doi:10.3389/frobt.2020.561660

  5. [13]

    Anderson Keller, Lyle Muller, Terrence J

    T. Anderson Keller, Lyle Muller, Terrence J. Sejnowski, and Max Welling. A Spacetime Perspective on Dynamical Computation in Neural Information Processing Systems , September 2024. Preprint at https://arxiv.org/abs/2409.13669

  6. [14]

    Hamza Keurti, Bernhard Sch \"o lkopf, Pau Vilimelis Aceituno, and Benjamin F. Grewe. Stitching Manifolds : Leveraging Interaction to Compose Object Representations into Scenes . In ICML 2024 Workshop on Geometry-grounded Representation Learning and Generative Modeling , June 2024

  7. [15]

    Romero and Suhas Lohit

    David W. Romero and Suhas Lohit. Learning Partial Equivariances From Data . Advances in Neural Information Processing Systems, 35: 0 36466--36478, December 2022

  8. [16]

    Learning and generalization with the information bottleneck

    Ohad Shamir, Sivan Sabato, and Naftali Tishby. Learning and generalization with the information bottleneck. Theoretical Computer Science, 411 0 (29): 0 2696--2711, 2010. ISSN 0304-3975. doi:10.1016/j.tcs.2010.04.006

  9. [17]

    Flow Factorized Representation Learning

    Yue Song, Andy Keller, Nicu Sebe, and Max Welling. Flow Factorized Representation Learning . Advances in Neural Information Processing Systems, 36: 0 49761--49782, December 2023

  10. [18]

    Pereira, and William Bialek

    Naftali Tishby, Fernando C. Pereira, and William Bialek. The information bottleneck method, April 2000. Preprint at https://arxiv.org/abs/physics/0004057

  11. [19]

    van der Ouderaa , Mark van der Wilk , and Pim de Haan

    Tycho F. van der Ouderaa , Mark van der Wilk , and Pim de Haan . Noether's Razor : Learning Conserved Quantities . Advances in Neural Information Processing Systems, 37: 0 135943--135965, December 2024

  12. [20]

    Approximately Equivariant Networks for Imperfectly Symmetric Dynamics

    Rui Wang, Robin Walters, and Rose Yu. Approximately Equivariant Networks for Imperfectly Symmetric Dynamics . In Proceedings of the 39th International Conference on Machine Learning , pages 23078--23091. PMLR, June 2022

  13. [21]

    Raymond W. Yeung. Information Theory and Network Coding . Springer, 2008

  14. [22]

    On the Information Bottleneck Problems : Models , Connections , Applications and Information Theoretic Views

    Abdellatif Zaidi, I \ n aki Estella-Aguerri , and Shlomo Shamai (Shitz). On the Information Bottleneck Problems : Models , Connections , Applications and Information Theoretic Views . Entropy, 22 0 (2), 2020. ISSN 1099-4300. doi:10.3390/e22020151

  15. [23]

    Deterministic annealing and the evolution of Information Bottleneck representations

    Noga Zaslavsky and Naftali Tishby. Deterministic annealing and the evolution of Information Bottleneck representations. August 2019. Preprint at https://www.nogsky.com/publication/2019-evo-ib/2019-evo-IB.pdf

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.