Pith. sign in

REVIEW 2 major objections 3 minor 121 references

Learning Optimal Representations with the Decodable Information Bottleneck

T0 review · 2 major / 3 minor · reviewed 2026-08-27 · deepseek-v4-flash

Pith's one-line read A representation that is V-minimal and V-sufficient for the classifier family V forces every empirical risk minimizer on every dataset to hit the best achievable test risk.

desk verdict The DIB framework and generalization probe are genuinely useful, but the main optimality theorem rests on a monotonic-biasing assumption that fails for multiclass softmax networks, so the theorem does not cover the paper's own experiments as stated. read the letter →

arxiv 2009.12789 v2 pith:WZ2NDTKV submitted 2020-09-27 cs.LG cs.ITmath.ITstat.ML

classification cs.LGcs.ITmath.ITstat.ML MSC 68T0568T0794A17
keywords decodableinformationbottleneckV-informationrepresentationlearninggeneralizationgapempiricalriskminimizationminimalsufficientrepresentationsPACestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks what makes a learned representation optimal for supervised learning, and answers that optimality must be defined relative to the family V of classifiers that will consume the representation, not by raw Shannon compression. It defines a representation as V-sufficient when classifiers in V can reach the best achievable loss from it, and V-minimal when no classifier in V can distinguish examples that share a label. The central result, Theorem 1, states that a V-minimal V-sufficient representation forces every empirical risk minimizer in V on every dataset to reach the best achievable test risk. That would mean a representation trained with the proposed decodable information bottleneck objective can enforce generalization regardless of how complex or unregularized the downstream classifier is.

What carries the argument

The machinery is $V$-information, $I_V[Z\to Y] = H[Y] - \inf_{f\in V} \mathbb{E}[-\log f[Z](Y)]$, which measures how much of the label can actually be decoded by predictors in $V$, and its extension to decompositions: for each label $y$, $Dec(X,y)$ is the set of all deterministic labelings of the conditional inputs $X_y$, and $I_V[Z\to Dec(X,Y)]$ averages $I_V[Z_y\to N]$ over those labelings. A representation is $V$-minimal exactly when this quantity is zero, meaning no decoder in $V$ can distinguish same-label examples. The proof of Theorem 1 uses a closure property called monotonic biasing of $V$ to convert a hypothetical ERM that overfits into a decoder that separates training from test examples, contradicting $V$-minimality.

What would settle it

To test the theorem's applicability, train a fixed-width MLP encoder with the DIB objective and then explicitly search for an empirical risk minimizer with positive test risk by optimizing train loss minus a small multiple of test loss; a worst-case ERM whose test loss exceeds the $V$-sufficient lower bound would break the claim. Separately, for a concrete $V$ such as a single hidden layer MLP, take any $f\in V$ and check whether there exists a $g\in V$ that changes $f$ at one representation $z_0$ to an arbitrary probability while preserving all pairwise prediction signs; if no such $g$ exists, the monotonic biasing assumption fails and the proof's construction is unavailable.

Watch

Extended reading notes

Core claim

In the paper's own terms, the discovery is a characterization of optimal representations: under finite sample spaces, log loss, and deterministic labels, a representation $Z$ that is $V$-minimal and $V$-sufficient guarantees that any empirical risk minimizer $\hat f \in V$ satisfies $R(\hat f,Z) = \min_Z \min_{f\in V} R(f,Z)$. $V$-sufficiency ensures that some classifier in $V$ achieves the best possible loss, and $V$-minimality ensures that no classifier in $V$ can carve the training set out of the test set, because any classifier that did would carry positive $V$-information about a decomposition of the inputs within a label class. The decodable information bottleneck objective $L_{DIB} = -I_V[Z\to Y] + \beta I_V[Z\to Dec(X,Y)]$ is the Lagrangian relaxation of this characterization, and the paper shows it can be estimated from finite samples with PAC-style guarantees.

Load-bearing premise

The theorem relies on the assumption that the classifier family $V$ is closed under monotonic biasing: from any classifier one can build another classifier in $V$ whose prediction at a single representation point can be set to any probability while preserving the ordering of every other pair of predictions.

Editorial extensions

If this is right

  • Training an encoder with DIB for a known downstream family $V$ should remove the need for classifier-side regularization: every ERM in $V$, including adversarially selected worst-case ones, reaches the best achievable test loss.
  • The DIB objective is estimable from finite data with PAC-style bounds that grow with the Rademacher complexity of $\log V$, so it becomes easier to estimate when $V$ is small and harder in the universal limit where it recovers the classical information bottleneck objective.
  • $V$-minimality of a network's hidden representation correlates with its generalization gap more strongly than output entropy, path norm, gradient variance, and sharpness across 562 trained convolutional networks.
  • Penalizing $I_V[Z\to Dec(X,Y)]$ removes decodable information about spurious correlates, such as the overlaid MNIST digits in the paper's CIFAR experiments, so DIB-trained representations are robust to shortcut features.
  • When the downstream family is the universal family $U$, $V$-minimal $V$-sufficiency collapses to ordinary minimal sufficiency, recovering the classical information bottleneck as the special case where the decoder is unconstrained.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is a quantitative generalization bound for approximate $V$-minimality: since its experiments show the train-test gap shrinking monotonically as $I_V[Z\to Dec(X,Y)]$ decreases, one could try to prove $\text{gap} \le C\, I_V[Z\to Dec(X,Y)]^\alpha$, which would make the framework useful when perfect $V$-minimality is unreachable.
  • The monotonic biasing closure assumption suggests the proof techniques transfer most directly to families that are closed under output-bias manipulation, such as softmax classifiers with adjustable final biases; architectures without such closure would need a different proof or a relaxed notion of minimality.
  • The framework's logic implies a recipe for transfer learning: measure $V$-minimality of a frozen encoder's representation against the target family, and use it to decide whether a candidate pretrained encoder will generalize before training a downstream head.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper proposes the Decodable Information Bottleneck (DIB), a representation-learning objective that replaces Shannon mutual information in the classical IB with V-information relative to a prescribed decoder family V. It defines V-sufficient and V-minimal representations, states an optimality theorem (Theorem 1) claiming that any V-minimal V-sufficient representation makes every empirical risk minimizer in V achieve the best achievable test risk under deterministic labels, proves structural properties of the proposed notions (recoverability, monotonicity, characterization, existence), and gives an estimation bound for the DIB objective. The experimental sections study V-sufficiency and V-minimality on CIFAR variants and MNIST, including a worst-ERM protocol and a large correlation study across 562 trained CNNs.

Significance. If Theorem 1 holds as stated, the paper is a genuinely useful conceptual contribution: it transfers the classical notions of sufficient and minimal statistics from an unconstrained universal family to a pragmatic predictive family V, and it gives a plausible mechanism by which representations, rather than hypothesis classes, can enforce generalization. The proof structure is transparent, the properties in Proposition 2 are nontrivial, and the empirical work is extensive; in particular, the worst-ERM evaluation in Section 4.2 and the 562-model correlation study in Section 4.3 are valuable strengths, and the authors provide code for reproducibility. The significance is currently conditional, however: the main theorem relies on a closure assumption on V whose stated justification for neural networks is not correct, so the theorem's applicability to the paper's own experimental setting is not established as written.

major comments (2)
  1. [Appendix B.2 and Appendix C.3.3] The 'monotonic biasing of V' assumption is load-bearing for Theorem 1, but the justification given for neural networks is incorrect for multiclass softmax classifiers. The assumption requires that for any f' and any label y, one can construct g by 'modifying the bias term of the final softmax layer' while preserving the sign of every pairwise difference f'[z'](y')-f'[z''](y') for all y'. Adding a bias to one logit changes the normalization of all other classes, so the order of the predicted probabilities for an unmodified label can flip. Concretely, for a 3-class softmax with logits z=(0,0,ln 100) and z'=(ln 10,0,0), adding a bias b to class 0 gives p_1(z)=1/102 and p_1(z')=1/12 at b=0, so p_1(z')>p_1(z), while at b=4, p_1(z)=1/155.6 and p_1(z')=1/548.0, so p_1(z)>p_1(z'); the sign flips. The proof of Theorem 1 uses the assumption to construct the decoder g and also uses it in the multiclass-to-binary reduction, so the theorem as stated does not currently apply to the 10-class CIFAR-10 networks used in Section 4.2. The theorem may be repairable with a weaker single-label order-preservation condition, which does appear to hold for softmax outputs, but the current statement of Assumption B.2 and the proof must be revised.
  2. [Proposition 7 / Lemma 11] The claimed PAC estimation guarantee for the V-minimality term is not a vanishing bound. Lemma 11 bounds the total estimation error of the V-minimality term by beta*log|Y|, which is independent of the sample size M, and the proof of Proposition 7 carries this constant into the final bound. Consequently, the bound in Eq. (14) does not tend to zero as M grows, so the statement that DIB "can be estimated with guarantees" should be qualified. The authors themselves note in the paragraph after Lemma 11 that the bound is loose and 'is not a PAC-style bound,' which is in tension with the title of Proposition 7 and with the abstract's claim of estimation guarantees.
minor comments (3)
  1. [Appendix D.1.2] The description of V+-DIB says it uses an MLP with 8192 hidden units '(instead of 8192)'; this should presumably read '(instead of 128)', matching the preceding V- and V entries.
  2. [Appendix C.4, Lemma 7] In the last line of the proof of Lemma 7, the condition is written as I[Y; Z|Y]=0; this should be I[X; Z|Y]=0 to be consistent with the conclusion Z⊥X|Y.
  3. [Section 3.3] The main text says LDIB 'inherits V-information's estimation bounds' and refers to Appendix C.5, but the bound for the V-minimality term is a non-vanishing constant; the main text should point the reader to this caveat.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: Theorem 1 is derived from the V-minimality definition and external V-information results, not from its conclusions.

full rationale

The paper's central claim, Theorem 1, is a genuine derivation rather than a restatement of its inputs. V-minimality is defined as V-sufficiency plus minimal average V-information with y-decompositions of X, and the proof in Appx. C.3.3 supplies the nontrivial intermediate steps: it assumes a zero-empirical-risk predictor with positive test risk, constructs a nuisance labeling N in Dec(X,y), uses the monotonic-biasing assumption to build a decoder g that beats the marginal predictor on that labeling, and concludes IV[Z->N] > 0, contradicting V-minimality via Lemma 5. This is a substantive implication, not a definitional equivalence. The paper does not fit parameters and then relabel them as predictions; the DIB objective and its PAC bound in Prop. 7 are evaluated against the external V-information framework of Xu et al., with no self-citation chain carrying the argument. The 'monotonic biasing of V' assumption is load-bearing and its stated softmax justification is questionable for multiclass networks, but an unsupported or false assumption is a correctness or soundness risk, not circularity. The experimental generalization probe is an empirical correlation study, and no fitted constant is presented as a derived prediction. Overall the derivation chain is self-contained and not circular.

Assumptions & free parameters 3 free parameters · 8 assumptions · 1 invented entities

The central theorem depends on several explicit assumptions about the predictive family V, the most fragile being monotonic biasing. The deterministic-label and finite-space assumptions limit scope. Free parameters beta, K, and gamma are tuning knobs for the objective and evaluation procedure, not fitted constants that determine the theorem's validity.

free parameters (3)
  • beta (Lagrangian multiplier for DIB) = tuned per experiment (e.g., beta=10 for CIFAR10 DIB)
    Controls the trade-off between V-sufficiency and V-minimality in the DIB objective. The theorem applies only at exact V-minimality, but in practice a finite beta is tuned for best performance.
  • K (number of Monte Carlo y-decompositions) = 4 in most experiments
    Used to estimate the V-minimality term; the paper shows results are not very sensitive to K.
  • gamma (worst-ERM search coefficient) = 0.1
    Used in the Lagrangian relaxation to find an ERM that generalizes poorly; chosen so that the found predictor is approximately an ERM.
assumptions (8)
  • ad hoc to paper Monotonic biasing of V (Appx. B.2)
    Crucial for the proof of Theorem 1 to construct a decoder g that separates training and test examples within a class. It is a strong closure property; the paper only sketches why it holds for neural networks.
  • domain assumption Deterministic labeling Y = t(X) (Appx. B.3)
    Assumed for Theorem 1 and Prop. 2. Authors state it holds for common ML datasets but may not hold in general.
  • domain assumption Invariance of V to label permutations (Appx. B.2)
    Used to simplify the proof of Theorem 1; assumed to hold for practical families.
  • domain assumption Non-empty preimage of labels (Appx. B.2)
    Assumes each label has a representation that predicts it with probability 1; used in Prop. 6 to show best achievable risk is 0.
  • domain assumption Arbitrary constant prediction of V (Appx. B.2)
    Assumes V can always predict any constant distribution; used to equate HV[Y|∅] with H[Y].
  • standard math Finite sample spaces and |Z| >= |Y| (Appx. B.1)
    Restricts the setting to finite r.v.s and ensures sufficient capacity for label representations.
  • domain assumption Log-loss scoring rule (Appx. B.1)
    The theory and experiments focus on log loss; authors claim extension to other proper scoring rules is likely.
  • standard math V-information properties from Xu et al. [23] (non-negativity, monotonicity, estimation bounds)
    The paper relies on these prior results as external benchmarks, which are cited and not re-derived.
invented entities (1)
  • y-decomposition of X (Dec(X,y))
    purpose: Defines V-minimality by averaging V-information over all labelings within a class.
    A mathematical construction introduced to measure decodable information within a class; it is a definition, not an empirically falsifiable entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Learning Optimal Representations with the Decodable Information Bottleneck." pith.science (2026). https://pith.science/paper/WZ2NDTKV

@misc{pith2026200912789,
  author       = {Pith},
  title        = {Pith review of: Learning Optimal Representations with the Decodable Information Bottleneck},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WZ2NDTKV}},
  note         = {Machine review of arXiv:2009.12789}
}
read the original abstract

We address the question of characterizing and finding optimal representations for supervised learning. Traditionally, this question has been tackled using the Information Bottleneck, which compresses the inputs while retaining information about the targets, in a decoder-agnostic fashion. In machine learning, however, our goal is not compression but rather generalization, which is intimately linked to the predictive family or decoder of interest (e.g. linear classifier). We propose the Decodable Information Bottleneck (DIB) that considers information retention and compression from the perspective of the desired predictive family. As a result, DIB gives rise to representations that are optimal in terms of expected test performance and can be estimated with guarantees. Empirically, we show that the framework can be used to enforce a small generalization gap on downstream classifiers and to predict the generalization ability of neural networks.

Figures

Figures reproduced from arXiv: 2009.12789 by the authors.

Figure 1
Figure 1. Left two plots: illustration of representations learned by joint empirical risk minimization (J-ERM) and our decodable information bottleneck (DIB), for classifiers V with linear vertical decision boundaries. (a) For representations learned by J-ERM, there may exist an ERM that does not generalize; (b) Representations learned by DIB ensure that any ERM will generalize to the test set (V-minimality) . Right two plots… view at source ↗
Figure 2
Figure 2. Practical DIB (a) Pseudo-code to compute the LˆDIB(D); (b) Illustration of DIB to train a neural encoder. V-sufficiency corresponds to the standard log loss. V-minimality heads are trained to classify K arbitrary labeling within each class but the gradients w.r.t. the encoder are reversed so that Z cannot be used for that task. Each head has different parameters but the same architecture V. Proposition 2. Let V ⊆ V+… view at source ↗
Figure 3
Figure 3. Optimality of V-sufficiency. Plots of Alice’s best possible performance with different VBob-sufficient representations. The log likelihood is column-wise scaled from 0 to 100, and vertical separators are present to discourage between-column comparison. The predictive families are MLPs with varying widths. (a) Samples of Bob’s 2D representations along with Alice’s decision boundaries for an odd-even CIFAR100 binary c… view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Effect of DIB on generalization Left two plots (CIFAR10): Impact of DIB’s β on: (a) the train and test performance of Alice worst ERM; (b) the ˆIV [Z → Dec(X, Y); D] and ˆIV [Y → Z; D] of Bob’s representation Z. As β increases, Z becomes V-minimal which increases the A…
Figure 5
Figure 5. Figure 5: Intuition behind the construction of g ∈ V in the proof of Theorem 1. The plot schematically represents the all the representations Zy associated with a label y = 1. The representations Z t y (green) are associated with training examples, Z e y (yellow) are associated …
Figure 6
Figure 6. Figure 6: Sweeping over functional families. Each plot shows how the complexity of a functional [PITH_FULL_IMAGE:figures/full_fig_p036_6.png]
Figure 7
Figure 7. Figure 7: Effect of taking multiple inner optimization steps (over [PITH_FULL_IMAGE:figures/full_fig_p037_7.png]
Figure 8
Figure 8. Figure 8: Consequences of not performing the internal optimization to convergence on (a) the average [PITH_FULL_IMAGE:figures/full_fig_p038_8.png]
Figure 9
Figure 9. Figure 9: Effect of number of Monte Carlo Samples on (a) the worst case generalization gap; (b) the [PITH_FULL_IMAGE:figures/full_fig_p039_9.png]
Figure 10
Figure 10. Figure 10: Effect of using Base |Y| expansion (labeled “DIB”) vs. randomly selecting N from the set of Y decompositions of X (labeled “DIB Random”) on (a) the worst case generalization gap; (b) the terms estimated by DIB. All other hyper-parameters are the same as for [PITH_FUL…
Figure 11
Figure 11. Figure 11: (a) Schematic illustration of using the predictor / [PITH_FULL_IMAGE:figures/full_fig_p041_11.png]
Figure 12
Figure 12. Figure 12: Sweeping over γ to find a poorly generalizing ERM. As γ increases, the test performance decreases without having much effect on the training performance, until approximately γ = 0.1. In these experiments, Bob learns representations with either joint ERM (labeled ERM) …
Figure 13
Figure 13. Figure 13: Optimality of V-sufficiency for different hyperparameters. Both plots show the compara￾tive training performance of VBob-sufficient representations for classifiers in VAlice. As in the main text, the log likelihood is scaled to lie in the range [0, . . . , 100] for ea…
Figure 14
Figure 14. Figure 14: Optimality of V-sufficiency in larger functional families. Plots are the same as in [PITH_FULL_IMAGE:figures/full_fig_p044_14.png]
Figure 15
Figure 15. Figure 15: Effect of β on the worst case generalization gap, as well as the terms estimated by DIB (using different Vs for the minimality term) and VIB. (a) Train log likelihood and worst case test log likelihood of the different predictors; (b) Estimated IV [Z → Y] and IV [Z → …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 80 canonical work pages

  1. [1]

    Salton and M

    G. Salton and M. McGill, Introduction to Modern Information Retrieval. McGraw-Hill Book Company, 1984

  2. [2]

    Support-vector networks

    C. Cortes and V . Vapnik, “Support-vector networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995. [Online]. Available: https://doi.org/10.1007/BF00994018

  3. [3]

    Representing and recognizing the visual appearance of materials using three-dimensional textons,

    T. K. Leung and J. Malik, “Representing and recognizing the visual appearance of materials using three-dimensional textons,”Int. J. Comput. Vis., vol. 43, no. 1, pp. 29–44, 2001. [Online]. Available: https://doi.org/10.1023/A:1011126920638

  4. [4]

    Distinctive image features from scale-invariant keypoints,

    D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” Int. J. Comput. Vis., vol. 60, no. 2, pp. 91–110, 2004. [Online]. Available: https://doi.org/10.1023/B: VISI.0000029664.99615.94

  5. [5]

    Random features for large-scale kernel machines,

    A. Rahimi and B. Recht, “Random features for large-scale kernel machines,” in Advances in Neural Information Processing Systems 20, Proceedings of the Twenty-First Annual Conference on Neural Information Processing Systems, Vancouver, British Columbia, Canada, December 3-6, 2007 , J. C. Platt, D. Koller, Y . Singer, and S. T. Roweis, Eds. Curran Associate...

  6. [6]

    Representation Learning: A Review and New Perspectives

    Y . Bengio, A. C. Courville, and P. Vincent, “Representation learning: A review and new perspectives,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 35, no. 8, pp. 1798–1828, 2013. [Online]. Available: https://doi.org/10.1109/TPAMI.2013.50

  7. [7]

    An Overview on Data Representation Learning: From Traditional Feature Learning to Recent Deep Learning

    G. Zhong, L. Wang, and J. Dong, “An overview on data representation learning: From traditional feature learning to recent deep learning,” CoRR, vol. abs/1611.08331, 2016. [Online]. Available: http://arxiv.org/abs/1611.08331

  8. [8]

    Shalev-Shwartz and S

    S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning - From Theory to Algorithms. Cambridge University Press, 2014. [Online]. Available: http://www.cambridge. org/de/academic/subjects/computer-science/pattern-recognition-and-machine-learning/ understanding-machine-learning-theory-algorithms

Show all 121 references
  1. [9]

    Probabilistic supervised learning,

    F. Gressmann, F. J. Király, B. A. Mateen, and H. Oberhauser, “Probabilistic supervised learning,” CoRR, vol. abs/1801.00753, 2018. [Online]. Available: http: //arxiv.org/abs/1801.00753

  2. [10]

    Multiclass classification, information, divergence and surrogate risk,

    J. Duchi, K. Khosravi, F. Ruan et al., “Multiclass classification, information, divergence and surrogate risk,”The Annals of Statistics, vol. 46, no. 6B, pp. 3246–3275, 2018. 10

  3. [11]

    The information bottleneck method,

    N. Tishby, F. C. N. Pereira, and W. Bialek, “The information bottleneck method,”CoRR, vol. physics/0004057, 2000. [Online]. Available: http://arxiv.org/abs/physics/0004057

  4. [12]

    Learning and generalization with the information bottleneck,

    O. Shamir, S. Sabato, and N. Tishby, “Learning and generalization with the information bottleneck,” Theor. Comput. Sci., vol. 411, no. 29-30, pp. 2696–2711, 2010. [Online]. Available: https://doi.org/10.1016/j.tcs.2010.04.006

  5. [13]

    Document clustering using word clusters via the information bottleneck method,

    N. Slonim and N. Tishby, “Document clustering using word clusters via the information bottleneck method,” inSIGIR 2000: Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, July 24-28, 2000, Athens, Greece, E. ...

  6. [14]

    Opening the black box of deep neural networks via information,

    R. Shwartz-Ziv and N. Tishby, “Opening the black box of deep neural networks via information,” CoRR, vol. abs/1703.00810, 2017. [Online]. Available: http: //arxiv.org/abs/1703.00810

  7. [15]

    Information bottleneck methods for distributed learning,

    P. Farajiparvar, A. Beirami, and M. S. Nokleby, “Information bottleneck methods for distributed learning,” in56th Annual Allerton Conference on Communication, Control, and Computing, Allerton 2018, Monticello, IL, USA, October 2-5, 2018. IEEE, 2018, pp. 24–31. [Online]. Availa...

  8. [16]

    Infobot: Transfer and exploration via the information bottleneck,

    A. Goyal, R. Islam, D. Strouse, Z. Ahmed, H. Larochelle, M. Botvinick, Y . Bengio, and S. Levine, “Infobot: Transfer and exploration via the information bottleneck,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. Op...

  9. [17]

    A mathematical theory of communication,

    C. E. Shannon, “A mathematical theory of communication,” Bell Syst. Tech. J., vol. 27, no. 4, pp. 623–656, 1948. [Online]. Available: https://doi.org/10.1002/j.1538-7305.1948.tb00917.x

  10. [18]

    Relevant sparse codes with varia- tional information bottleneck,

    M. Chalk, O. Marre, and G. Tkacik, “Relevant sparse codes with varia- tional information bottleneck,” in Advances in Neural Information Processing Sys- tems 29: Annual Conference on Neural Information Processing Systems 2016, Decem- ber 5-10, 2016, Barcelona, Spain , D. D. Lee...

  11. [19]

    Deep variational information bottleneck,

    A. A. Alemi, I. Fischer, J. V . Dillon, and K. Murphy, “Deep variational information bottleneck,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. [Online]. Available:...

  12. [20]

    Information dropout: Learning optimal representations through noisy computation,

    A. Achille and S. Soatto, “Information dropout: Learning optimal representations through noisy computation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 12, pp. 2897–2905,

  13. [21]

    Nonlinear information bottleneck,

    A. Kolchinsky, B. D. Tracey, and D. H. Wolpert, “Nonlinear information bottleneck,”Entropy, vol. 21, no. 12, p. 1181, 2019. [Online]. Available: https://doi.org/10.3390/e21121181

  14. [22]

    Backpropagation applied to handwritten zip code recognition,

    Y . LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural Computation, vol. 1, no. 4, pp. 541–551, 1989. [Online]. Available: https://doi.org/10.1162/neco.1989.1.4.541

  15. [23]

    A theory of usable information under computational constraints,

    Y . Xu, S. Zhao, J. Song, R. Stewart, and S. Ermon, “A theory of usable information under computational constraints,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net, 2020. [Online]. Available: h...

  16. [24]

    Emergence of invariance and disentanglement in deep representations,

    A. Achille and S. Soatto, “Emergence of invariance and disentanglement in deep representations,” J. Mach. Learn. Res., vol. 19, pp. 50:1–50:34, 2018. [Online]. Available: http://jmlr.org/papers/v19/17-646.html

  17. [25]

    On the mathematical foundations of theoretical statistics,

    R. A. Fisher, “On the mathematical foundations of theoretical statistics,”Philosophical Trans- actions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character, vol. 222, no. 594-604, pp. 309–368, 1922. 11

  18. [26]

    The role of information complexity and randomization in representation learning,

    M. Vera, P. Piantanida, and L. R. Vega, “The role of information complexity and randomization in representation learning,” CoRR, vol. abs/1802.05355, 2018. [Online]. Available: http://arxiv.org/abs/1802.05355

  19. [27]

    The information bottleneck: Connections to other problems, learning and exploration of the ib curve,

    B. Rodriguez Galvez, “The information bottleneck: Connections to other problems, learning and exploration of the ib curve,” 2019

  20. [28]

    Reversible architectures for arbitrarily deep residual neural networks,

    B. Chang, L. Meng, E. Haber, L. Ruthotto, D. Begert, and E. Holtham, “Reversible architectures for arbitrarily deep residual neural networks,” inProceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial...

  21. [29]

    i-revnet: Deep invertible networks,

    J. Jacobsen, A. W. M. Smeulders, and E. Oyallon, “i-revnet: Deep invertible networks,” in6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018. [Online]. Available:...

  22. [30]

    Formal limitations on the measurement of mutual information,

    D. McAllester and K. Stratos, “Formal limitations on the measurement of mutual information,” in The 23rd International Conference on Artificial Intelligence and Statistics, AISTATS 2020, 26-28 August 2020, Online [Palermo, Sicily, Italy], ser. Proceedings of Machine Learning Re...

  23. [31]

    Information bottleneck for gaussian variables,

    G. Chechik, A. Globerson, N. Tishby, and Y . Weiss, “Information bottleneck for gaussian variables,” J. Mach. Learn. Res., vol. 6, pp. 165–188, 2005. [Online]. Available: http://jmlr.org/papers/v6/chechik05a.html

  24. [32]

    Meta-gaussian information bottleneck,

    M. Rey and V . Roth, “Meta-gaussian information bottleneck,” in Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6, 2012, Lake Tahoe, Nevada, United States , P. ...

  25. [33]

    Learning representations for neural network-based classifica- tion using the information bottleneck principle,

    R. A. Amjad and B. C. Geiger, “Learning representations for neural network-based classifica- tion using the information bottleneck principle,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019

  26. [34]

    Strictly proper scoring rules, prediction, and estimation,

    T. Gneiting and A. E. Raftery, “Strictly proper scoring rules, prediction, and estimation,” Journal of the American Statistical Association, vol. 102, no. 477, pp. 359–378, 2007

  27. [35]

    A theory of the learnable,

    L. G. Valiant, “A theory of the learnable,”Commun. ACM, vol. 27, no. 11, pp. 1134–1142,

  28. [36]

    Rademacher and gaussian complexities: Risk bounds and structural results,

    P. L. Bartlett and S. Mendelson, “Rademacher and gaussian complexities: Risk bounds and structural results,” J. Mach. Learn. Res., vol. 3, pp. 463–482, 2002. [Online]. Available: http://jmlr.org/papers/v3/bartlett02a.html

  29. [37]

    Understanding deep learning requires rethinking generalization,

    C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2...

  30. [38]

    On distinguishability criteria for estimating generative models,

    I. J. Goodfellow, “On distinguishability criteria for estimating generative models,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Workshop Track Proceedings, Y . Bengio and Y . LeCun, Eds., 2015. [Online]. Available...

  31. [39]

    Unrolled generative adversarial networks,

    L. Metz, B. Poole, D. Pfau, and J. Sohl-Dickstein, “Unrolled generative adversarial networks,” in5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. [Online]. Available: htt...

  32. [40]

    Approximation by superpositions of a sigmoidal function,

    G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of Control, Signals and Systems, vol. 2, no. 4, pp. 303–314, Dec. 1989. [Online]. Available: https://doi.org/10.1007/BF02551274 12

  33. [41]

    Approximation capabilities of multilayer feedforward networks,

    K. Hornik, “Approximation capabilities of multilayer feedforward networks,” Neural Networks, vol. 4, no. 2, pp. 251–257, 1991. [Online]. Available: https://doi.org/10.1016/ 0893-6080(91)90009-T

  34. [42]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016. IEEE Computer Society, 2016, pp. 770–778. [Online]. Available: https://doi....

  35. [43]

    Learning multiple layers of features from tiny images,

    A. Krizhevsky, G. Hinton et al., “Learning multiple layers of features from tiny images,” 2009

  36. [44]

    Towards explaining the regularization effect of initial large learning rate in training neural networks,

    Y . Li, C. Wei, and T. Ma, “Towards explaining the regularization effect of initial large learning rate in training neural networks,” inAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8-14 Dec...

  37. [45]

    Dropout: a simple way to prevent neural networks from overfitting,

    N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,”J. Mach. Learn. Res., vol. 15, no. 1, pp. 1929–1958, 2014. [Online]. Available: http://dl.acm.org/citation.cfm?id=2670313

  38. [46]

    Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data,

    G. K. Dziugaite and D. M. Roy, “Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data,” inProceedings of the Thirty-Third Conference on Uncertainty in Artificial Intelligence, UAI 2017, Sydney, Australia, A...

  39. [47]

    Stronger generalization bounds for deep nets via a compression approach,

    S. Arora, R. Ge, B. Neyshabur, and Y . Zhang, “Stronger generalization bounds for deep nets via a compression approach,” in Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, ser. Proceedings ...

  40. [48]

    On the importance of single directions for generalization,

    A. S. Morcos, D. G. T. Barrett, N. C. Rabinowitz, and M. Botvinick, “On the importance of single directions for generalization,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings....

  41. [49]

    Predicting the generalization gap in deep networks with margin distributions,

    Y . Jiang, D. Krishnan, H. Mobahi, and S. Bengio, “Predicting the generalization gap in deep networks with margin distributions,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. [Online]. Availa...

  42. [50]

    Fantastic generalization measures and where to find them,

    Y . Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio, “Fantastic generalization measures and where to find them,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net,

  43. [51]

    Sensitivity and generalization in neural networks: an empirical study,

    R. Novak, Y . Bahri, D. A. Abolafia, J. Pennington, and J. Sohl-Dickstein, “Sensitivity and generalization in neural networks: an empirical study,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference ...

  44. [52]

    Train faster, generalize better: Stability of stochastic gradient descent,

    M. Hardt, B. Recht, and Y . Singer, “Train faster, generalize better: Stability of stochastic gradient descent,” in Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, ser. JMLR Workshop and Conference Proc...

  45. [53]

    A bayesian perspective on generalization and stochastic gradient descent,

    S. L. Smith and Q. V . Le, “A bayesian perspective on generalization and stochastic gradient descent,” in6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018. [Onl...

  46. [54]

    Spectrally-normalized margin bounds for neural networks,

    P. L. Bartlett, D. J. Foster, and M. Telgarsky, “Spectrally-normalized margin bounds for neural networks,” in Advances in Neural Information Processing Systems 30: Annual 13 Conference on Neural Information Processing Systems 2017, 4-9 December 2017, Long Beach, CA, USA , I. G...

  47. [55]

    Flat minima,

    S. Hochreiter and J. Schmidhuber, “Flat minima,”Neural Computation, vol. 9, no. 1, pp. 1–42,

  48. [56]

    Entropy-sgd: Biasing gradient descent into wide valleys,

    P. Chaudhari, A. Choromanska, S. Soatto, Y . LeCun, C. Baldassi, C. Borgs, J. T. Chayes, L. Sagun, and R. Zecchina, “Entropy-sgd: Biasing gradient descent into wide valleys,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2...

  49. [57]

    On large-batch training for deep learning: Generalization gap and sharp minima,

    N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Tra...

  50. [58]

    Data-dependent sample complexity of deep neural networks via lipschitz augmentation,

    C. Wei and T. Ma, “Data-dependent sample complexity of deep neural networks via lipschitz augmentation,” inAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8-14 December 2019, Vancou- ver, BC, ...

  51. [59]

    Path-sgd: Path-normalized optimization in deep neural networks,

    B. Neyshabur, R. Salakhutdinov, and N. Srebro, “Path-sgd: Path-normalized optimization in deep neural networks,” in Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Can...

  52. [60]

    A NEW MEASURE OF RANK CORRELATION,

    M. G. KENDALL, “A NEW MEASURE OF RANK CORRELATION,”Biometrika, vol. 30, no. 1-2, pp. 81–93, 06 1938. [Online]. Available: https://doi.org/10.1093/biomet/30.1-2.81

  53. [61]

    Regularizing neural networks by penalizing confident output distributions,

    G. Pereyra, G. Tucker, J. Chorowski, L. Kaiser, and G. E. Hinton, “Regularizing neural networks by penalizing confident output distributions,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Workshop Track Proceedings. ...

  54. [62]

    Norm-based capacity control in neural networks,

    B. Neyshabur, R. Tomioka, and N. Srebro, “Norm-based capacity control in neural networks,” in Proceedings of The 28th Conference on Learning Theory, COLT 2015, Paris, France, July 3-6, 2015, ser. JMLR Workshop and Conference Proceedings, P. Grünwald, E. Hazan, and S. Kale, Eds...

  55. [63]

    The lottery ticket hypothesis: Finding sparse, trainable neural networks,

    J. Frankle and M. Carbin, “The lottery ticket hypothesis: Finding sparse, trainable neural networks,” in7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. [Online]. Available: https://openreview.net/...

  56. [64]

    E. T. Jaynes and R. D. Rosenkrantz, E.T. Jaynes : papers on probability, statistics, and statistical physics. D. Reidel ; Sold and distributed in the U.S.A. and Canada by Kluwer Boston Dordrecht, Holland ; Boston : Hingham, MA, 1983

  57. [65]

    Why least squares and maximum entropy? an axiomatic approach to inference for linear inverse problems,

    I. Csiszar et al., “Why least squares and maximum entropy? an axiomatic approach to inference for linear inverse problems,”The annals of statistics, vol. 19, no. 4, pp. 2032–2066, 1991

  58. [66]

    Information-theoretical optimization techniques,

    F. Topsøe, “Information-theoretical optimization techniques,”Kybernetika, vol. 15, no. 1, pp. 8–27, 1979. [Online]. Available: http://www.kybernetika.cz/content/1979/1/8

  59. [67]

    Walley, Statistical Reasoning with Imprecise Probabilities

    P. Walley, Statistical Reasoning with Imprecise Probabilities. Chapman & Hall, 1991

  60. [68]

    Game theory, maximum entropy, minimum discrepancy and robust bayesian decision theory,

    P. D. Grünwald, A. P. Dawidet al., “Game theory, maximum entropy, minimum discrepancy and robust bayesian decision theory,”the Annals of Statistics, vol. 32, no. 4, pp. 1367–1433, 2004. 14

  61. [69]

    A robust minimax approach to classification,

    G. R. G. Lanckriet, L. E. Ghaoui, C. Bhattacharyya, and M. I. Jordan, “A robust minimax approach to classification,” J. Mach. Learn. Res., vol. 3, pp. 555–582, 2002. [Online]. Available: http://jmlr.org/papers/v3/lanckriet02a.html

  62. [70]

    The minimum information principle for discriminative learning,

    A. Globerson and N. Tishby, “The minimum information principle for discriminative learning,” in Proceedings of the 20th conference on Uncertainty in artificial intelligence, ser. UAI ’04. Banff, Canada: AUAI Press, Jul. 2004, pp. 193–200

  63. [71]

    A minimax approach to supervised learning,

    F. Farnia and D. Tse, “A minimax approach to supervised learning,” inAdvances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, D. D. Lee, M. Sugiyama, U. von Luxburg, I. Guyon, ...

  64. [72]

    Estimating divergence functionals and the likelihood ratio by convex risk minimization,

    X. Nguyen, M. J. Wainwright, and M. I. Jordan, “Estimating divergence functionals and the likelihood ratio by convex risk minimization,” IEEE Trans. Inf. Theory, vol. 56, no. 11, pp. 5847–5861, 2010. [Online]. Available: https://doi.org/10.1109/TIT.2010.2068870

  65. [73]

    Sufficiency and completeness in the general gauss-markov model,

    H. Drygas, “Sufficiency and completeness in the general gauss-markov model,”Sankhy¯a: The Indian Journal of Statistics, Series A (1961-2002), vol. 45, no. 1, pp. 88–98, 1983. [Online]. Available: http://www.jstor.org/stable/25050416

  66. [74]

    Linear transformations preserving best linear unbiased estimators in a general gauss-markoff model,

    J. K. Baksalary and R. Kala, “Linear transformations preserving best linear unbiased estimators in a general gauss-markoff model,”The Annals of Statistics, pp. 913–916, 1981

  67. [75]

    Some notes on linear sufficiency,

    R. Kala, S. Puntanen, and Y . Tian, “Some notes on linear sufficiency,” Statistical Papers, vol. 58, no. 1, pp. 1–17, 2017

  68. [76]

    Minimal achievable sufficient statistic learning,

    M. Cvitkovic and G. Koliander, “Minimal achievable sufficient statistic learning,” in Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, ser. Proceedings of Machine Learning Research, K. Chaudhuri and R....

  69. [77]

    Learning the kernel matrix with semidefinite programming,

    G. R. G. Lanckriet, N. Cristianini, P. L. Bartlett, L. E. Ghaoui, and M. I. Jordan, “Learning the kernel matrix with semidefinite programming,” J. Mach. Learn. Res., vol. 5, pp. 27–72, 2004. [Online]. Available: http://jmlr.org/papers/v5/lanckriet04a.html

  70. [78]

    Multiple kernel learning, conic duality, and the SMO algorithm,

    F. R. Bach, G. R. G. Lanckriet, and M. I. Jordan, “Multiple kernel learning, conic duality, and the SMO algorithm,” in Machine Learning, Proceedings of the Twenty-first International Conference (ICML 2004), Banff, Alberta, Canada, July 4-8, 2004, ser. ACM International Conferen...

  71. [79]

    Large scale multiple kernel learning,

    S. Sonnenburg, G. Rätsch, C. Schäfer, and B. Schölkopf, “Large scale multiple kernel learning,” J. Mach. Learn. Res. , vol. 7, pp. 1531–1565, 2006. [Online]. Available: http://jmlr.org/papers/v7/sonnenburg06a.html

  72. [80]

    Multiple kernel learning algorithms,

    M. Gönen and E. Alpaydin, “Multiple kernel learning algorithms,” J. Mach. Learn. Res., vol. 12, pp. 2211–2268, 2011. [Online]. Available: http://dl.acm.org/citation.cfm?id=2021071

  73. [81]

    Feature space perspectives for learning the kernel,

    C. A. Micchelli and M. Pontil, “Feature space perspectives for learning the kernel,” Mach. Learn., vol. 66, no. 2-3, pp. 297–319, 2007. [Online]. Available: https://doi.org/10.1007/s10994-006-0679-0

  74. [82]

    Feature selection for svms,

    J. Weston, S. Mukherjee, O. Chapelle, M. Pontil, T. A. Poggio, and V . Vapnik, “Feature selection for svms,” inAdvances in Neural Information Processing Systems 13, Papers from Neural Information Processing Systems (NIPS) 2000, Denver, CO, USA, T. K. Leen, T. G. Dietterich, an...

  75. [83]

    Choosing multiple parameters for support vector machines,

    O. Chapelle, V . Vapnik, O. Bousquet, and S. Mukherjee, “Choosing multiple parameters for support vector machines,” Mach. Learn., vol. 46, no. 1-3, pp. 131–159, 2002. [Online]. Available: https://doi.org/10.1023/A:1012450327387

  76. [84]

    Learning bounds for support vector machines with learned kernels,

    N. Srebro and S. Ben-David, “Learning bounds for support vector machines with learned kernels,” in Learning Theory, 19th Annual Conference on Learning Theory, COLT 2006, Pittsburgh, PA, USA, June 22-25, 2006, Proceedings, ser. Lecture Notes in Computer Science, G. Lugosi and H...

  77. [85]

    Generalization bounds for learning kernels,

    C. Cortes, M. Mohri, and A. Rostamizadeh, “Generalization bounds for learning kernels,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10), June 21-24, 2010, Haifa, Israel , J. Fürnkranz and T. Joachims, Eds. Omnipress, 2010, pp. 247–254. [Online...

  78. [86]

    On the convergence rate of lp-norm multiple kernel learning,

    M. Kloft and G. Blanchard, “On the convergence rate of lp-norm multiple kernel learning,” J. Mach. Learn. Res. , vol. 13, pp. 2465–2502, 2012. [Online]. Available: http://dl.acm.org/citation.cfm?id=2503321

  79. [87]

    Learning kernels using local rademacher complexity,

    C. Cortes, M. Kloft, and M. Mohri, “Learning kernels using local rademacher complexity,” in Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe,...

  80. [88]

    Infinite kernel learning: Generalization bounds and algorithms,

    Y . Liu, S. Liao, H. Lin, Y . Yue, and W. Wang, “Infinite kernel learning: Generalization bounds and algorithms,” inProceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, California, USA, S. P. Singh and S. Markovitch, Eds....

  81. [89]

    On mutual information maximization for representation learning,

    M. Tschannen, J. Djolonga, P. K. Rubenstein, S. Gelly, and M. Lucic, “On mutual information maximization for representation learning,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net,

  82. [90]

    Self-organization in a perceptual network,

    R. Linsker, “Self-organization in a perceptual network,”IEEE Computer, vol. 21, no. 3, pp. 105–117, 1988. [Online]. Available: https://doi.org/10.1109/2.36

  83. [91]

    Learning deep representations by mutual information estimation and maximization,

    R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y . Bengio, “Learning deep representations by mutual information estimation and maximization,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, ...

  84. [92]

    Representation learning with contrastive predictive coding,

    A. van den Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” CoRR, vol. abs/1807.03748, 2018. [Online]. Available: http: //arxiv.org/abs/1807.03748

  85. [93]

    On the information bottleneck theory of deep learning,

    A. M. Saxe, Y . Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox, “On the information bottleneck theory of deep learning,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference...

  86. [94]

    Caveats for information bottleneck in deterministic scenarios,

    A. Kolchinsky, B. D. Tracey, and S. V . Kuyk, “Caveats for information bottleneck in deterministic scenarios,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. [Online]. Available: https://openre...

  87. [95]

    Available: https://openreview.net/forum?id=rkxoh24FPH

    [Online]. Available: https://openreview.net/forum?id=rkxoh24FPH

  88. [96]

    Parmigiani, L

    G. Parmigiani, L. Inoue, and H. Lopes, Decision Theory: Principles and Approaches. Wiley Blackwell, Dec. 2010

  89. [97]

    Coherent measures of discrepancy, uncertainty and dependence, with applica- tions to bayesian predictive experimental design,

    A. P. Dawid, “Coherent measures of discrepancy, uncertainty and dependence, with applica- tions to bayesian predictive experimental design,”Department of Statistical Science, University College London. http://www. ucl. ac. uk/Stats/research/abs94. html, Tech. Rep, vol. 139, 1998

  90. [98]

    J. M. Bernardo and A. F. Smith, Bayesian theory. John Wiley & Sons, 2009, vol. 405

  91. [99]

    Adam: A method for stochastic optimization,

    D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Y . Bengio and Y . LeCun, Eds., 2015. [Online]. Available: http://arxiv...

  92. [100]

    Pytorch: An imperative style, high-performance deep learning library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-pe...

  93. [101]

    Deep learning and the information bottleneck principle,

    N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in2015 IEEE Information Theory Workshop, ITW 2015, Jerusalem, Israel, April 26 - May 1, 2015 . IEEE, 2015, pp. 1–5. [Online]. Available: https://doi.org/10.1109/ITW.2015.7133169

  94. [102]

    Unsupervised domain adaptation by backpropagation,

    Y . Ganin and V . S. Lempitsky, “Unsupervised domain adaptation by backpropagation,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, ser. JMLR Workshop and Conference Proceedings, F. R. Bach and D. M. Blei, Eds...

  95. [103]

    Reverse-mode AD in a functional framework: Lambda the ultimate backpropagator,

    B. A. Pearlmutter and J. M. Siskind, “Reverse-mode AD in a functional framework: Lambda the ultimate backpropagator,”ACM Trans. Program. Lang. Syst., vol. 30, no. 2, pp. 7:1–7:36,

  96. [104]

    Gradient-based hyperparameter optimization through reversible learning,

    D. Maclaurin, D. Duvenaud, and R. P. Adams, “Gradient-based hyperparameter optimization through reversible learning,” inProceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, ser. JMLR Workshop and Conference Proceedings...

  97. [105]

    Generalized inner loop meta-learning,

    E. Grefenstette, B. Amos, D. Yarats, P. M. Htut, A. Molchanov, F. Meier, D. Kiela, K. Cho, and S. Chintala, “Generalized inner loop meta-learning,”CoRR, vol. abs/1910.01727, 2019. [Online]. Available: http://arxiv.org/abs/1910.01727

  98. [106]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift,

    S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 , ser. JMLR Workshop and Conference Procee...

  99. [107]

    Envelope theorems for arbitrary choice sets,

    P. Milgrom and I. Segal, “Envelope theorems for arbitrary choice sets,”Econometrica, vol. 70, no. 2, pp. 583–601, 2002

  100. [108]

    SGD on neural networks learns functions of increasing complexity,

    D. Kalimeris, G. Kaplun, P. Nakkiran, B. L. Edelman, T. Yang, B. Barak, and H. Zhang, “SGD on neural networks learns functions of increasing complexity,” in Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, N...

  101. [114]

    Understanding generalization through visualizations,

    W. R. Huang, Z. Emam, M. Goldblum, L. Fowl, J. K. Terry, F. Huang, and T. Goldstein, “Understanding generalization through visualizations,” CoRR, vol. abs/1906.03291, 2019. [Online]. Available: http://arxiv.org/abs/1906.03291

  102. [116]

    As the theorem is about all ERMs, use a proof by contrapositive to only talk about a single ERMf′ that performs optimally on train but not on test

  103. [117]

    Show that ˜N∈ Dec(X,y )

    Construct a random variable ˜N which labels train examples as 1 and test as 0. Show that ˜N∈ Dec(X,y )

  104. [118]

    monotonic biasing ofV

    Using the “monotonic biasing ofV” assumption, constructg∈V fromf′∈V by monoton- ically biasing the predictions towards the “test” label ˜N = 0 s.t.g[z](0) =P˜N(0) for every representationz(c) that are perfectly labelled byf′ (as shown in blue in Fig. 5)

  105. [119]

    5), while predicting as well as P˜N forz(c)

    Show thatg predicts ˜N better than the marginal distributionP˜N for representationsz(w) that are not perfectly labelled byf′ (as shown in orange in Fig. 5), while predicting as well as P˜N forz(c). Conclude thatg predicts ˜N better thanP˜N

  106. [120]

    selector

    Show that the previous point entails IV [ Z→ ˜N ] ⁄= 0. Conclude by Lemma 5 that Z⁄∈MV as desired. C.3.3 Formal Proof Proof. If Z∈M V isV-minimalV-sufficient then by definition it is also V-sufficient, we thus restrict our discussion toV-sufficient representations. As Z isV-suffici...

  107. [121]

    for loop

    Using the monotonocity of V-information in our setting (Lemma 9) we have 0 = IV[Zy→ N] ≤ IV +[Zy→ N]. As V-information is always positive, we conclude that ∀y ∈Y ,∀N∈ Dec(X,y ), IV[Zy→ N] = 0 . By Lemma 5, we conclude that Z isV- minimalV-sufficient as desired. Recoverability U...

  108. [1984]

    Available: https://doi.org/10.1145/1968.1972

    [Online]. Available: https://doi.org/10.1145/1968.1972

  109. [1997]

    Available: https://doi.org/10.1162/neco.1997.9.1.1

    [Online]. Available: https://doi.org/10.1162/neco.1997.9.1.1

  110. [2008]

    Available: https://doi.org/10.1145/1330017.1330018

    [Online]. Available: https://doi.org/10.1145/1330017.1330018

  111. [2017]

    Available: http://auai.org/uai2017/proceedings/papers/173.pdf

    [Online]. Available: http://auai.org/uai2017/proceedings/papers/173.pdf

  112. [2018]

    Available: https://doi.org/10.1109/TPAMI.2017.2784440

    [Online]. Available: https://doi.org/10.1109/TPAMI.2017.2784440

  113. [2020]

    Available: https://openreview.net/forum?id=SJgIPJBFvH

    [Online]. Available: https://openreview.net/forum?id=SJgIPJBFvH

Pith tools

Reviewed August 27, 2026 · model on record in the stance chip above.