REVIEW 2 major objections 3 minor 121 references
Learning Optimal Representations with the Decodable Information Bottleneck
T0 review · 2 major / 3 minor · reviewed 2026-08-27 · deepseek-v4-flash
Pith's one-line read A representation that is V-minimal and V-sufficient for the classifier family V forces every empirical risk minimizer on every dataset to hit the best achievable test risk.
desk verdict The DIB framework and generalization probe are genuinely useful, but the main optimality theorem rests on a monotonic-biasing assumption that fails for multiclass softmax networks, so the theorem does not cover the paper's own experiments as stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is $V$-information, $I_V[Z\to Y] = H[Y] - \inf_{f\in V} \mathbb{E}[-\log f[Z](Y)]$, which measures how much of the label can actually be decoded by predictors in $V$, and its extension to decompositions: for each label $y$, $Dec(X,y)$ is the set of all deterministic labelings of the conditional inputs $X_y$, and $I_V[Z\to Dec(X,Y)]$ averages $I_V[Z_y\to N]$ over those labelings. A representation is $V$-minimal exactly when this quantity is zero, meaning no decoder in $V$ can distinguish same-label examples. The proof of Theorem 1 uses a closure property called monotonic biasing of $V$ to convert a hypothetical ERM that overfits into a decoder that separates training from test examples, contradicting $V$-minimality.
What would settle it
To test the theorem's applicability, train a fixed-width MLP encoder with the DIB objective and then explicitly search for an empirical risk minimizer with positive test risk by optimizing train loss minus a small multiple of test loss; a worst-case ERM whose test loss exceeds the $V$-sufficient lower bound would break the claim. Separately, for a concrete $V$ such as a single hidden layer MLP, take any $f\in V$ and check whether there exists a $g\in V$ that changes $f$ at one representation $z_0$ to an arbitrary probability while preserving all pairwise prediction signs; if no such $g$ exists, the monotonic biasing assumption fails and the proof's construction is unavailable.
Extended reading notes
Core claim
In the paper's own terms, the discovery is a characterization of optimal representations: under finite sample spaces, log loss, and deterministic labels, a representation $Z$ that is $V$-minimal and $V$-sufficient guarantees that any empirical risk minimizer $\hat f \in V$ satisfies $R(\hat f,Z) = \min_Z \min_{f\in V} R(f,Z)$. $V$-sufficiency ensures that some classifier in $V$ achieves the best possible loss, and $V$-minimality ensures that no classifier in $V$ can carve the training set out of the test set, because any classifier that did would carry positive $V$-information about a decomposition of the inputs within a label class. The decodable information bottleneck objective $L_{DIB} = -I_V[Z\to Y] + \beta I_V[Z\to Dec(X,Y)]$ is the Lagrangian relaxation of this characterization, and the paper shows it can be estimated from finite samples with PAC-style guarantees.
Load-bearing premise
The theorem relies on the assumption that the classifier family $V$ is closed under monotonic biasing: from any classifier one can build another classifier in $V$ whose prediction at a single representation point can be set to any probability while preserving the ordering of every other pair of predictions.
Editorial extensions
If this is right
- Training an encoder with DIB for a known downstream family $V$ should remove the need for classifier-side regularization: every ERM in $V$, including adversarially selected worst-case ones, reaches the best achievable test loss.
- The DIB objective is estimable from finite data with PAC-style bounds that grow with the Rademacher complexity of $\log V$, so it becomes easier to estimate when $V$ is small and harder in the universal limit where it recovers the classical information bottleneck objective.
- $V$-minimality of a network's hidden representation correlates with its generalization gap more strongly than output entropy, path norm, gradient variance, and sharpness across 562 trained convolutional networks.
- Penalizing $I_V[Z\to Dec(X,Y)]$ removes decodable information about spurious correlates, such as the overlaid MNIST digits in the paper's CIFAR experiments, so DIB-trained representations are robust to shortcut features.
- When the downstream family is the universal family $U$, $V$-minimal $V$-sufficiency collapses to ordinary minimal sufficiency, recovering the classical information bottleneck as the special case where the decoder is unconstrained.
Reading between the lines
- A natural extension the paper leaves implicit is a quantitative generalization bound for approximate $V$-minimality: since its experiments show the train-test gap shrinking monotonically as $I_V[Z\to Dec(X,Y)]$ decreases, one could try to prove $\text{gap} \le C\, I_V[Z\to Dec(X,Y)]^\alpha$, which would make the framework useful when perfect $V$-minimality is unreachable.
- The monotonic biasing closure assumption suggests the proof techniques transfer most directly to families that are closed under output-bias manipulation, such as softmax classifiers with adjustable final biases; architectures without such closure would need a different proof or a relaxed notion of minimality.
- The framework's logic implies a recipe for transfer learning: measure $V$-minimality of a frozen encoder's representation against the target family, and use it to decide whether a candidate pretrained encoder will generalize before training a downstream head.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the Decodable Information Bottleneck (DIB), a representation-learning objective that replaces Shannon mutual information in the classical IB with V-information relative to a prescribed decoder family V. It defines V-sufficient and V-minimal representations, states an optimality theorem (Theorem 1) claiming that any V-minimal V-sufficient representation makes every empirical risk minimizer in V achieve the best achievable test risk under deterministic labels, proves structural properties of the proposed notions (recoverability, monotonicity, characterization, existence), and gives an estimation bound for the DIB objective. The experimental sections study V-sufficiency and V-minimality on CIFAR variants and MNIST, including a worst-ERM protocol and a large correlation study across 562 trained CNNs.
Significance. If Theorem 1 holds as stated, the paper is a genuinely useful conceptual contribution: it transfers the classical notions of sufficient and minimal statistics from an unconstrained universal family to a pragmatic predictive family V, and it gives a plausible mechanism by which representations, rather than hypothesis classes, can enforce generalization. The proof structure is transparent, the properties in Proposition 2 are nontrivial, and the empirical work is extensive; in particular, the worst-ERM evaluation in Section 4.2 and the 562-model correlation study in Section 4.3 are valuable strengths, and the authors provide code for reproducibility. The significance is currently conditional, however: the main theorem relies on a closure assumption on V whose stated justification for neural networks is not correct, so the theorem's applicability to the paper's own experimental setting is not established as written.
major comments (2)
- [Appendix B.2 and Appendix C.3.3] The 'monotonic biasing of V' assumption is load-bearing for Theorem 1, but the justification given for neural networks is incorrect for multiclass softmax classifiers. The assumption requires that for any f' and any label y, one can construct g by 'modifying the bias term of the final softmax layer' while preserving the sign of every pairwise difference f'[z'](y')-f'[z''](y') for all y'. Adding a bias to one logit changes the normalization of all other classes, so the order of the predicted probabilities for an unmodified label can flip. Concretely, for a 3-class softmax with logits z=(0,0,ln 100) and z'=(ln 10,0,0), adding a bias b to class 0 gives p_1(z)=1/102 and p_1(z')=1/12 at b=0, so p_1(z')>p_1(z), while at b=4, p_1(z)=1/155.6 and p_1(z')=1/548.0, so p_1(z)>p_1(z'); the sign flips. The proof of Theorem 1 uses the assumption to construct the decoder g and also uses it in the multiclass-to-binary reduction, so the theorem as stated does not currently apply to the 10-class CIFAR-10 networks used in Section 4.2. The theorem may be repairable with a weaker single-label order-preservation condition, which does appear to hold for softmax outputs, but the current statement of Assumption B.2 and the proof must be revised.
- [Proposition 7 / Lemma 11] The claimed PAC estimation guarantee for the V-minimality term is not a vanishing bound. Lemma 11 bounds the total estimation error of the V-minimality term by beta*log|Y|, which is independent of the sample size M, and the proof of Proposition 7 carries this constant into the final bound. Consequently, the bound in Eq. (14) does not tend to zero as M grows, so the statement that DIB "can be estimated with guarantees" should be qualified. The authors themselves note in the paragraph after Lemma 11 that the bound is loose and 'is not a PAC-style bound,' which is in tension with the title of Proposition 7 and with the abstract's claim of estimation guarantees.
minor comments (3)
- [Appendix D.1.2] The description of V+-DIB says it uses an MLP with 8192 hidden units '(instead of 8192)'; this should presumably read '(instead of 128)', matching the preceding V- and V entries.
- [Appendix C.4, Lemma 7] In the last line of the proof of Lemma 7, the condition is written as I[Y; Z|Y]=0; this should be I[X; Z|Y]=0 to be consistent with the conclusion Z⊥X|Y.
- [Section 3.3] The main text says LDIB 'inherits V-information's estimation bounds' and refers to Appendix C.5, but the bound for the V-minimality term is a non-vanishing constant; the main text should point the reader to this caveat.
Circularity Check
No circularity: Theorem 1 is derived from the V-minimality definition and external V-information results, not from its conclusions.
full rationale
The paper's central claim, Theorem 1, is a genuine derivation rather than a restatement of its inputs. V-minimality is defined as V-sufficiency plus minimal average V-information with y-decompositions of X, and the proof in Appx. C.3.3 supplies the nontrivial intermediate steps: it assumes a zero-empirical-risk predictor with positive test risk, constructs a nuisance labeling N in Dec(X,y), uses the monotonic-biasing assumption to build a decoder g that beats the marginal predictor on that labeling, and concludes IV[Z->N] > 0, contradicting V-minimality via Lemma 5. This is a substantive implication, not a definitional equivalence. The paper does not fit parameters and then relabel them as predictions; the DIB objective and its PAC bound in Prop. 7 are evaluated against the external V-information framework of Xu et al., with no self-citation chain carrying the argument. The 'monotonic biasing of V' assumption is load-bearing and its stated softmax justification is questionable for multiclass networks, but an unsupported or false assumption is a correctness or soundness risk, not circularity. The experimental generalization probe is an empirical correlation study, and no fitted constant is presented as a derived prediction. Overall the derivation chain is self-contained and not circular.
Assumptions & free parameters
free parameters (3)
- beta (Lagrangian multiplier for DIB) =
tuned per experiment (e.g., beta=10 for CIFAR10 DIB)
- K (number of Monte Carlo y-decompositions) =
4 in most experiments
- gamma (worst-ERM search coefficient) =
0.1
assumptions (8)
- ad hoc to paper Monotonic biasing of V (Appx. B.2)
- domain assumption Deterministic labeling Y = t(X) (Appx. B.3)
- domain assumption Invariance of V to label permutations (Appx. B.2)
- domain assumption Non-empty preimage of labels (Appx. B.2)
- domain assumption Arbitrary constant prediction of V (Appx. B.2)
- standard math Finite sample spaces and |Z| >= |Y| (Appx. B.1)
- domain assumption Log-loss scoring rule (Appx. B.1)
- standard math V-information properties from Xu et al. [23] (non-negativity, monotonicity, estimation bounds)
invented entities (1)
-
y-decomposition of X (Dec(X,y))
Cite this review
Pith. "Pith review of Learning Optimal Representations with the Decodable Information Bottleneck." pith.science (2026). https://pith.science/paper/WZ2NDTKV
@misc{pith2026200912789,
author = {Pith},
title = {Pith review of: Learning Optimal Representations with the Decodable Information Bottleneck},
year = {2026},
howpublished = {\url{https://pith.science/paper/WZ2NDTKV}},
note = {Machine review of arXiv:2009.12789}
}
read the original abstract
We address the question of characterizing and finding optimal representations for supervised learning. Traditionally, this question has been tackled using the Information Bottleneck, which compresses the inputs while retaining information about the targets, in a decoder-agnostic fashion. In machine learning, however, our goal is not compression but rather generalization, which is intimately linked to the predictive family or decoder of interest (e.g. linear classifier). We propose the Decodable Information Bottleneck (DIB) that considers information retention and compression from the perspective of the desired predictive family. As a result, DIB gives rise to representations that are optimal in terms of expected test performance and can be estimated with guarantees. Empirically, we show that the framework can be used to enforce a small generalization gap on downstream classifiers and to predict the generalization ability of neural networks.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Salton and M
G. Salton and M. McGill, Introduction to Modern Information Retrieval. McGraw-Hill Book Company, 1984
1984
-
[2]
C. Cortes and V . Vapnik, “Support-vector networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995. [Online]. Available: https://doi.org/10.1007/BF00994018
-
[3]
Representing and recognizing the visual appearance of materials using three-dimensional textons,
T. K. Leung and J. Malik, “Representing and recognizing the visual appearance of materials using three-dimensional textons,”Int. J. Comput. Vis., vol. 43, no. 1, pp. 29–44, 2001. [Online]. Available: https://doi.org/10.1023/A:1011126920638
-
[4]
Distinctive image features from scale-invariant keypoints,
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” Int. J. Comput. Vis., vol. 60, no. 2, pp. 91–110, 2004. [Online]. Available: https://doi.org/10.1023/B: VISI.0000029664.99615.94
-
[5]
Random features for large-scale kernel machines,
A. Rahimi and B. Recht, “Random features for large-scale kernel machines,” in Advances in Neural Information Processing Systems 20, Proceedings of the Twenty-First Annual Conference on Neural Information Processing Systems, Vancouver, British Columbia, Canada, December 3-6, 2007 , J. C. Platt, D. Koller, Y . Singer, and S. T. Roweis, Eds. Curran Associate...
2007
-
[6]
Representation Learning: A Review and New Perspectives
Y . Bengio, A. C. Courville, and P. Vincent, “Representation learning: A review and new perspectives,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 35, no. 8, pp. 1798–1828, 2013. [Online]. Available: https://doi.org/10.1109/TPAMI.2013.50
-
[7]
G. Zhong, L. Wang, and J. Dong, “An overview on data representation learning: From traditional feature learning to recent deep learning,” CoRR, vol. abs/1611.08331, 2016. [Online]. Available: http://arxiv.org/abs/1611.08331
work page Pith review arXiv 2016
-
[8]
Shalev-Shwartz and S
S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning - From Theory to Algorithms. Cambridge University Press, 2014. [Online]. Available: http://www.cambridge. org/de/academic/subjects/computer-science/pattern-recognition-and-machine-learning/ understanding-machine-learning-theory-algorithms
2014
Show all 121 references
-
[9]
Probabilistic supervised learning,
F. Gressmann, F. J. Király, B. A. Mateen, and H. Oberhauser, “Probabilistic supervised learning,” CoRR, vol. abs/1801.00753, 2018. [Online]. Available: http: //arxiv.org/abs/1801.00753
2018 arXiv
-
[10]
Multiclass classification, information, divergence and surrogate risk,
J. Duchi, K. Khosravi, F. Ruan et al., “Multiclass classification, information, divergence and surrogate risk,”The Annals of Statistics, vol. 46, no. 6B, pp. 3246–3275, 2018. 10
2018
-
[11]
The information bottleneck method,
N. Tishby, F. C. N. Pereira, and W. Bialek, “The information bottleneck method,”CoRR, vol. physics/0004057, 2000. [Online]. Available: http://arxiv.org/abs/physics/0004057
2000 arXiv
-
[12]
Learning and generalization with the information bottleneck,
O. Shamir, S. Sabato, and N. Tishby, “Learning and generalization with the information bottleneck,” Theor. Comput. Sci., vol. 411, no. 29-30, pp. 2696–2711, 2010. [Online]. Available: https://doi.org/10.1016/j.tcs.2010.04.006
2010 doi
-
[13]
Document clustering using word clusters via the information bottleneck method,
N. Slonim and N. Tishby, “Document clustering using word clusters via the information bottleneck method,” inSIGIR 2000: Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, July 24-28, 2000, Athens, Greece, E. ...
2000
-
[14]
Opening the black box of deep neural networks via information,
R. Shwartz-Ziv and N. Tishby, “Opening the black box of deep neural networks via information,” CoRR, vol. abs/1703.00810, 2017. [Online]. Available: http: //arxiv.org/abs/1703.00810
2017 arXiv
-
[15]
Information bottleneck methods for distributed learning,
P. Farajiparvar, A. Beirami, and M. S. Nokleby, “Information bottleneck methods for distributed learning,” in56th Annual Allerton Conference on Communication, Control, and Computing, Allerton 2018, Monticello, IL, USA, October 2-5, 2018. IEEE, 2018, pp. 24–31. [Online]. Availa...
2018
-
[16]
Infobot: Transfer and exploration via the information bottleneck,
A. Goyal, R. Islam, D. Strouse, Z. Ahmed, H. Larochelle, M. Botvinick, Y . Bengio, and S. Levine, “Infobot: Transfer and exploration via the information bottleneck,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. Op...
2019
-
[17]
A mathematical theory of communication,
C. E. Shannon, “A mathematical theory of communication,” Bell Syst. Tech. J., vol. 27, no. 4, pp. 623–656, 1948. [Online]. Available: https://doi.org/10.1002/j.1538-7305.1948.tb00917.x
1948
-
[18]
Relevant sparse codes with varia- tional information bottleneck,
M. Chalk, O. Marre, and G. Tkacik, “Relevant sparse codes with varia- tional information bottleneck,” in Advances in Neural Information Processing Sys- tems 29: Annual Conference on Neural Information Processing Systems 2016, Decem- ber 5-10, 2016, Barcelona, Spain , D. D. Lee...
2016
-
[19]
Deep variational information bottleneck,
A. A. Alemi, I. Fischer, J. V . Dillon, and K. Murphy, “Deep variational information bottleneck,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. [Online]. Available:...
2017
-
[20]
Information dropout: Learning optimal representations through noisy computation,
A. Achille and S. Soatto, “Information dropout: Learning optimal representations through noisy computation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 12, pp. 2897–2905,
-
[21]
Nonlinear information bottleneck,
A. Kolchinsky, B. D. Tracey, and D. H. Wolpert, “Nonlinear information bottleneck,”Entropy, vol. 21, no. 12, p. 1181, 2019. [Online]. Available: https://doi.org/10.3390/e21121181
2019 doi
-
[22]
Backpropagation applied to handwritten zip code recognition,
Y . LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural Computation, vol. 1, no. 4, pp. 541–551, 1989. [Online]. Available: https://doi.org/10.1162/neco.1989.1.4.541
1989 doi
-
[23]
A theory of usable information under computational constraints,
Y . Xu, S. Zhao, J. Song, R. Stewart, and S. Ermon, “A theory of usable information under computational constraints,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net, 2020. [Online]. Available: h...
2020
-
[24]
Emergence of invariance and disentanglement in deep representations,
A. Achille and S. Soatto, “Emergence of invariance and disentanglement in deep representations,” J. Mach. Learn. Res., vol. 19, pp. 50:1–50:34, 2018. [Online]. Available: http://jmlr.org/papers/v19/17-646.html
2018
-
[25]
On the mathematical foundations of theoretical statistics,
R. A. Fisher, “On the mathematical foundations of theoretical statistics,”Philosophical Trans- actions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character, vol. 222, no. 594-604, pp. 309–368, 1922. 11
1922
-
[26]
The role of information complexity and randomization in representation learning,
M. Vera, P. Piantanida, and L. R. Vega, “The role of information complexity and randomization in representation learning,” CoRR, vol. abs/1802.05355, 2018. [Online]. Available: http://arxiv.org/abs/1802.05355
2018 arXiv
-
[27]
The information bottleneck: Connections to other problems, learning and exploration of the ib curve,
B. Rodriguez Galvez, “The information bottleneck: Connections to other problems, learning and exploration of the ib curve,” 2019
2019
-
[28]
Reversible architectures for arbitrarily deep residual neural networks,
B. Chang, L. Meng, E. Haber, L. Ruthotto, D. Begert, and E. Holtham, “Reversible architectures for arbitrarily deep residual neural networks,” inProceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial...
2018
-
[29]
i-revnet: Deep invertible networks,
J. Jacobsen, A. W. M. Smeulders, and E. Oyallon, “i-revnet: Deep invertible networks,” in6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018. [Online]. Available:...
2018
-
[30]
Formal limitations on the measurement of mutual information,
D. McAllester and K. Stratos, “Formal limitations on the measurement of mutual information,” in The 23rd International Conference on Artificial Intelligence and Statistics, AISTATS 2020, 26-28 August 2020, Online [Palermo, Sicily, Italy], ser. Proceedings of Machine Learning Re...
2020
-
[31]
Information bottleneck for gaussian variables,
G. Chechik, A. Globerson, N. Tishby, and Y . Weiss, “Information bottleneck for gaussian variables,” J. Mach. Learn. Res., vol. 6, pp. 165–188, 2005. [Online]. Available: http://jmlr.org/papers/v6/chechik05a.html
2005
-
[32]
Meta-gaussian information bottleneck,
M. Rey and V . Roth, “Meta-gaussian information bottleneck,” in Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6, 2012, Lake Tahoe, Nevada, United States , P. ...
2012
-
[33]
Learning representations for neural network-based classifica- tion using the information bottleneck principle,
R. A. Amjad and B. C. Geiger, “Learning representations for neural network-based classifica- tion using the information bottleneck principle,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019
2019
-
[34]
Strictly proper scoring rules, prediction, and estimation,
T. Gneiting and A. E. Raftery, “Strictly proper scoring rules, prediction, and estimation,” Journal of the American Statistical Association, vol. 102, no. 477, pp. 359–378, 2007
2007
-
[35]
A theory of the learnable,
L. G. Valiant, “A theory of the learnable,”Commun. ACM, vol. 27, no. 11, pp. 1134–1142,
-
[36]
Rademacher and gaussian complexities: Risk bounds and structural results,
P. L. Bartlett and S. Mendelson, “Rademacher and gaussian complexities: Risk bounds and structural results,” J. Mach. Learn. Res., vol. 3, pp. 463–482, 2002. [Online]. Available: http://jmlr.org/papers/v3/bartlett02a.html
2002
-
[37]
Understanding deep learning requires rethinking generalization,
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2...
2017
-
[38]
On distinguishability criteria for estimating generative models,
I. J. Goodfellow, “On distinguishability criteria for estimating generative models,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Workshop Track Proceedings, Y . Bengio and Y . LeCun, Eds., 2015. [Online]. Available...
2015 arXiv
-
[39]
Unrolled generative adversarial networks,
L. Metz, B. Poole, D. Pfau, and J. Sohl-Dickstein, “Unrolled generative adversarial networks,” in5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. [Online]. Available: htt...
2017
-
[40]
Approximation by superpositions of a sigmoidal function,
G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of Control, Signals and Systems, vol. 2, no. 4, pp. 303–314, Dec. 1989. [Online]. Available: https://doi.org/10.1007/BF02551274 12
1989 doi
-
[41]
Approximation capabilities of multilayer feedforward networks,
K. Hornik, “Approximation capabilities of multilayer feedforward networks,” Neural Networks, vol. 4, no. 2, pp. 251–257, 1991. [Online]. Available: https://doi.org/10.1016/ 0893-6080(91)90009-T
1991
-
[42]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016. IEEE Computer Society, 2016, pp. 770–778. [Online]. Available: https://doi....
2016 doi
-
[43]
Learning multiple layers of features from tiny images,
A. Krizhevsky, G. Hinton et al., “Learning multiple layers of features from tiny images,” 2009
2009
-
[44]
Towards explaining the regularization effect of initial large learning rate in training neural networks,
Y . Li, C. Wei, and T. Ma, “Towards explaining the regularization effect of initial large learning rate in training neural networks,” inAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8-14 Dec...
2019
-
[45]
Dropout: a simple way to prevent neural networks from overfitting,
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,”J. Mach. Learn. Res., vol. 15, no. 1, pp. 1929–1958, 2014. [Online]. Available: http://dl.acm.org/citation.cfm?id=2670313
1929
-
[46]
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data,
G. K. Dziugaite and D. M. Roy, “Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data,” inProceedings of the Thirty-Third Conference on Uncertainty in Artificial Intelligence, UAI 2017, Sydney, Australia, A...
2017
-
[47]
Stronger generalization bounds for deep nets via a compression approach,
S. Arora, R. Ge, B. Neyshabur, and Y . Zhang, “Stronger generalization bounds for deep nets via a compression approach,” in Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, ser. Proceedings ...
2018
-
[48]
On the importance of single directions for generalization,
A. S. Morcos, D. G. T. Barrett, N. C. Rabinowitz, and M. Botvinick, “On the importance of single directions for generalization,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings....
2018
-
[49]
Predicting the generalization gap in deep networks with margin distributions,
Y . Jiang, D. Krishnan, H. Mobahi, and S. Bengio, “Predicting the generalization gap in deep networks with margin distributions,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. [Online]. Availa...
2019
-
[50]
Fantastic generalization measures and where to find them,
Y . Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio, “Fantastic generalization measures and where to find them,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net,
2020
-
[51]
Sensitivity and generalization in neural networks: an empirical study,
R. Novak, Y . Bahri, D. A. Abolafia, J. Pennington, and J. Sohl-Dickstein, “Sensitivity and generalization in neural networks: an empirical study,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference ...
2018
-
[52]
Train faster, generalize better: Stability of stochastic gradient descent,
M. Hardt, B. Recht, and Y . Singer, “Train faster, generalize better: Stability of stochastic gradient descent,” in Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, ser. JMLR Workshop and Conference Proc...
2016
-
[53]
A bayesian perspective on generalization and stochastic gradient descent,
S. L. Smith and Q. V . Le, “A bayesian perspective on generalization and stochastic gradient descent,” in6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018. [Onl...
2018
-
[54]
Spectrally-normalized margin bounds for neural networks,
P. L. Bartlett, D. J. Foster, and M. Telgarsky, “Spectrally-normalized margin bounds for neural networks,” in Advances in Neural Information Processing Systems 30: Annual 13 Conference on Neural Information Processing Systems 2017, 4-9 December 2017, Long Beach, CA, USA , I. G...
2017
-
[55]
Flat minima,
S. Hochreiter and J. Schmidhuber, “Flat minima,”Neural Computation, vol. 9, no. 1, pp. 1–42,
-
[56]
Entropy-sgd: Biasing gradient descent into wide valleys,
P. Chaudhari, A. Choromanska, S. Soatto, Y . LeCun, C. Baldassi, C. Borgs, J. T. Chayes, L. Sagun, and R. Zecchina, “Entropy-sgd: Biasing gradient descent into wide valleys,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2...
2017
-
[57]
On large-batch training for deep learning: Generalization gap and sharp minima,
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Tra...
2017
-
[58]
Data-dependent sample complexity of deep neural networks via lipschitz augmentation,
C. Wei and T. Ma, “Data-dependent sample complexity of deep neural networks via lipschitz augmentation,” inAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8-14 December 2019, Vancou- ver, BC, ...
2019
-
[59]
Path-sgd: Path-normalized optimization in deep neural networks,
B. Neyshabur, R. Salakhutdinov, and N. Srebro, “Path-sgd: Path-normalized optimization in deep neural networks,” in Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Can...
2015
-
[60]
A NEW MEASURE OF RANK CORRELATION,
M. G. KENDALL, “A NEW MEASURE OF RANK CORRELATION,”Biometrika, vol. 30, no. 1-2, pp. 81–93, 06 1938. [Online]. Available: https://doi.org/10.1093/biomet/30.1-2.81
1938 doi
-
[61]
Regularizing neural networks by penalizing confident output distributions,
G. Pereyra, G. Tucker, J. Chorowski, L. Kaiser, and G. E. Hinton, “Regularizing neural networks by penalizing confident output distributions,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Workshop Track Proceedings. ...
2017
-
[62]
Norm-based capacity control in neural networks,
B. Neyshabur, R. Tomioka, and N. Srebro, “Norm-based capacity control in neural networks,” in Proceedings of The 28th Conference on Learning Theory, COLT 2015, Paris, France, July 3-6, 2015, ser. JMLR Workshop and Conference Proceedings, P. Grünwald, E. Hazan, and S. Kale, Eds...
2015
-
[63]
The lottery ticket hypothesis: Finding sparse, trainable neural networks,
J. Frankle and M. Carbin, “The lottery ticket hypothesis: Finding sparse, trainable neural networks,” in7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. [Online]. Available: https://openreview.net/...
2019
-
[64]
E. T. Jaynes and R. D. Rosenkrantz, E.T. Jaynes : papers on probability, statistics, and statistical physics. D. Reidel ; Sold and distributed in the U.S.A. and Canada by Kluwer Boston Dordrecht, Holland ; Boston : Hingham, MA, 1983
1983
-
[65]
Why least squares and maximum entropy? an axiomatic approach to inference for linear inverse problems,
I. Csiszar et al., “Why least squares and maximum entropy? an axiomatic approach to inference for linear inverse problems,”The annals of statistics, vol. 19, no. 4, pp. 2032–2066, 1991
1991
-
[66]
Information-theoretical optimization techniques,
F. Topsøe, “Information-theoretical optimization techniques,”Kybernetika, vol. 15, no. 1, pp. 8–27, 1979. [Online]. Available: http://www.kybernetika.cz/content/1979/1/8
1979
-
[67]
Walley, Statistical Reasoning with Imprecise Probabilities
P. Walley, Statistical Reasoning with Imprecise Probabilities. Chapman & Hall, 1991
1991
-
[68]
Game theory, maximum entropy, minimum discrepancy and robust bayesian decision theory,
P. D. Grünwald, A. P. Dawidet al., “Game theory, maximum entropy, minimum discrepancy and robust bayesian decision theory,”the Annals of Statistics, vol. 32, no. 4, pp. 1367–1433, 2004. 14
2004
-
[69]
A robust minimax approach to classification,
G. R. G. Lanckriet, L. E. Ghaoui, C. Bhattacharyya, and M. I. Jordan, “A robust minimax approach to classification,” J. Mach. Learn. Res., vol. 3, pp. 555–582, 2002. [Online]. Available: http://jmlr.org/papers/v3/lanckriet02a.html
2002
-
[70]
The minimum information principle for discriminative learning,
A. Globerson and N. Tishby, “The minimum information principle for discriminative learning,” in Proceedings of the 20th conference on Uncertainty in artificial intelligence, ser. UAI ’04. Banff, Canada: AUAI Press, Jul. 2004, pp. 193–200
2004
-
[71]
A minimax approach to supervised learning,
F. Farnia and D. Tse, “A minimax approach to supervised learning,” inAdvances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, D. D. Lee, M. Sugiyama, U. von Luxburg, I. Guyon, ...
2016
-
[72]
Estimating divergence functionals and the likelihood ratio by convex risk minimization,
X. Nguyen, M. J. Wainwright, and M. I. Jordan, “Estimating divergence functionals and the likelihood ratio by convex risk minimization,” IEEE Trans. Inf. Theory, vol. 56, no. 11, pp. 5847–5861, 2010. [Online]. Available: https://doi.org/10.1109/TIT.2010.2068870
2010
-
[73]
Sufficiency and completeness in the general gauss-markov model,
H. Drygas, “Sufficiency and completeness in the general gauss-markov model,”Sankhy¯a: The Indian Journal of Statistics, Series A (1961-2002), vol. 45, no. 1, pp. 88–98, 1983. [Online]. Available: http://www.jstor.org/stable/25050416
1961
-
[74]
Linear transformations preserving best linear unbiased estimators in a general gauss-markoff model,
J. K. Baksalary and R. Kala, “Linear transformations preserving best linear unbiased estimators in a general gauss-markoff model,”The Annals of Statistics, pp. 913–916, 1981
1981
-
[75]
Some notes on linear sufficiency,
R. Kala, S. Puntanen, and Y . Tian, “Some notes on linear sufficiency,” Statistical Papers, vol. 58, no. 1, pp. 1–17, 2017
2017
-
[76]
Minimal achievable sufficient statistic learning,
M. Cvitkovic and G. Koliander, “Minimal achievable sufficient statistic learning,” in Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, ser. Proceedings of Machine Learning Research, K. Chaudhuri and R....
2019
-
[77]
Learning the kernel matrix with semidefinite programming,
G. R. G. Lanckriet, N. Cristianini, P. L. Bartlett, L. E. Ghaoui, and M. I. Jordan, “Learning the kernel matrix with semidefinite programming,” J. Mach. Learn. Res., vol. 5, pp. 27–72, 2004. [Online]. Available: http://jmlr.org/papers/v5/lanckriet04a.html
2004
-
[78]
Multiple kernel learning, conic duality, and the SMO algorithm,
F. R. Bach, G. R. G. Lanckriet, and M. I. Jordan, “Multiple kernel learning, conic duality, and the SMO algorithm,” in Machine Learning, Proceedings of the Twenty-first International Conference (ICML 2004), Banff, Alberta, Canada, July 4-8, 2004, ser. ACM International Conferen...
2004
-
[79]
Large scale multiple kernel learning,
S. Sonnenburg, G. Rätsch, C. Schäfer, and B. Schölkopf, “Large scale multiple kernel learning,” J. Mach. Learn. Res. , vol. 7, pp. 1531–1565, 2006. [Online]. Available: http://jmlr.org/papers/v7/sonnenburg06a.html
2006
-
[80]
Multiple kernel learning algorithms,
M. Gönen and E. Alpaydin, “Multiple kernel learning algorithms,” J. Mach. Learn. Res., vol. 12, pp. 2211–2268, 2011. [Online]. Available: http://dl.acm.org/citation.cfm?id=2021071
2011
-
[81]
Feature space perspectives for learning the kernel,
C. A. Micchelli and M. Pontil, “Feature space perspectives for learning the kernel,” Mach. Learn., vol. 66, no. 2-3, pp. 297–319, 2007. [Online]. Available: https://doi.org/10.1007/s10994-006-0679-0
2007 doi
-
[82]
Feature selection for svms,
J. Weston, S. Mukherjee, O. Chapelle, M. Pontil, T. A. Poggio, and V . Vapnik, “Feature selection for svms,” inAdvances in Neural Information Processing Systems 13, Papers from Neural Information Processing Systems (NIPS) 2000, Denver, CO, USA, T. K. Leen, T. G. Dietterich, an...
2000
-
[83]
Choosing multiple parameters for support vector machines,
O. Chapelle, V . Vapnik, O. Bousquet, and S. Mukherjee, “Choosing multiple parameters for support vector machines,” Mach. Learn., vol. 46, no. 1-3, pp. 131–159, 2002. [Online]. Available: https://doi.org/10.1023/A:1012450327387
2002 doi
-
[84]
Learning bounds for support vector machines with learned kernels,
N. Srebro and S. Ben-David, “Learning bounds for support vector machines with learned kernels,” in Learning Theory, 19th Annual Conference on Learning Theory, COLT 2006, Pittsburgh, PA, USA, June 22-25, 2006, Proceedings, ser. Lecture Notes in Computer Science, G. Lugosi and H...
2006 doi
-
[85]
Generalization bounds for learning kernels,
C. Cortes, M. Mohri, and A. Rostamizadeh, “Generalization bounds for learning kernels,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10), June 21-24, 2010, Haifa, Israel , J. Fürnkranz and T. Joachims, Eds. Omnipress, 2010, pp. 247–254. [Online...
2010
-
[86]
On the convergence rate of lp-norm multiple kernel learning,
M. Kloft and G. Blanchard, “On the convergence rate of lp-norm multiple kernel learning,” J. Mach. Learn. Res. , vol. 13, pp. 2465–2502, 2012. [Online]. Available: http://dl.acm.org/citation.cfm?id=2503321
2012
-
[87]
Learning kernels using local rademacher complexity,
C. Cortes, M. Kloft, and M. Mohri, “Learning kernels using local rademacher complexity,” in Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe,...
2013
-
[88]
Infinite kernel learning: Generalization bounds and algorithms,
Y . Liu, S. Liao, H. Lin, Y . Yue, and W. Wang, “Infinite kernel learning: Generalization bounds and algorithms,” inProceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, California, USA, S. P. Singh and S. Markovitch, Eds....
2017
-
[89]
On mutual information maximization for representation learning,
M. Tschannen, J. Djolonga, P. K. Rubenstein, S. Gelly, and M. Lucic, “On mutual information maximization for representation learning,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net,
2020
-
[90]
Self-organization in a perceptual network,
R. Linsker, “Self-organization in a perceptual network,”IEEE Computer, vol. 21, no. 3, pp. 105–117, 1988. [Online]. Available: https://doi.org/10.1109/2.36
1988 doi
-
[91]
Learning deep representations by mutual information estimation and maximization,
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y . Bengio, “Learning deep representations by mutual information estimation and maximization,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, ...
2019
-
[92]
Representation learning with contrastive predictive coding,
A. van den Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” CoRR, vol. abs/1807.03748, 2018. [Online]. Available: http: //arxiv.org/abs/1807.03748
2018 arXiv
-
[93]
On the information bottleneck theory of deep learning,
A. M. Saxe, Y . Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox, “On the information bottleneck theory of deep learning,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference...
2018
-
[94]
Caveats for information bottleneck in deterministic scenarios,
A. Kolchinsky, B. D. Tracey, and S. V . Kuyk, “Caveats for information bottleneck in deterministic scenarios,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. [Online]. Available: https://openre...
2019
-
[95]
Available: https://openreview.net/forum?id=rkxoh24FPH
[Online]. Available: https://openreview.net/forum?id=rkxoh24FPH
-
[96]
Parmigiani, L
G. Parmigiani, L. Inoue, and H. Lopes, Decision Theory: Principles and Approaches. Wiley Blackwell, Dec. 2010
2010
-
[97]
Coherent measures of discrepancy, uncertainty and dependence, with applica- tions to bayesian predictive experimental design,
A. P. Dawid, “Coherent measures of discrepancy, uncertainty and dependence, with applica- tions to bayesian predictive experimental design,”Department of Statistical Science, University College London. http://www. ucl. ac. uk/Stats/research/abs94. html, Tech. Rep, vol. 139, 1998
1998
-
[98]
J. M. Bernardo and A. F. Smith, Bayesian theory. John Wiley & Sons, 2009, vol. 405
2009
-
[99]
Adam: A method for stochastic optimization,
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Y . Bengio and Y . LeCun, Eds., 2015. [Online]. Available: http://arxiv...
2015 arXiv
-
[100]
Pytorch: An imperative style, high-performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-pe...
2019
-
[101]
Deep learning and the information bottleneck principle,
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in2015 IEEE Information Theory Workshop, ITW 2015, Jerusalem, Israel, April 26 - May 1, 2015 . IEEE, 2015, pp. 1–5. [Online]. Available: https://doi.org/10.1109/ITW.2015.7133169
2015
-
[102]
Unsupervised domain adaptation by backpropagation,
Y . Ganin and V . S. Lempitsky, “Unsupervised domain adaptation by backpropagation,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, ser. JMLR Workshop and Conference Proceedings, F. R. Bach and D. M. Blei, Eds...
2015
-
[103]
Reverse-mode AD in a functional framework: Lambda the ultimate backpropagator,
B. A. Pearlmutter and J. M. Siskind, “Reverse-mode AD in a functional framework: Lambda the ultimate backpropagator,”ACM Trans. Program. Lang. Syst., vol. 30, no. 2, pp. 7:1–7:36,
-
[104]
Gradient-based hyperparameter optimization through reversible learning,
D. Maclaurin, D. Duvenaud, and R. P. Adams, “Gradient-based hyperparameter optimization through reversible learning,” inProceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, ser. JMLR Workshop and Conference Proceedings...
2015
-
[105]
Generalized inner loop meta-learning,
E. Grefenstette, B. Amos, D. Yarats, P. M. Htut, A. Molchanov, F. Meier, D. Kiela, K. Cho, and S. Chintala, “Generalized inner loop meta-learning,”CoRR, vol. abs/1910.01727, 2019. [Online]. Available: http://arxiv.org/abs/1910.01727
1910 arXiv
-
[106]
Batch normalization: Accelerating deep network training by reducing internal covariate shift,
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 , ser. JMLR Workshop and Conference Procee...
2015
-
[107]
Envelope theorems for arbitrary choice sets,
P. Milgrom and I. Segal, “Envelope theorems for arbitrary choice sets,”Econometrica, vol. 70, no. 2, pp. 583–601, 2002
2002
-
[108]
SGD on neural networks learns functions of increasing complexity,
D. Kalimeris, G. Kaplun, P. Nakkiran, B. L. Edelman, T. Yang, B. Barak, and H. Zhang, “SGD on neural networks learns functions of increasing complexity,” in Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, N...
2019
-
[114]
Understanding generalization through visualizations,
W. R. Huang, Z. Emam, M. Goldblum, L. Fowl, J. K. Terry, F. Huang, and T. Goldstein, “Understanding generalization through visualizations,” CoRR, vol. abs/1906.03291, 2019. [Online]. Available: http://arxiv.org/abs/1906.03291
1906 arXiv
-
[116]
As the theorem is about all ERMs, use a proof by contrapositive to only talk about a single ERMf′ that performs optimally on train but not on test
-
[117]
Show that ˜N∈ Dec(X,y )
Construct a random variable ˜N which labels train examples as 1 and test as 0. Show that ˜N∈ Dec(X,y )
-
[118]
monotonic biasing ofV
Using the “monotonic biasing ofV” assumption, constructg∈V fromf′∈V by monoton- ically biasing the predictions towards the “test” label ˜N = 0 s.t.g[z](0) =P˜N(0) for every representationz(c) that are perfectly labelled byf′ (as shown in blue in Fig. 5)
-
[119]
5), while predicting as well as P˜N forz(c)
Show thatg predicts ˜N better than the marginal distributionP˜N for representationsz(w) that are not perfectly labelled byf′ (as shown in orange in Fig. 5), while predicting as well as P˜N forz(c). Conclude thatg predicts ˜N better thanP˜N
-
[120]
selector
Show that the previous point entails IV [ Z→ ˜N ] ⁄= 0. Conclude by Lemma 5 that Z⁄∈MV as desired. C.3.3 Formal Proof Proof. If Z∈M V isV-minimalV-sufficient then by definition it is also V-sufficient, we thus restrict our discussion toV-sufficient representations. As Z isV-suffici...
-
[121]
for loop
Using the monotonocity of V-information in our setting (Lemma 9) we have 0 = IV[Zy→ N] ≤ IV +[Zy→ N]. As V-information is always positive, we conclude that ∀y ∈Y ,∀N∈ Dec(X,y ), IV[Zy→ N] = 0 . By Lemma 5, we conclude that Z isV- minimalV-sufficient as desired. Recoverability U...
-
[1984]
Available: https://doi.org/10.1145/1968.1972
[Online]. Available: https://doi.org/10.1145/1968.1972
1968
-
[1997]
Available: https://doi.org/10.1162/neco.1997.9.1.1
[Online]. Available: https://doi.org/10.1162/neco.1997.9.1.1
1997 doi
-
[2008]
Available: https://doi.org/10.1145/1330017.1330018
[Online]. Available: https://doi.org/10.1145/1330017.1330018
-
[2017]
Available: http://auai.org/uai2017/proceedings/papers/173.pdf
[Online]. Available: http://auai.org/uai2017/proceedings/papers/173.pdf
-
[2018]
Available: https://doi.org/10.1109/TPAMI.2017.2784440
[Online]. Available: https://doi.org/10.1109/TPAMI.2017.2784440
2017
-
[2020]
Available: https://openreview.net/forum?id=SJgIPJBFvH
[Online]. Available: https://openreview.net/forum?id=SJgIPJBFvH
Reviewed August 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.