Pith. sign in

REVIEW 2 cited by

Learning Representations for Neural Network-Based Classification Using the Information Bottleneck Principle

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.09766 v6 pith:54VLWV52 submitted 2018-02-27 cs.LG cs.CVcs.ITmath.IT

classification cs.LGcs.CVcs.ITmath.IT
keywords dnnsfunctionalclassificationoptimizationbottleneckframeworkinformationissues
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this theory paper, we investigate training deep neural networks (DNNs) for classification via minimizing the information bottleneck (IB) functional. We show that the resulting optimization problem suffers from two severe issues: First, for deterministic DNNs, either the IB functional is infinite for almost all values of network parameters, making the optimization problem ill-posed, or it is piecewise constant, hence not admitting gradient-based optimization methods. Second, the invariance of the IB functional under bijections prevents it from capturing properties of the learned representation that are desirable for classification, such as robustness and simplicity. We argue that these issues are partly resolved for stochastic DNNs, DNNs that include a (hard or soft) decision rule, or by replacing the IB functional with related, but more well-behaved cost functions. We conclude that recent successes reported about training DNNs using the IB framework must be attributed to such solutions. As a side effect, our results indicate limitations of the IB framework for the analysis of DNNs. We also note that rather than trying to repair the inherent problems in the IB functional, a better approach may be to design regularizers on latent representation enforcing the desired properties directly.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feature selection of neural networks is skewed towards the less abstract cue

    cs.CV 2019-08 conditional novelty 6.0 of 10

    Fully-connected networks trained on images with two equally valid class cues systematically learn the low-level statistical pattern and ignore the symbolic code, unless the pattern is partially corrupted.

  2. The HSIC Bottleneck: Deep Learning without Back-Propagation

    cs.LG 2019-08 conditional novelty 6.0 of 10

    An HSIC-based information-bottleneck objective trains deep networks layer-by-layer without backpropagation and matches backpropagation accuracy on small image benchmarks in the reported runs.

Pith tools