Pith. sign in

REVIEW 7 cited by

On Mutual Information Maximization for Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.13625 v2 pith:7JFORSTF submitted 2019-07-31 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningmethodsrepresentationargueestimatefeatureinformationmutual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many recent methods for unsupervised or self-supervised representation learning train feature extractors by maximizing an estimate of the mutual information (MI) between different views of the data. This comes with several immediate problems: For example, MI is notoriously hard to estimate, and using it as an objective for representation learning may lead to highly entangled representations due to its invariance under arbitrary invertible transformations. Nevertheless, these methods have been repeatedly shown to excel in practice. In this paper we argue, and provide empirical evidence, that the success of these methods cannot be attributed to the properties of MI alone, and that they strongly depend on the inductive bias in both the choice of feature extractor architectures and the parametrization of the employed MI estimators. Finally, we establish a connection to deep metric learning and argue that this interpretation may be a plausible explanation for the success of the recently introduced methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 220 citations worldwide. Full citation record

  1. Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning

    cs.CV 2025-06 reject novelty 6.0 of 10

    AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.

  2. DeInfoReg: A Decoupled Learning Framework for Better Training Throughput

    cs.LG 2025-06 conditional novelty 6.0 of 10

    DeInfoReg trains deep networks with per-module local losses so gradients flow only within each module, improving accuracy and enabling pipeline parallelism, with speedups of up to 1.47x over single-GPU backpropagation.

  3. A Mathematical Perspective On Contrastive Learning

    stat.ML 2025-05 conditional novelty 6.0 of 10

    A probabilistic tilting framework for contrastive learning yields closed-form Gaussian results showing which conditional statistics each loss can recover.

  4. Language-Aware Information Maximization for Transductive Few-Shot CLIP

    cs.CV 2025-08 conditional novelty 5.0 of 10

    LIMO, a transductive loss combining mutual information, zero-shot KL regularization, and LoRA, sets new state-of-the-art few-shot accuracy for CLIP on 11 datasets.

  5. Structure Maintained Representation Learning Neural Network for Causal Inference

    stat.ML 2025-08 reject novelty 5.0 of 10

    SMRLNN combines an adversarial discriminator with a canonical-correlation structure keeper to improve individual treatment effect estimation.

  6. Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Contrastive Successor Features recover ground-truth RL states up to a linear map whenever the skill-conditioned transition differences follow a von Mises-Fisher distribution and policies are diverse.

  7. C-LEAD: Contrastive Learning for Enhanced Adversarial Defense

    cs.CV 2025-10 reject novelty 2.0 of 10

    Contrastive learning with adversarial perturbations as positive pairs improves robustness of ResNet models on CIFAR-10, but evidence is weakened by missing baselines and inconsistent reporting.

Pith tools