Pith. sign in

REVIEW 8 cited by

MINE: Mutual Information Neural Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1801.04062 v5 pith:63QVOCCV submitted 2018-01-12 cs.LG stat.ML

classification cs.LGstat.ML
keywords informationminemutualneuralestimationachievedadversariallyapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We argue that the estimation of mutual information between high dimensional continuous random variables can be achieved by gradient descent over neural networks. We present a Mutual Information Neural Estimator (MINE) that is linearly scalable in dimensionality as well as in sample size, trainable through back-prop, and strongly consistent. We present a handful of applications on which MINE can be used to minimize or maximize mutual information. We apply MINE to improve adversarially trained generative models. We also use MINE to implement Information Bottleneck, applying it to supervised classification; our results demonstrate substantial improvement in flexibility and performance in these settings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 310 citations worldwide. Full citation record

  1. Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Statistical alignment with human ratings does not imply strategic alignment: LLM-as-a-Judge is highly correlated but easily manipulated, while a new statement-level mutual-information metric is robust to manipulation.

  2. Sky Background Building of Multi-objective Fiber spectra Based on Mutual Information Network

    cs.CV 2025-08 reject novelty 6.0 of 10

    A mutual-information neural network estimates per-fiber sky backgrounds from all spectra in a LAMOST exposure, outperforming the standard Super sky on sky fibers, though object-fiber performance is not validated.

  3. Score-based Diffusion Model for Unpaired Virtual Histology Staining

    eess.IV 2025-06 conditional novelty 6.0 of 10

    An unpaired, mutual-information-guided diffusion model translates H&E histology images into IHC images with improved structural and staining fidelity.

  4. DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    DVSA combines bidirectional attention, MI-based contrastive learning, and dynamic label disambiguation to improve zero-shot learning performance under ambiguous (noisy) labels.

  5. Towards Comprehensive Information-theoretic Multi-view Learning

    cs.LG 2025-09 reject novelty 4.0 of 10

    CIML combines a Gács-Körner-style common-representation objective with per-view information-bottleneck unique representations and independence constraints, reporting state-of-the-art accuracy on six multi-view datasets.

  6. Domain Adaptation via Feature Refinement

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    DAFR2 combines target-data batch normalization adaptation, feature distillation, and hypothesis transfer to make models robust to image corruption without target labels.

  7. Causal-Inspired Multi-Agent Decision-Making via Graph Reinforcement Learning

    cs.MA 2025-07 reject novelty 4.0 of 10

    A graph-RL driving agent using VGAE-based causal feature extraction achieves lower collision rates and higher rewards at a simulated unsignalized intersection than graph-RL baselines.

  8. Platonic Projection Structures: Operator-Induced Observability in Representation Learning

    cs.LG 2026-07 reject novelty 2.0 of 10

    The paper introduces 'Platonic Projection Structures,' a reformulation of standard PSD operator theory applied to representation learning, with experiments that verify definitions rather than test predictions.

Pith tools