REVIEW 8 cited by
MINE: Mutual Information Neural Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We argue that the estimation of mutual information between high dimensional continuous random variables can be achieved by gradient descent over neural networks. We present a Mutual Information Neural Estimator (MINE) that is linearly scalable in dimensionality as well as in sample size, trainable through back-prop, and strongly consistent. We present a handful of applications on which MINE can be used to minimize or maximize mutual information. We apply MINE to improve adversarially trained generative models. We also use MINE to implement Information Bottleneck, applying it to supervised classification; our results demonstrate substantial improvement in flexibility and performance in these settings.
Forward citations
Cited by 8 Pith papers
-
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
Statistical alignment with human ratings does not imply strategic alignment: LLM-as-a-Judge is highly correlated but easily manipulated, while a new statement-level mutual-information metric is robust to manipulation.
-
Sky Background Building of Multi-objective Fiber spectra Based on Mutual Information Network
A mutual-information neural network estimates per-fiber sky backgrounds from all spectra in a LAMOST exposure, outperforming the standard Super sky on sky fibers, though object-fiber performance is not validated.
-
Score-based Diffusion Model for Unpaired Virtual Histology Staining
An unpaired, mutual-information-guided diffusion model translates H&E histology images into IHC images with improved structural and staining fidelity.
-
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
DVSA combines bidirectional attention, MI-based contrastive learning, and dynamic label disambiguation to improve zero-shot learning performance under ambiguous (noisy) labels.
-
Towards Comprehensive Information-theoretic Multi-view Learning
CIML combines a Gács-Körner-style common-representation objective with per-view information-bottleneck unique representations and independence constraints, reporting state-of-the-art accuracy on six multi-view datasets.
-
Domain Adaptation via Feature Refinement
DAFR2 combines target-data batch normalization adaptation, feature distillation, and hypothesis transfer to make models robust to image corruption without target labels.
-
Causal-Inspired Multi-Agent Decision-Making via Graph Reinforcement Learning
A graph-RL driving agent using VGAE-based causal feature extraction achieves lower collision rates and higher rewards at a simulated unsignalized intersection than graph-RL baselines.
-
Platonic Projection Structures: Operator-Induced Observability in Representation Learning
The paper introduces 'Platonic Projection Structures,' a reformulation of standard PSD operator theory applied to representation learning, with experiments that verify definitions rather than test predictions.
Discussion (0). Sign in to comment.