ALU uses public data to suppress unlearning cost quadratically while characterizing distribution mismatch effects, enabling mass unlearning with maintained utility.
MINE: Mutual information neural estimation
12 Pith papers cite this work, alongside 310 external citations. Polarity classification is still indexing.
abstract
We argue that the estimation of mutual information between high dimensional continuous random variables can be achieved by gradient descent over neural networks. We present a Mutual Information Neural Estimator (MINE) that is linearly scalable in dimensionality as well as in sample size, trainable through back-prop, and strongly consistent. We present a handful of applications on which MINE can be used to minimize or maximize mutual information. We apply MINE to improve adversarially trained generative models. We also use MINE to implement Information Bottleneck, applying it to supervised classification; our results demonstrate substantial improvement in flexibility and performance in these settings.
citation-role summary
citation-polarity summary
representative citing papers
NEO is a probabilistic neural model that induces compositional programs as a learned Language of Thought from non-textual observations and executes them via a shared transition model to enable explanation-driven generalization.
MP-IB uses an 8x information asymmetry via FP16 trait heads and INT4 state heads to disentangle speaker identity from agitation in voice biomarkers, outperforming larger models on edge devices with low latency and suppressed identity leakage.
On-policy self-distillation with sampled demonstrations reduces rollout diversity by amplifying existing probability gaps in the base model, unlike ideal RL which preserves ratios among correct outputs.
InfoShield uses TimeAwareMINE to minimize mutual information between speech representations and sensitive attributes, cutting gender inference from 92.6% to 55.5% and age inference from 55.7% to 30.3% while dropping depression F1 by only 6%.
DVSA combines bidirectional attention, MI-based contrastive learning, and dynamic label disambiguation to improve zero-shot learning performance under ambiguous (noisy) labels.
An iterative bootstrapped self-filtering approach selects balanced clean and diverse subsets from noisy vision-language datasets to train improved CLIP models.
Proposes a semantic information theory for LLMs that substitutes the token for the bit as the atomic carrier of meaning, recasts the Transformer as an energy-based model, and derives directed rate-distortion and rate-reward functions using Massey's directed information.
DVSA improves zero-shot learning under ambiguous labels by mutually calibrating visual features and attributes with attention and dynamic disambiguation.
Introduces IFM loss regularization for CNNs to learn correlated discriminative features, tested on shiftedMNIST dataset.
The paper introduces Semantic Variational Bayes (SVB) derived from the author's Semantic Information G Theory, claiming simpler computation than standard VB for latent variable inference via maximum G/R criterion.
The paper introduces 'Platonic Projection Structures,' a reformulation of standard PSD operator theory applied to representation learning, with experiments that verify definitions rather than test predictions.
citing papers explorer
-
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
ALU uses public data to suppress unlearning cost quadratically while characterizing distribution mismatch effects, enabling mass unlearning with maintained utility.
-
Learning to Theorize the World from Observation
NEO is a probabilistic neural model that induces compositional programs as a learned Language of Thought from non-textual observations and executes them via a shared transition model to enable explanation-driven generalization.
-
Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection
MP-IB uses an 8x information asymmetry via FP16 trait heads and INT4 state heads to disentangle speaker identity from agitation in voice biomarkers, outperforming larger models on edge devices with low latency and suppressed identity leakage.
-
On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity
On-policy self-distillation with sampled demonstrations reduces rollout diversity by amplifying existing probability gaps in the base model, unlike ideal RL which preserves ratios among correct outputs.
-
InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization
InfoShield uses TimeAwareMINE to minimize mutual information between speech representations and sensitive attributes, cutting gender inference from 92.6% to 55.5% and age inference from 55.7% to 30.3% while dropping depression F1 by only 6%.
-
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
DVSA combines bidirectional attention, MI-based contrastive learning, and dynamic label disambiguation to improve zero-shot learning performance under ambiguous (noisy) labels.
-
Data Selection Through Iterative Self-Filtering for Vision-Language Settings
An iterative bootstrapped self-filtering approach selects balanced clean and diverse subsets from noisy vision-language datasets to train improved CLIP models.
-
Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
Proposes a semantic information theory for LLMs that substitutes the token for the bit as the atomic carrier of meaning, recasts the Transformer as an energy-based model, and derives directed rate-distortion and rate-reward functions using Massey's directed information.
-
Dynamic Visual-semantic Alignment for Zero-shot Learning with Ambiguous Labels
DVSA improves zero-shot learning under ambiguous labels by mutually calibrating visual features and attributes with attention and dynamic disambiguation.
-
Learning to Find Correlated Features by Maximizing Information Flow in Convolutional Neural Networks
Introduces IFM loss regularization for CNNs to learn correlated discriminative features, tested on shiftedMNIST dataset.
-
Semantic Variational Bayes Based on Semantic Information G Theory for Solving Latent Variables
The paper introduces Semantic Variational Bayes (SVB) derived from the author's Semantic Information G Theory, claiming simpler computation than standard VB for latent variable inference via maximum G/R criterion.
-
Platonic Projection Structures: Operator-Induced Observability in Representation Learning
The paper introduces 'Platonic Projection Structures,' a reformulation of standard PSD operator theory applied to representation learning, with experiments that verify definitions rather than test predictions.