Pith. sign in

REVIEW 20 cited by

Estimating Mutual Information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv cond-mat/0305641 v1 pith:HXJ4YJH4 submitted 2003-05-28 cond-mat.stat-mech cond-mat.dis-nn

Estimating Mutual Information

classification cond-mat.stat-mech cond-mat.dis-nn
keywords estimatorsalgorithmsbiasdatadensitydistributionsentropyestimates
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present two classes of improved estimators for mutual information $M(X,Y)$, from samples of random points distributed according to some joint probability density $\mu(x,y)$. In contrast to conventional estimators based on binnings, they are based on entropy estimates from $k$-nearest neighbour distances. This means that they are data efficient (with $k=1$ we resolve structures down to the smallest possible scales), adaptive (the resolution is higher where data are more numerous), and have minimal bias. Indeed, the bias of the underlying entropy estimates is mainly due to non-uniformity of the density at the smallest resolved scale, giving typically systematic errors which scale as functions of $k/N$ for $N$ points. Numerically, we find that both families become {\it exact} for independent distributions, i.e. the estimator $\hat M(X,Y)$ vanishes (up to statistical fluctuations) if $\mu(x,y) = \mu(x) \mu(y)$. This holds for all tested marginal distributions and for all dimensions of $x$ and $y$. In addition, we give estimators for redundancies between more than 2 random variables. We compare our algorithms in detail with existing algorithms. Finally, we demonstrate the usefulness of our estimators for assessing the actual independence of components obtained from independent component analysis (ICA), for improving ICA, and for estimating the reliability of blind source separation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

    cs.AI 2026-04 conditional novelty 8.0

    Above a computable autonomy threshold with mixed human–AI feedback cycles, no single-locus accountability framework can jointly satisfy causal attribution, foreseeability, non-vacuity, and full allocation.

  2. Same Signal, Different Story: Demystifying Receiver Effects in Wi-Fi Channel State Information

    eess.SP 2026-05 conditional novelty 7.0

    Receiver-specific CSI differences in Wi-Fi sensing stem primarily from AGC and subcarrier nonlinearities and are largely mitigated by gain alignment, restoring most cross-device model accuracy.

  3. Task Relevance Is Not Local Replaceability: A Two-Axis View of Channel Information

    cs.CV 2026-05 unverdicted novelty 7.0

    Channel importance splits into task relevance and local replaceability; local-axis metrics predict safe removal under pruning better than target-axis metrics across multiple CNNs and datasets.

  4. Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection

    cs.LG 2026-05 unverdicted novelty 7.0

    MP-IB uses an 8x information asymmetry via FP16 trait heads and INT4 state heads to disentangle speaker identity from agitation in voice biomarkers, outperforming larger models on edge devices with low latency and sup...

  5. Quantum Causal Discovery via Amplitude Estimation of Kullback-Leibler Divergence

    quant-ph 2026-04 unverdicted novelty 7.0

    QKLA achieves quadratic query-complexity improvement for clipped KL estimation, yielding 2.7-7.4x fewer oracle queries than classical methods when embedded in the PC causal-discovery algorithm at moderate precision.

  6. The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

    cs.AI 2026-04 unverdicted novelty 7.0

    The Accountability Incompleteness Theorem demonstrates that human-AI collectives above the Accountability Horizon with feedback cycles cannot simultaneously meet attributability, foreseeability, non-vacuity, and compl...

  7. DegenDetector: Symbolic Recovery of Parameter Degeneracies in Bayesian Posteriors

    astro-ph.IM 2026-07 conditional novelty 6.0

    A pipeline combining mutual information screening with alternating symbolic regression recovers closed-form degeneracy equations from posterior samples.

  8. A Self-Supervised Approach for Minimal-Annotation Hydroacoustic Data Exploration

    cs.SD 2026-07 conditional novelty 6.0

    Event-level MAE embeddings plus UMAP/HDBSCAN or K-Means clustering recover 15 hydroacoustic classes from multi-year Mayotte data with ~1 hour of annotation and detector-comparable F1.

  9. Information-Theoretic Reliability is Robust to Analytic Choice: A 24-Specification Multiverse on Public Cognitive Test-Retest Data

    stat.ME 2026-05 unverdicted novelty 6.0

    NLRΔ applied to 50 primary measures from five cognitive task families on public test-retest data yields median -0.138 nats with zero cells passing the headline rule across a 24-specification multiverse.

  10. Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

    cs.AI 2026-05 unverdicted novelty 6.0

    Spectral partitioning on pairwise mutual-information graphs from agent hidden states detects representational coalitions that behavioral measures miss in multi-agent AI.

  11. Efficient characterization of general Gottesman-Kitaev-Preskill qubits

    quant-ph 2026-04 unverdicted novelty 6.0

    A family of positive semidefinite operators is introduced that witnesses arbitrary logical GKP qubit states and enables their efficient characterization via three quadrature measurements.

  12. Efficient characterization of general Gottesman-Kitaev-Preskill qubits

    quant-ph 2026-04 unverdicted novelty 6.0

    A new set of positive semidefinite operators for GKP qubits gives logical infidelity via three quadrature measurements and yields finite-dimensional approximations as ground states.

  13. Feature Selection via Mutual Information: New Theoretical Insights

    cs.LG 2019-07 unverdicted novelty 6.0

    Conditional mutual information bounds ideal prediction errors for feature subsets and supplies a stopping condition for greedy selection algorithms.

  14. Investigating causality between principal components in protein dynamics

    physics.chem-ph 2026-06 unverdicted novelty 5.0

    Applies causal inference to PCs from MD trajectories of two proteins to construct directed influence networks complementary to PCA and TICA.

  15. Inferring stellar metallicity and elemental abundances from kinematic and spectroscopic data using machine learning -- Implications for exoplanet host stars

    astro-ph.EP 2026-05 unverdicted novelty 5.0

    ML regressors trained on APOGEE DR17 red giants predict C, O, Mg, Si abundances from kinematics and [Fe/H] more accurately than [Fe/H] baseline, with external validation on HARPS FGK dwarfs and reproduction of Galacti...

  16. Efficient Benchmarking Is Just Feature Selection and Multiple Regression

    stat.ML 2026-05 unverdicted novelty 4.0

    Kernel ridge regression combined with mRMR feature selection improves prediction of full benchmark scores from question subsets over existing efficient benchmarking techniques.

  17. Data-Driven Reduction of Fault Location Errors in Onshore Wind Farm Collectors

    eess.SY 2025-11 unverdicted novelty 4.0

    A Gated Residual Network correction model reduces fault location error by 76% in simulated onshore wind farm collector networks compared to state-of-the-art methods.

  18. Information-Theoretic Measures in AI: A Practical Decision Framework

    cs.AI 2026-04 unverdicted novelty 3.0

    A practical guide that organizes seven IT measures around three questions each—what it answers in AI, suitable estimators, and dangerous misuses—complete with flowchart, table, and worked examples.

  19. Information-Theoretic Measures in AI: A Practical Decision Framework

    cs.AI 2026-04 conditional novelty 3.0

    The paper is a review that packages existing IT-estimator guidance into a seven-measure decision framework; a claimed validation case study is absent from the full text.

  20. Information-Theoretic Measures in AI: A Practical Decision Framework

    cs.AI 2026-04 unverdicted novelty 3.0

    A synthesis paper offering a practical decision guide, flowchart, and table for choosing among seven established information-theoretic measures in AI and agent applications.