{"id":"73b60f25-1a30-4397-96e1-1a5267c8c9cb","arxiv_id":"2505.14696","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"infomeasure is a new Python library unifying information-theoretic measures and estimators with local values, p-values, and t-scores, validated against known analytic and numerical benchmarks.","lead":"The authors present infomeasure, an open-source Python package for computing entropy, mutual information, transfer entropy, divergences, and related quantities with several estimation techniques. It is validated against known analytical cases and demonstrated on EEG data, offering a unified tool for reproducible information-theoretic analysis.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's robustness claim rests on only a few validated cells of the Table 2 measure–estimator matrix; the unvalidated conditional, divergence, bias-corrected, and hypothesis-testing branches are exactly where a bug would matter.","rationale":"I read the paper as a software contribution whose primary advertised value is breadth: many measures, many estimators, and a unified interface. The strongest claim in the abstract is that the package provides 'robust tools' for a wide variety of information-theoretic measures. For that claim to hold, the implementations behind the entire Table 2 matrix should be correct, not merely the four cases displayed in Figure 2. The reader identified this same generalization as the weakest assumption, and the preprint text itself supports the concern: the Validation section describes only Gaussian entropy, Gaussian MI, Schreiber's tent-map TE, and Ulam-map TE/MI, while Section 2 and Table 2 promise conditional variants, divergences, bias-corrected estimators, Rényi/Tsallis entropies, p-values, and t-scores. The Ulam-map example is also qualitative rather than an analytic check, so it does not independently confirm numerical accuracy. The authors point to a comprehensive test suite and CI/CD pipelines as evidence of correctness, but those tests are not reproduced in the preprint and no coverage report is given. This is not an internal inconsistency or a mathematical error in the shown validations; it is a gap between the breadth of the claim and the evidence presented. The proposed concrete test directly addresses this gap by exposing the untested branches to exact ground truths. Since the reader's conditional verdict already accounts for this limitation, my stress-test does not move the verdict; it reinforces the need for a public validation matrix before the broadest robustness claim is fully accepted.","tokens_in":11353,"tokens_out":3305,"duration_ms":36631,"concrete_test":"Add to the repository a validation script that loops over every (measure, estimator) entry in Table 2, plus the advertised extras (p-values, t-scores, local values), using small synthetic datasets whose true values are computable exactly (e.g. enumerated discrete distributions for KLD/JSD/conditional measures and known Gaussian integrals for continuous estimators), and report the full pass/fail matrix in the preprint or Zenodo archive. A minimal decisive first step is to run this sweep for bias-corrected conditional MI, which combines two untested features; if it passes within a stated tolerance, the generalization concern is substantially reduced.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that infomeasure provides robust tools across the full measure/estimator matrix in Table 2. The four shown validations cover Gaussian Shannon entropy, Gaussian mutual information, Schreiber's tent-map transfer entropy, and qualitative Ulam-map TE/MI. They do not exercise several advertised capabilities that are central to the package's claimed advantage: KLD and JSD, conditional MI and conditional TE, Rényi/Tsallis variants beyond the q=1.05 MI point, bias-corrected estimators, and the p-value/t-score machinery. The text asserts a unit-test suite in Section 3 and links to documentation demos, but the preprint does not show per-cell test results or coverage. A normalization or conditioning error in any of those unshown branches would weaken the 'robust tools' claim exactly for a user who selected that branch. Since the paper's contribution is breadth plus unification, this coverage gap is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents infomeasure, an open-source Python package for information-theoretic analysis. It implements Shannon, Rényi, and Tsallis entropies; joint and cross entropy; mutual information and conditional mutual information; transfer entropy and conditional transfer entropy; and Kullback-Leibler and Jensen-Shannon divergences. These measures are available through discrete, kernel, metric/kNN, ordinal, and bias-corrected estimators, with local values, p-values, t-scores, confidence intervals, and effective transfer entropy, as summarized in Table 2. The manuscript describes the architecture, gives basic usage examples, reports computational timings, and validates a subset of the functionality against analytical results for Gaussian entropy and mutual information and against Schreiber's tent-map and Ulam-map transfer entropy experiments. It concludes with an EEG case study that compares three TE estimators and control/patient differences.","tokens_in":11506,"tokens_out":9686,"duration_ms":88093,"significance":"If its claims hold, infomeasure would be a useful contribution to the information-theory software landscape by unifying measures, estimators, and hypothesis-testing machinery behind a common Python interface. The package is publicly available and versioned on Zenodo, which is a genuine reproducibility asset. The four reported validation experiments are well chosen and match the expected benchmarks: the Gaussian entropy and MI curves follow the closed-form expressions, and the tent-map fit gives alpha = 0.760 +/- 0.003, close to Schreiber's value. The main limitation is that these experiments cover only a small fraction of the matrix advertised in Table 2, leaving the 'robust tools' claim less supported than the abstract suggests.","major_comments":[{"comment":"The numerical validation covers only three measure rows (Shannon entropy, mutual information, transfer entropy, plus a qualitative Ulam-map check); conditional MI, conditional TE, KLD, JSD, Rényi/Tsallis variants beyond one point, bias-corrected estimators, local values, and the p-value/t-score machinery are not checked against known answers anywhere in the manuscript. Because the abstract's central claim is that infomeasure provides 'robust tools' across the full measure/estimator matrix, the evidence is incomplete for exactly the branches where a conditioning or normalization bug would be most dangerous. I ask the authors either to add analytic or cross-package validations for cMI/cTE, KLD/JSD, one bias-corrected entropy estimator, and the permutation null, or to restate the claim so that it explicitly covers only the validated subset.","section":"Validation / Table 2"},{"comment":"The package advertises p-values, t-scores, and confidence intervals for MI and TE, but no calibration test is reported. It is easy to implement a permutation test that returns numbers yet has the wrong null distribution; without a type-I-error check on synthetic null data, the hypothesis-testing feature cannot be considered validated. Please add a short experiment (e.g., false-positive rate at nominal alpha on independent Gaussian or surrogate data) or explicitly label this functionality as unvalidated.","section":"Measures, estimators, and features / Validation"}],"minor_comments":[{"comment":"The text says the box-kernel estimator deviates for small sigma 'resulting from the kernel bandwidth being too small', but with the fixed bandwidth=2 used in Listing 1 the issue is more naturally described as the bandwidth being large relative to the data spread; please clarify or correct the explanation.","section":"Validation (Gaussian entropy)"},{"comment":"Please state the range of epsilon used in the least-squares fit and the number of independent realizations per point; Eq. (3) is a small-epsilon approximation, so the fit range matters for interpreting alpha = 0.760 +/- 0.003.","section":"Validation (tent-map TE)"},{"comment":"The parameter noise_level=0.001 is used without being defined in the text; a one-sentence explanation of its role in the metric TE estimator would help new users.","section":"Listing 3"},{"comment":"The formula and the figure legend render the tent-map expression ambiguously ('2 2/ln(2)'); it should be typeset as alpha^2 epsilon^2 / ln 2.","section":"Eq. (3) and Figure 2 legend"},{"comment":"The bullet 'Rényi and Tsallis Estimations' should read 'Rényi and Tsallis entropies' for consistency with the surrounding list of measures.","section":"Measures, estimators, and features"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a competent software-description paper and the core validations are sound. I am recommending major revision rather than rejection because the gap between the abstract's 'robust tools' claim and the validated subset is real but fixable: the authors can add targeted validations or temper the claim. If the journal has space for supplementary material, a compressed test-coverage report would address most of the concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Carlson -- quick take: this is a solid software paper, better than most package announcements. What's new is not the individual estimators (all from the literature) but the breadth in one consistent API: entropy, MI, TE, conditional variants, divergences, with kernel, metric/kNN, ordinal, and bias-corrected estimators, plus local values, p-values, t-scores and effective TE. That breadth is the actual contribution, and it's genuinely useful for empirical fields.\n\nThe validation is the strong part. Gaussian entropy and MI match closed forms across estimators; the tent-map TE gives alpha = 0.760 +/- 0.003 against Schreiber's 0.77; the Ulam map gives the correct directional asymmetry. Those are external benchmarks, not circular checks. The code, docs, and Zenodo archive are real; the mention of a unit test suite is credible but not shown in the paper.\n\nThe soft spot is exactly what the stress test flags: Table 2 claims support across a wide matrix, but the paper only shows four cells (plus the qualitative Ulam result). The unvalidated branches--conditional MI/TE, KLD/JSD, Renyi/Tsallis variants, bias-corrected estimators, and the permutation-test machinery--are places where an implementation error would silently hurt a user. The paper says unit tests exist; the preprint doesn't show them. That is a load-bearing gap for the 'robust tools' claim, but it's fixable: provide per-cell validation curves or point reviewers to the CI/test output. I don't think the core argument is broken; this is a coverage question.\n\nAlso minor: the EEG case study is explicitly preliminary and the estimator disagreements are not statistically tested, but the authors say so themselves, so it's not a hidden flaw. The performance figure is on one machine; fine as indicative.\n\nWho's this for? Anyone currently cobbling together JIDT, scipy, and custom code for information-theoretic measures; this could plausibly replace that pile for many use cases. The paper deserves a serious referee--not a desk reject. If I were editor, I'd send it out and ask the referee to push for the missing test evidence before accepting. I would cite it if I needed a single package for this.\n\nRecommendation: accept conditional on a supplement that maps Table 2 cells to validation results or machine-readable test reports.","headline":"A genuinely useful unified information-theory package, validated on key analytic benchmarks but with a coverage gap across its full measure-estimator matrix; worth refereeing with a request for per-cell tests.","tokens_in":12032,"tokens_out":1767,"would_cite":true,"duration_ms":17020,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A17","62B10"],"pacs":[],"model":"deepseek-v4-flash","headline":"One Python package unifies info-theory measures and estimators.","keywords":["information theory","Python package","transfer entropy","mutual information","entropy estimation","KSG estimator","ordinal patterns","reproducibility"],"falsifier":"Take a supported combination that is not shown in Figure 2, such as conditional transfer entropy with the bias-corrected estimator or Jensen-Shannon divergence between two Gaussian distributions, and compare infomeasure's output against an independent analytic or high-precision numerical reference; one clear mismatch beyond expected estimator error would show that the package's breadth claim overreaches its validation.","tokens_in":11150,"feed_emoji":"📦","tokens_out":5346,"duration_ms":50099,"temperature":0.7,"pith_summary":"infomeasure is a Python package that puts a broad family of information-theoretic measures — Shannon, Rényi and Tsallis entropies, mutual information, transfer entropy, cross-entropy, and KL and Jensen-Shannon divergences — behind a single interface, for both discrete and continuous data. The paper's central claim is that this unification makes information-theoretic analysis practical and reproducible: a user can switch among kernel, nearest-neighbour (KL/KSG), ordinal, and bias-corrected estimators by changing one argument, and can obtain local values, p-values, t-scores, and confidence intervals within the same framework. The claim is supported by validation against known analytic solutions for Gaussian entropy and mutual information, by reproduction of the canonical coupled-map-lattice transfer-entropy benchmarks, and by a schizophrenia EEG case study showing how the choice of estimator changes the observed pattern of information flow. A sympathetic reader would take the paper to be establishing that the package's full measure-estimator matrix is correct and usable, and that this breadth is itself the contribution.","feed_headline":"One Python package unifies info-theory measures and estimators","feed_subtitle":"infomeasure puts discrete and continuous estimation, local values, p-values, and t-scores behind one interface.","key_machinery":"The load-bearing object is the `Estimator` base class and its inheritance-based design: every measure-estimator combination is an estimator object exposing the same interface, with mixins such as `PValueMixin` and `EffectiveValueMixin` adding permutation tests and effective transfer entropy. High-level functional entry points (`im.entropy`, `im.mutual_information`, `im.transfer_entropy`) map user-facing names like \"kernel\" or \"ordinal\" to the right estimator class through a dynamic import mechanism. This architecture is what lets the paper claim breadth without an explosion of special cases: the same machinery handles discrete, continuous, conditional, and local variants, and it is what makes the validation results transfer across the table of supported measures.","core_discovery":"The discovery presented is a working, open-source implementation that covers essentially all combinations of a substantial measure set and several estimator families: discrete plug-in, kernel density, KL/kNN/KSG, and ordinal/permutation estimators, with bias-corrected entropy estimators for small samples. On the paper's own terms, the package's defining achievement is not any single new estimator but the integration: consistent slicing for transfer entropy, a shared interface for hypothesis testing, an effective-transfer-entropy variant, and a modular design where a central estimator base class and mixins add p-values and local values without code duplication. The validation section demonstrates that the package reproduces analytical values for Gaussian entropy and mutual information, and recovers the theoretical coupling-strength scaling of transfer entropy in coupled tent-map lattices with fitted coefficient 0.760 ± 0.003 against a literature value near 0.77. The EEG case study then shows the framework's value in practice, where different estimators give different and partly complementary pictures of information transfer.","pith_inferences":["Beyond the paper: because the package exposes a uniform estimator interface, one could systematically benchmark estimator families against each other on synthetic ground-truth processes, producing an estimator-selection guide that the current paper leaves to user judgment.","Beyond the paper: the unvalidated cells of the measure-estimator matrix are a concrete testing target; a natural follow-up would be a formal cross-validation of every combination in Table 2 against analytic or high-precision numerical references.","Beyond the paper: the EEG results suggest that disagreement among estimators is itself informative; one testable extension is to use the pattern of estimator disagreement as a feature for classifying patient versus control recordings."],"forward_implications":["A practitioner can test several estimators on one dataset by changing a single argument, making estimator-sensitivity analysis a routine step instead of a toolbox-integration project.","Because transfer entropy, mutual information, and entropy all share one interface with p-values, t-scores, and local values, hypothesis testing and time-resolved information flow can be reported consistently across studies.","The coupled-map-lattice benchmarks indicate that the package reproduces known theoretical scaling, so it can serve as a reference implementation for new estimators proposed in the literature.","The EEG case illustrates that estimator choice can change the qualitative conclusion about brain-network differences, implying that multi-estimator comparison should be part of any applied information-theoretic analysis."],"supporting_citations":[{"why":"Defines transfer entropy and supplies the canonical coupled-map-lattice benchmarks used for validation.","marker":"[32]"},{"why":"Provides the Kozachenko-Leonenko nearest-neighbour entropy estimator family implemented in the package.","marker":"[35]"},{"why":"Provides the Kraskov-Stögbauer-Grassberger mutual-information estimator that the package's metric/kNN branch builds on.","marker":"[36]"},{"why":"Provides the ordinal/permutation symbolisation used by the ordinal estimator family.","marker":"[37]"},{"why":"One of the comparative reviews used to select the bias-corrected entropy estimators.","marker":"[38]"},{"why":"Second comparative review used to select bias-corrected estimators for short sequences.","marker":"[39]"},{"why":"Establishes local information measures, which the package implements for entropy, mutual information, and transfer entropy.","marker":"[41]"},{"why":"Introduces effective transfer entropy, implemented as the eTE variant for finite-sample bias reduction.","marker":"[42]"}],"fun_headline_variants":["infomeasure: one Python package for info-theory measures","Open-source infomeasure unifies info-theory estimation","Python library infomeasure covers entropies and transfer entropy","infomeasure simplifies info-theory analysis in Python","New Python package integrates info-theory estimators and p-values"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole 'robust tools' claim rests on the assumption that the four validation experiments in Figure 2, together with the unit tests and documentation examples, guarantee that every supported measure-estimator combination is correctly implemented.","fun_headline_variants_meta":{"raw":{"variants":["infomeasure: one Python package for info-theory measures","Open-source infomeasure unifies info-theory estimation","Python library infomeasure covers entropies and transfer entropy","infomeasure simplifies info-theory analysis in Python","New Python package integrates info-theory estimators and p-values"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1426,"prompt_tokens":929,"completion_tokens":497,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":415}},"tokens_in":545,"tokens_out":497,"duration_ms":5080,"temperature":1.0,"reasoning_tokens":415,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:35:57.767286+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a supported combination that is not shown in Figure 2, such as conditional transfer entropy with the bias-corrected estimator or Jensen-Shannon divergence between two Gaussian distributions, and compare infomeasure's output against an independent analytic or high-precision numerical reference; one clear mismatch beyond expected estimator error would show that the package's breadth claim overreaches its validation.","supporting_citations":[{"cited_title":"& Leonenko, N","cited_arxiv_id":null,"evidence_quote":"Provides the Kozachenko-Leonenko nearest-neighbour entropy estimator family implemented in the package."},{"cited_title":"J., Legón-Pérez, C","cited_arxiv_id":null,"evidence_quote":"Second comparative review used to select bias-corrected estimators for short sequences."},{"cited_title":"T.Measuring the Dynamics of Information Processing on a Local Scale in Time and Space, 161–193 (Springer Berlin Heidelberg, Berlin, Heidelberg, 2014)","cited_arxiv_id":null,"evidence_quote":"Establishes local information measures, which the package implements for entropy, mutual information, and transfer entropy."},{"cited_title":"& Kantz, H","cited_arxiv_id":null,"evidence_quote":"Introduces effective transfer entropy, implemented as the eTE variant for finite-sample bias reduction."}],"review_version":1}