{"id":"f7990434-35d7-4abf-98b8-99d7cdaee7b4","arxiv_id":"2505.01607","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Starting from minimum description length, the paper derives and reviews the minimax entropy principle: the optimal features are those that yield the maximum entropy model with the lowest entropy.","lead":"The paper derives the minimax entropy principle: the best features to include in a model are the ones that make the maximum entropy model as certain as possible, with the lowest entropy. It then reviews applications in neuroscience, texture modeling, and handwritten digits, and outlines the computational challenges for high-dimensional biological data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Minimax entropy selection is exact only under noiseless feature averages and large sample limits; with finite data, minimizing model entropy can overfit sampling fluctuations, so the claimed 'optimal features' lack finite-sample justification.","rationale":"The reader's weakest-assumption analysis identified exactly the load-bearing condition: exact feature averages (Sec. III.A, near Eq. 15). My stress-test concurs and adds that the Appendix A MDL derivation itself is a large-T approximation, so finite-sample issues affect the derivation of the maximum-entropy model from MDL, not just the subsequent KL-divergence interpretation. The paper is internally consistent under its stated idealizations, and the theoretical core is a valid restatement of the minimax entropy principle previously proposed in refs. 6-7. The main gap is that the abstract and central claim frame the result as yielding 'optimal features' for real systems, while the derivation requires noiseless, effectively infinite data. The reader's CONDITIONAL verdict already captures this limitation, so no change in verdict is needed. My concrete test would settle whether the finite-sample failure is quantitatively severe enough to demand a penalty term or alternative selection criterion in practical applications.","tokens_in":15878,"tokens_out":4902,"duration_ms":60563,"concrete_test":"Run a synthetic experiment with a known N-spin Ising model (e.g., N = 20) and a fixed feature budget M. Compute the true optimal M features using exact averages from Ptrue, then compute the minimax-optimal M features from empirical averages estimated with T samples for T = 100, 1000, 10000. Compare the selected feature sets and the resulting DKL(Ptrue||PF) for the finite-sample optimal model versus the true optimal model. If the finite-sample minimax choice has substantially larger KL divergence to Ptrue (or if its entropy decreases while its KL divergence increases with added features), the concern that finite-sample fluctuations break the optimality claim is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central equivalence in Sec. II.B, L(PF) = S(PF) (Eqs. 10-11), and the accuracy interpretation in Sec. III.A, DKL(Ptrue||PF) = S(PF) - S(Ptrue) (Eq. 15), both rely on the assumption that the empirical feature averages are exact: <f_mu(x)>_exp = <f_mu(x)>_true. The paper states this assumption near Eq. 15, but it is load-bearing for the headline claim that the minimum-entropy maximum-entropy model is the optimal model for the system. With finite samples, empirical averages differ from true averages by fluctuations of order 1/sqrt(T). A maximum-entropy model fitted to noisy averages will have an entropy that is biased downward by overfitting, and minimizing S(PF) over feature sets will preferentially select features that explain sampling noise rather than genuine structure. Consequently, the equality between minimizing model entropy and minimizing KL divergence to the true distribution breaks down: DKL(Ptrue||PF) is no longer S(PF) - S(Ptrue), because the model does not match the true feature averages. Furthermore, the Appendix A derivation of maximum entropy from MDL is explicitly a large-T limit; for finite T, the description length minimizer is not generally the maximum-entropy distribution. The paper does not provide a finite-sample correction, a model-complexity penalty, or a holdout-based stopping rule, so the minimax entropy criterion as stated cannot justify calling a feature set 'optimal' for finite experimental data. This concern does not invalidate the idealized derivation, but it limits the central claim to noiseless, large-sample settings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper claims to derive the minimax entropy principle from the minimum description length (MDL) principle. The central argument is that for a maximum-entropy model P_F(x) built from feature set F, the description length equals the model entropy, L(P_F)=S(P_F) (Eqs. 10-11), so the optimal feature set is the one whose maximum-entropy model has the smallest entropy (Eqs. 12-13). Under the assumption that empirical feature averages equal the true averages, the paper also shows that minimizing S(P_F) is equivalent to minimizing the KL divergence to the true distribution and to maximizing the information gain relative to an independent model. The authors illustrate the principle with a three-spin Ising example, then survey applications to Gaussian graphical models, Ising models of neural populations, texture modeling, and multiple-distribution models of handwritten digits. They discuss greedy algorithms, exact solutions for trees and generalized series-parallel networks, and open challenges such as submodularity and parameterized features.","tokens_in":16118,"tokens_out":7167,"duration_ms":83831,"significance":"If the idealized claim holds, the paper provides a clean unification of feature selection with maximum-entropy modeling, and the exactly solvable tree and GSP results are valuable practical contributions. The derivation is self-contained and the main equalities are easy to verify, which is a strength. However, the central claim is stated without the finite-sample qualifications that the formal argument actually requires, and the applications to real neural and brain data inherit that gap. The paper is therefore a useful perspective with a correct core, but its headline claim overreaches as written.","major_comments":[{"comment":"The equalities L(P_F)=S(P_F) and D_KL(P_true||P_F)=S(P_F)-S(P_true) rely on the assumption, stated near Eq. (15), that the measured feature averages are exact: <f_mu(x)>_exp = <f_mu(x)>_true. For finite or noisy samples this assumption fails. If the fitted model matches the empirical averages but not the true averages, the KL divergence gains an additional term sum_mu lambda_mu (<f_mu>_exp - <f_mu>_true), and minimizing S(P_F) is no longer equivalent to minimizing D_KL(P_true||P_F). Because all applications in Section IV use real, finite data without a finite-sample correction or holdout validation, the statement that F* is 'optimal for the system' is not justified for those datasets. The authors should either restrict the claim to the noiseless limit or add a model-complexity/regularization term and discuss overfitting.","section":"Section III.A, Eq. (15); Section II.B, Eqs. (9)-(11)"},{"comment":"The quantity L(P) in Eq. (4) is the expected per-symbol code length averaged over datasets consistent with the observed features, not the full MDL stochastic complexity of the observed data. Standard MDL also charges for the cost of describing the model, including the feature set F and the parameters lambda_mu, and no such term appears in Eqs. (9)-(13). The paper's claim that the minimax entropy criterion is derived 'starting only from the MDL principle' (Introduction) is therefore overstated: the derivation shows that, among models with a fixed feature budget, the maximum-entropy model minimizes the data-coding cost, and that among feature sets of equal size the cross-entropy is minimized by minimizing model entropy. The authors should either reformulate the claim as a fixed-budget cross-entropy minimization or incorporate a model-complexity term and explain how the minimax criterion would change.","section":"Section II.A-B; Appendix A"},{"comment":"The applications to fMRI and neuronal recordings minimize entropy using empirical covariances and correlations from finite samples, but the paper does not quantify the sample sizes relative to the number of constraints or test predictive performance on held-out data. As a result, the reported 'optimal' networks could be selecting features that fit sampling fluctuations rather than genuine structure. I ask the authors to add a brief discussion of how the exact-averages assumption is approximated in these datasets and to state clearly that the reported improvements are measured for the empirical distributions, not for the true distributions.","section":"Section IV.B-C; Figs. 3-4"}],"minor_comments":[{"comment":"The phrase 'computing the the gradients' contains a duplicated article and should read 'computing the gradients'.","section":"Section V.C"},{"comment":"The spelling 'na¨ıve' uses a nonstandard dieresis; the standard English spelling is 'naive'.","section":"Section IV.A"},{"comment":"Several references lack complete publication data: Ref. [3] has no year, Ref. [7] has no year or page range, Ref. [62] has no year, and Ref. [63] has no year. These should be completed.","section":"References"},{"comment":"The multiple-distribution formulation in Eq. (26) does not state the objective being optimized over the shared feature set; the text should specify whether the criterion is the sum, average, or another combination of the entropies S(P_alpha).","section":"Section IV.E"},{"comment":"The caption says 'the optimal bound indicates the minimum entropy' but does not describe how the bound is computed; a sentence on the search procedure would improve reproducibility.","section":"Figure 1(b)"}],"recommendation":"major_revision","confidential_remarks":"The formal core of the paper is sound under the stated idealization of exact feature averages, and the tree/GSP results are a genuine strength. The main risk is that the abstract and applications overclaim optimality for real data without a finite-sample analysis. I recommend major revision rather than rejection, with the expectation that the authors add explicit caveats and, ideally, a finite-sample correction or complexity penalty."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a clean, honest review with a compact derivation. The minimax entropy principle isn't new — Zhu, Wu, and Mumford stated it in 1997, and the authors say so. The contribution is that they derive it from a minimum description length argument and then survey recent applications, including their own solvable tree/GSP results. That's a useful synthesis, not a breakthrough. Citation pattern is fine: prior work is credited, and self-citations point to actually relevant prior results.\n\nThe derivation in Sec II is internally consistent. L(P_F)=S(P_F) follows from the maxent constraints and the definition of chi; Eq (15) correctly gives D_KL(P_true||P_F)=S(P_F)-S(P_true) under the stated assumption that the experimental feature averages equal the true ones. The three-spin example and the discussion of alternative divergence/information derivations are clear. Credit where due: the paper is explicit about the large-T limit in Appendix A and about the greedy algorithm not being globally optimal.\n\nWhere the soft spots are: the finite-sample issue is real, and it's load-bearing for how you read the word 'optimal.' The MDL derivation is explicitly a large-T limit, and the KL-divergence accuracy claim assumes exact feature averages. With finite data, minimizing entropy can overfit sampling fluctuations; the model entropy is biased downward by fitting noise, and the equality D_KL = S(P_F) - S(P_true) no longer holds because the fitted model doesn't match the true averages. The paper doesn't provide a finite-sample correction, a model-complexity penalty, or a holdout rule. That doesn't invalidate the idealized derivation, but the abstract's 'optimal features' is stronger than the assumptions justify. A referee should ask the authors to either add a caveat to the headline claims or show how the principle extends to finite samples.\n\nSecond soft spot: the lead fMRI application (Fig 3) is drawn from an in-preparation manuscript (ref 11). A referee can't verify it. The paper's other applications are from published or arXiv work, so this is a minor issue, but it should be fixed before publication — either the data/analysis should be made available or the figure should be flagged as preliminary.\n\nWho this is for: someone working on maximum entropy models in neuroscience or biophysics who wants a compact statement of minimax entropy and a map of where it can be applied. It's also a good discussion piece for a journal club on model selection. It deserves peer review, not a desk reject.\n\nMy recommendation: send it to a serious referee, with the finite-sample caveat and the in-preparation application as the two questions to press.","headline":"A clean, honest review that derives the known minimax entropy principle from MDL; worth a serious referee, but the 'optimal' claim only holds under exact feature averages and large T.","tokens_in":16757,"tokens_out":7034,"would_cite":true,"duration_ms":65335,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Feature selection reduces to minimizing the entropy of the maximum entropy model.","keywords":["minimax entropy","maximum entropy principle","minimum description length","feature selection","statistical physics","Ising model","Gaussian graphical model","neural activity"],"falsifier":"Take a system with a known ground-truth distribution, draw finite samples, compute empirical feature averages, apply the minimax entropy rule to select features, and compare the resulting model's KL divergence to the truth against a model selected with a finite-sample complexity penalty. If the minimax-selected model is consistently worse on held-out samples, the exact-averages assumption fails in realistic settings.","tokens_in":15625,"feed_emoji":"🧠","tokens_out":6798,"duration_ms":58974,"temperature":0.7,"pith_summary":"This paper derives a rule for choosing which measured features to include in a statistical model: among all allowed feature sets, the best one is the set whose maximum entropy model has the lowest entropy. The derivation starts from the minimum description length (MDL) principle, which equates good models with short encodings of data. For maximum entropy models, the description length turns out to equal the model's entropy, so minimizing description length over feature sets is the same as minimizing entropy. The paper shows that this minimax entropy principle also selects the model closest to the true distribution and the one capturing the most information, assuming the measured features are exact. If correct, it gives a principled, parameter-free way to find compressed models of complex systems, from neural populations to textures.","feed_headline":"Feature selection reduces to minimizing model entropy","feed_subtitle":"Minimum description length shows the best features are those yielding the lowest-entropy max-entropy model.","key_machinery":"The central object is the maximum entropy distribution $P_F(x) = \\frac{1}{Z} \\exp\\left(\\sum_{\\mu} \\lambda_\\mu f_\\mu(x)\\right)$, a Boltzmann distribution whose Lagrange multipliers $\\lambda_\\mu$ are fit so that the model reproduces the measured feature averages $\\langle f_\\mu(x)\\rangle_{exp}$. The load-bearing identity is $L(P_F) = S(P_F)$: for these models, the minimum description length of a dataset consistent with the features equals the model's entropy, so minimizing description length across feature sets is identical to minimizing entropy. In the case of tree-structured correlations, the information $I_G$ decomposes into a sum of pairwise mutual informations, which reduces the optimization to a minimum spanning tree problem and yields an exact, efficient solution.","core_discovery":"The central claim is that the optimal set of features $F^*$ is the one for which the maximum entropy model $P_{F^*}(x)$ has the minimum entropy $S(P_{F^*})$ among all allowed sets of features. The key identity is that the description length of any maximum entropy model equals its entropy, $L(P_F) = S(P_F)$ (Eqs. 10-11). Since the MDL principle says the best model is the one with minimum description length, and each feature set yields exactly one maximum entropy model, feature selection becomes the minimax problem $F^* = \\arg\\min_F S(P_F)$, with $P_F$ itself the entropy-maximizing distribution subject to the empirical feature averages. The same argument shows that $F^*$ minimizes the Kullback-Leibler divergence $D_{KL}(P_{true} \\| P_F)$ and maximizes the information $I_F = S(P_{ind}) - S(P_F)$ contained in the selected features, provided the measured averages are exact. This unifies the maximum entropy principle (choosing the model for given features) with the minimax entropy principle (choosing the features themselves).","pith_inferences":["With finite data, the minimax entropy optimum will overfit, so a practical extension is to add a finite-sample correction or complexity penalty to the description length; the paper notes this term is absent.","The identity $L(P_F) = S(P_F)$ suggests that model entropy alone could serve as a model-selection criterion in other maximum entropy applications, such as species distribution modeling or natural language processing.","The continuous, parameterized-feature version sketched in the paper could link minimax entropy to representation learning, where learned features are optimized rather than selected from a discrete set.","Whether entropy reduction is submodular for real datasets is an open empirical question; if it is, greedy feature selection would be near-optimal in practice."],"forward_implications":["Feature selection for maximum entropy models reduces to a single objective—minimizing the fitted model's entropy—with no separate regularization term needed.","In neural recordings, a small fraction of optimally chosen correlations can capture most of the information, enabling highly compressed descriptions of brain activity.","For tree-structured correlation networks, the optimal model is found exactly via minimum spanning trees, giving an efficient solution for high-dimensional Ising models.","The same principle applies across multiple distributions at once, allowing a shared set of features to compress data sets such as handwritten digits.","When entropy reduction is submodular, the greedy algorithm is guaranteed to be within $1 - 1/e$ of the global optimum, providing a worst-case performance bound."],"supporting_citations":[{"why":"Supplies the information-theoretic definitions of entropy and code length that the derivation builds on.","marker":"[2]"},{"why":"Establishes that maximum entropy is a special case of the minimum description length criterion, anchoring the derivation.","marker":"[3]"},{"why":"Provides the minimum description length principle as the framework for model selection.","marker":"[4]"},{"why":"Discusses model selection via MDL, supporting the claim that shorter descriptions are better.","marker":"[5]"},{"why":"The original proposal of the minimax entropy principle for texture modeling, which this paper generalizes and re-derives.","marker":"[6]"},{"why":"The FRAME paper, which introduces filters, random fields, and minimax entropy as a unified texture model.","marker":"[7]"},{"why":"Jaynes' maximum entropy formalism, which defines the maximum entropy model used throughout.","marker":"[12]"}],"fun_headline_variants":["Optimal features minimize entropy in max-entropy models","Minimax entropy: choose features that minimize entropy","The best features yield the lowest-entropy max-entropy model","Feature selection reduces to entropy minimization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the experimentally measured feature averages are exact, with no sampling noise; if the data are finite or noisy, empirical fluctuations can be fit by extra features and the equivalence between low model entropy and closeness to the true distribution breaks down.","fun_headline_variants_meta":{"raw":{"variants":["Optimal features minimize entropy in max-entropy models","Minimax entropy: choose features that minimize entropy","The best features yield the lowest-entropy max-entropy model","Feature selection reduces to entropy minimization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":2939,"prompt_tokens":887,"completion_tokens":2052,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":1991}},"tokens_in":503,"tokens_out":2052,"duration_ms":14933,"temperature":1.0,"reasoning_tokens":1991,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:15:04.458913+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a system with a known ground-truth distribution, draw finite samples, compute empirical feature averages, apply the minimax entropy rule to select features, and compare the resulting model's KL divergence to the truth against a model selected with a finite-sample complexity penalty. If the minimax-selected model is consistently worse on held-out samples, the exact-averages assumption fails in realistic settings.","supporting_citations":[{"cited_title":"This algorithm, while prohibitively inefficient in most cases (see below), provides a general solution to the min- imax entropy problem","cited_arxiv_id":null,"evidence_quote":"Supplies the information-theoretic definitions of entropy and code length that the derivation builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that maximum entropy is a special case of the minimum description length criterion, anchoring the derivation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the minimum description length principle as the framework for model selection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Discusses model selection via MDL, supporting the claim that shorter descriptions are better."},{"cited_title":"Minimax Entropy Principle and Its Application to Texture Modeling,","cited_arxiv_id":null,"evidence_quote":"The original proposal of the minimax entropy principle for texture modeling, which this paper generalizes and re-derives."},{"cited_title":"A mathematical theory of commu- nication,","cited_arxiv_id":null,"evidence_quote":"The FRAME paper, which introduces filters, random fields, and minimax entropy as a unified texture model."},{"cited_title":"Information theory and statistical me- chanics,","cited_arxiv_id":null,"evidence_quote":"Jaynes' maximum entropy formalism, which defines the maximum entropy model used throughout."}],"review_version":1}