{"id":"41fcec03-5f08-4484-9b83-abef67b48b4a","arxiv_id":"2509.01064","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"It derives an exact growth-rate optimal e-variable for microcanonical maximum entropy tests and shows this e-variable is a valid, near-optimal approximation for canonical tests.","lead":"This paper constructs e-values for testing between maximum entropy models, giving an exact formula when models fix the observed statistics and an approximation when they only fix averages. The method applies to contingency tables and network models, including cases where the number of groups grows with the data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Canonical near-optimality of the microcanonical e-variable rests on r -> 0, which is unproven: Theorem 1's proof does not establish O(log m/m), and the paper explicitly labels the justification heuristic.","rationale":"The reader's CONDITIONAL verdict is exactly right. The strongest claim in the paper—the exact microcanonical e-variable formula and its validity as a canonical e-variable—is mathematically sound. The weak spot is the asymptotic justification of the microcanonical approximation for canonical tests: Theorem 1 does not imply r → 0, and its proof in SM S4 does not even establish the stated O(log m/m) rate, since the boundary terms are m^{−c k/2}. The paper honestly labels this reasoning heuristic, and the numerical evidence, while suggestive, has no error bars and shows some regimes (γ = 0.5) where regret slopes exceed the expected 0.5 at the sample sizes considered. This does not warrant rejection—the exact microcanonical contribution stands and the approximation is testable—but it does warrant conditioning acceptance on either a proof of r → 0 under explicit regularity conditions or a careful scoping of the claims to what the numerics actually support. Since the reader already reached CONDITIONAL for the same reason, no verdict change is needed.","tokens_in":984,"tokens_out":909,"duration_ms":113087,"concrete_test":"For the 2×2 canonical test with na = nb = m and independent Beta(0.5, 0.5) priors (the Jeffreys case, where the paper itself finds anomalous regret slopes in Figure 5), compute r_m in Eq. (49) for m ∈ {10^3, 10^4, 10^5, 10^6} by exact convolution for W*_0 and a high-resolution pseudo prior (resolution scale ≫ m). Fit log r_m against log m: if r_m does not vanish, or the estimated slope is not consistent with convergence to 0, the microcanonical approximation fails to be near-optimal in a regime the paper claims to cover; if it does vanish, the heuristic is empirically supported and only the formal-rate issue remains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The exact microcanonical derivation (Eqs. 29–33) and the fact that SGRO_mic is a valid canonical e-variable (Fact (40), via the total-expectation argument in SM S3.A) are sound. The load-bearing gap is the paper's core practical claim that SGRO_mic is near-optimal for canonical MEM tests, i.e. that r in Eq. (48) converges to 0. The only theoretical support is Theorem 1, but Theorem 1 controls probabilities of sets of normalized sufficient statistics, whereas r is an expectation of log-probability-masses (Eq. (49)); setwise convergence does not imply convergence of such log-expectations, especially when priors are unbounded (e.g. γ < 1) or when W*_0 has atoms. The paper itself concedes this: 'the convergence in (58) is too weak to formally imply r → 0' and 'All reasoning based on Theorem 1 should thus be understood as heuristic rather than fully formal.' Moreover, the proof of Theorem 1 in SM S4 only bounds the boundary terms via (S29)/(S31) by exp(−m · c k a log m / (2m)) = m^{−c k a/2}; setting a = 1 gives a polynomial rate m^{−c k/2}, not the stated O(log m/m). So the theorem's stated rate is not established by the supplied proof. Thus the microcanonical approximation's advertised near-optimality for canonical tests is not proven; it currently rests on the reported numerics and on a heuristic. This does not affect the exact microcanonical result, but it does affect the paper's central practical recommendation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops optimal e-variables (GRO e-variables) for hypothesis testing between maximum entropy models. For tests where both null and alternative are microcanonical MEMs with different sufficient statistics, it derives an exact closed-form GRO e-variable, Eq. (32), and proves directly that its null expectation is one, Eq. (33). It then shows that this microcanonical e-variable remains a valid e-variable for the corresponding canonical MEM test, Fact (40), and proposes it as an approximation to the usually intractable canonical GRO e-variable. The approximation is assessed through an interval width r between a 'pseudo' upper bound and the microcanonical lower bound, Eq. (48), with numerical evidence for 2×2 and 2×k contingency tables, and links to network and time-series models. The paper is honest in several places that the asymptotic justification of the canonical approximation is heuristic.","tokens_in":33976,"tokens_out":5076,"duration_ms":68257,"significance":"If the results hold, the exact microcanonical GRO formula is a genuine and useful contribution to e-value methodology for exponential-family-like models: it is explicit, simple, and directly checkable. The observation that every microcanonical e-variable is automatically a canonical e-variable is elegant and appears sound. The proposed approximation for canonical MEM tests addresses a real computational bottleneck and is supported by promising simulations. However, the advertised near-optimality of the microcanonical approximation in the canonical setting is not proven; the asymptotic argument rests on a theorem whose stated rate is not established by the supplied proof and, more fundamentally, on a mode of convergence too weak to control the quantity r. This gap affects the paper's central practical claim, though not the exact microcanonical derivation.","major_comments":[{"comment":"The proof of Theorem 1 does not establish the claimed O(log m/m) rate. The boundary terms in (S29) and (S31) are bounded by exp(-m * (1/2) c k a log m / m) = m^{-c k a/2}; setting a = 1 as instructed yields m^{-c k/2}, a polynomial rate, not O(log m/m). Thus Eq. (58) is unsupported by the furnished derivation. This is load-bearing because Example C and Section III.C invoke Theorem 1 to argue that r in Eq. (48) vanishes. The theorem should either be proved at the stated rate, or restated with the rate actually established and all downstream claims adjusted accordingly.","section":"SM S4, Theorem 1 (Eq. S27–S31)"},{"comment":"The theoretical support for the central claim that SGRO_mic is near-optimal for canonical tests is not merely missing a rate; it is missing the right mode of convergence. Theorem 1 controls probabilities of sets of normalized sufficient statistics, whereas r in Eq. (49) is an expectation of logarithms of probability masses. Setwise convergence does not imply convergence of such log-expectations, and the manuscript explicitly concedes in Example C that 'the convergence in (58) is too weak to formally imply r → 0'. Yet the abstract, introduction, and conclusion describe the approximation as 'excellent' and 'asymptotically exact' on the basis of 'theoretical arguments'. I request that either a stronger formal result be proved under explicit regularity conditions, or the asymptotic near-optimality be explicitly presented as a conjecture supported by simulations throughout the paper, includin","section":"III.C, Eqs. (48)–(49), Example C"},{"comment":"The regret bound (87), REG1 = ((d1-d0)/2) log m + O(1), is derived under the assumption that r' vanishes and that wpseudo,0 is a regular density. The paper itself notes that for beta priors with γ < 1 the convolution wpseudo is non-differentiable and that the bound fails, with Figure 5 showing fitted slopes exceeding 1/2 even on INECCSI sets. This is an important limitation of the proposed approximation for a practically relevant class of default priors, and it should be stated alongside the central claims rather than only in the discussion of the experiments.","section":"S6, Eq. (S33) and Section IV.A, Figure 5"}],"minor_comments":[{"comment":"The notation Wpseudo,0(c0(x)) is confusing: wpseudo,0 is introduced as a density on the mean-value space, while Wpseudo,0 appears to be the induced distribution on the sufficient statistic via (39). Please define the two objects explicitly and use distinct symbols throughout.","section":"Eq. (49)"},{"comment":"The displayed chain of inequalities in (S27) appears to contain a typographical corruption: the term Qw(B)+O(log m/m) appears inside a lower bound without clear justification. This should be carefully rewritten, since the proof is otherwise hard to follow.","section":"SM S4, around Eq. (S27)"},{"comment":"The Gaussian approximation writes σ2_k for the variance but then uses σ_k in the exponent. Please ensure the notation is consistent, e.g. define σ_k = sqrt(sum_i Var_{W^i_1}(n^i_1)).","section":"Section IV.B, Eq. (95)"},{"comment":"The caption says 'e-power difference' but the plotted quantity is not defined in the caption or surrounding text. Please state whether the difference is E[log SGRO_can] - E[log Sapprox] and whether it is absolute.","section":"Figure S2"}],"recommendation":"major_revision","confidential_remarks":"The exact microcanonical contribution (Eq. 32 and the E0[S]=1 proof) is sound and worth publishing. The main risk is that the paper's headline claim about the canonical approximation overstates what is currently proven. The authors already disclose the heuristic nature in the main text, but the abstract and conclusion are not aligned with that disclosure. A major revision that either supplies a stronger proof for r->0 under stated conditions, or systematically downgrades the asymptotic claim to a conjecture, would make the paper acceptable. I do not see the gap as fatal, since the exact result stands and the numerics are informative."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the exact microcanonical GRO e-variable is the real contribution, and it is solid. The canonical-side claims are softer than the abstract suggests, but the paper knows it and says so.\n\nWhat is new: a closed-form expression for the growth-rate optimal e-variable when testing between two microcanonical maximum entropy models with different sufficient statistics (Eq. 32). The derivation is self-contained, the proof that E0[S] = 1 is direct, and the fact that the same variable is also a valid e-variable for the corresponding canonical test follows cleanly from the law of total expectation. This genuinely extends prior e-variable work on exponential families, which either handled 2x2 tables with sequential e-variables or considered the same sufficient statistic with different carriers.\n\nThe paper also deserves credit for honesty. It explicitly labels the asymptotic justification for the microcanonical approximation as heuristic, and that label is accurate. The stress-test note is correct: the proof of Theorem 1 does not establish the stated O(log m/m) rate; the boundary arguments yield only a polynomial bound, and more importantly the convergence in (58) concerns probabilities of sets, which is too weak to formally imply r -> 0 for the expectation of log-densities in Eq. (49). So the advertised near-optimality of the microcanonical approximation for canonical tests is not proven; it rests on the numerics plus a heuristic. That matters because if the gap r does not vanish, the approximation could be substantially less powerful than the canonical GRO e-variable.\n\nMinor soft spots: the abstract and the 2xk section omit the per-group-size caveat — the approximation works when each group size m is large, not merely when the total sample size grows. The Jeffreys-prior suboptimality finding is suggestive and plausible, but the simulations lack error bars. Also, the pseudo approximation is not an e-variable except in special cases, which the paper states correctly.\n\nWho this is for: e-value theorists and network scientists who use MEMs as null models. The exact microcanonical result is a genuinely useful tool and likely to be cited. The canonical approximation is reasonable for practitioners who check r numerically, but it should be presented as a heuristic, not a theorem.\n\nRecommendation: send it to peer review. The main theorem's rate claim needs to be corrected or removed, and the canonical near-optimality section should be reframed as conjecture supported by numerics. The exact microcanonical part earns the referee time.","headline":"The exact microcanonical GRO e-variable is a clean, genuinely new result; the canonical near-optimality claim is honestly labeled heuristic but not proven, and the paper deserves a serious referee.","tokens_in":34424,"tokens_out":1483,"would_cite":true,"duration_ms":19546,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62B10","62F03","62H17"],"pacs":[],"model":"deepseek-v4-flash","headline":"Testing between maximum entropy models can be done with a closed-form e-variable","keywords":["e-values","growth-rate optimality","maximum entropy models","microcanonical approximation","canonical tests","contingency tables","exponential families","minimum description length"],"falsifier":"Run a 2×2 canonical test with independent beta priors having shape γ > 1, and fit the worst-case regret of the microcanonical approximation for m = 100 to 10,000 at interior parameter points away from the boundaries. If the fitted slope exceeds 1/2 by more than the O(1) remainder, or if the gap r fails to decay to zero, the claimed asymptotic optimality of the microcanonical approximation fails.","tokens_in":33463,"feed_emoji":"📊","tokens_out":6294,"duration_ms":71130,"temperature":0.7,"pith_summary":"The paper derives an exact, closed-form expression for the growth-rate optimal e-variable when both hypotheses are microcanonical maximum entropy models (hard constraints). The same e-variable remains valid when the models are canonical (soft constraints), a setting where the true optimal e-variable is usually intractable. This makes it practical to test whether a dataset's sufficient statistics are explained by a simpler or a richer constraint set, such as whether binary data from several groups are generated by one shared probability or by group-specific probabilities. The construction is applied to 2×k contingency tables, including cases where the number of groups k grows with the sample size, a regime relevant to network and time-series models.","feed_headline":"Exact e-variable found for maximum entropy model tests","feed_subtitle":"The closed-form microcanonical e-variable also works for canonical models, making group-comparison tests practical.","key_machinery":"The central object is the microcanonical GRO e-variable, constructed as a Bayes factor between a Bayesian mixture on the alternative and a Bayesian mixture on the null, where the null prior W*_0 is chosen to be the c0-marginal of the alternative mixture. The identity S(x) = (Omega0(c0)/W*_0(c0)) * (W1(c1)/Omega1(c1)) decomposes the e-variable into a combinatorial degeneracy ratio and a prior ratio, which is what makes exact computation possible. The duality fact that every microcanonical e-variable is also a canonical e-variable, because the canonical distribution averages the microcanonical conditional distribution, is what transfers the exact result to the canonical setting.","core_discovery":"For a microcanonical test with null statistic c0 and alternative statistic c1, the growth-rate optimal e-variable is S(x) = Omega0(c0(x))/Omega1(c1(x)) * W1(c1(x))/W*_0(c0(x)), where Omega_i counts configurations realizing a given constraint value, W1 is the prior on the alternative constraints, and W*_0 is the marginal distribution of c0 induced by the alternative model. The paper proves this form exactly and shows that the same variable is a valid e-variable for the corresponding canonical test, even though it is not generally the canonical GRO optimum. For canonical tests it proposes a microcanonical approximation, sandwiched between upper and lower bounds given by a pseudo approximation,","pith_inferences":["An implicit consequence is that the exact closed form extends beyond contingency tables to any pair of maximum entropy models where the alternative sufficient statistics determine the null statistic (Condition A), so a broad class of network and time-series tests inherit the same formula.","The paper's heuristic justification suggests a testable strengthening: if the concentration theorem could be proved for expectations of log densities rather than probabilities of sets, the microcanonical approximation would be formally asymptotically optimal; the current proof only establishes boundary-layer errors of order sqrt(log m/m) while the theorem states O(log m/m).","Since the pseudo approximation upper-bounds the canonical GRO e-power, iterating the high-resolution-limit construction could yield a numerical scheme to approximate the canonical optimal prior itself, beyond the two proposed bounds.","The numerical finding that Jeffreys-type priors (γ < 1) degrade the regret rate suggests that default priors chosen for minimax redundancy are not automatically good for e-value regret, pointing to a separate design criterion for default e-value priors."],"forward_implications":["For 2×k contingency tables, the microcanonical GRO e-variable has an explicit or easily computed form, and when k is large the optimal null prior W*_0 is well approximated by a discrete Gaussian.","Because the microcanonical e-variable is a valid canonical e-variable, it can be used directly in canonical tests, with its e-power lying between two computable bounds given by the pseudo approximation.","Under regularity conditions, both the canonical GRO e-variable and its microcanonical approximation achieve worst-case regret (d1-d0)/2 log m + O(1), so the approximation is asymptotically near-optimal in the canonical problem.","The framework covers non-Bayesian universal distributions such as NML, connecting e-values to Minimum Description Length model comparison and giving code-length differences a frequentist Type-I error guarantee.","The same test applies to network models such as Erdős–Rényi versus stochastic block models and homogeneity tests for degree sequences, by mapping them to 2×k contingency tables."],"supporting_citations":[{"why":"Supplies the definition of growth-rate optimal e-variables and the theorem that the optimal e-variable is a ratio of Bayesian marginal likelihoods with a special prior on the null.","marker":"[7]"},{"why":"Formalizes the misspecified redundancy bound used in the supplementary regret calculation and frames the contrast with exponential families having the same sufficient statistic but different carriers.","marker":"[23]"},{"why":"Provides the universal-distribution and redundancy framework used to justify low-regret e-variables and the (d/2) log m redundancy bounds.","marker":"[30]"},{"why":"Underpins the one-to-one mapping between canonical parameters and mean-value parameters and the steepness condition used throughout the paper.","marker":"[35]"},{"why":"Establishes that the NML microcanonical distribution equals a Bayesian uniform-prior distribution, the basis for treating microcanonical Bayesian mixtures as universal distributions.","marker":"[36]"},{"why":"Provides the closed-form convolution formula for k discrete uniform distributions used to compute W*_0 in the k-groups example.","marker":"[41]"},{"why":"Supplies a multivariate Chernoff concentration inequality used in the proof of Theorem 1 to control deviations of normalized sufficient statistics.","marker":"[50]"},{"why":"Extends the Chernoff-type bound used in the same proof to the multivariate setting needed for the theorem.","marker":"[51]"}],"fun_headline_variants":["Exact e-value formula for max entropy tests","Closed-form e-variable for max entropy models","Microcanonical e-variable works for canonical tests","Exact test for max entropy: microcanonical e-value","New e-variable for max entropy models, exact and practical"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the microcanonical approximation gap r shrinks to zero fast enough that the easily computed e-variable is nearly as powerful as the intractable canonical optimum; the paper supports this with heuristic reasoning and numerical experiments, not a complete proof.","fun_headline_variants_meta":{"raw":{"variants":["Exact e-value formula for max entropy tests","Closed-form e-variable for max entropy models","Microcanonical e-variable works for canonical tests","Exact test for max entropy: microcanonical e-value","New e-variable for max entropy models, exact and practical"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1095,"prompt_tokens":750,"completion_tokens":345,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":279}},"tokens_in":494,"tokens_out":345,"duration_ms":4570,"temperature":1.0,"reasoning_tokens":279,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:55:38.488154+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a 2×2 canonical test with independent beta priors having shape γ > 1, and fit the worst-case regret of the microcanonical approximation for m = 100 to 10,000 at interior parameter points away from the boundaries. If the fitted slope exceeds 1/2 by more than the O(1) remainder, or if the gap r fails to decay to zero, the claimed asymptotic optimality of the microcanonical approximation fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Formalizes the misspecified redundancy bound used in the supplementary regret calculation and frames the contrast with exponential families having the same sufficient statistic but different carriers."},{"cited_title":"(104) W ∗ 0 is then the convolution of all W i can,1(ni 1), which again can be computed numerically or by resorting to the Gaussian approximation (95)","cited_arxiv_id":null,"evidence_quote":"Provides the universal-distribution and redundancy framework used to justify low-regret e-variables and the (d/2) log m redundancy bounds."},{"cited_title":"thermodynamic limit","cited_arxiv_id":null,"evidence_quote":"Underpins the one-to-one mapping between canonical parameters and mean-value parameters and the steepness condition used throughout the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that the NML microcanonical distribution equals a Bayesian uniform-prior distribution, the basis for treating microcanonical Bayesian mixtures as universal distributions."},{"cited_title":"Asymptotically optimal data analysis for rejecting local realism","cited_arxiv_id":null,"evidence_quote":"Provides the closed-form convolution formula for k discrete uniform distributions used to compute W*_0 in the k-groups example."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies a multivariate Chernoff concentration inequality used in the proof of Theorem 1 to control deviations of normalized sufficient statistics."},{"cited_title":"Squartini and D","cited_arxiv_id":null,"evidence_quote":"Extends the Chernoff-type bound used in the same proof to the multivariate setting needed for the theorem."}],"review_version":1}