Pith. sign in

REVIEW 3 major objections 4 minor 88 references

Testing maximum entropy models with e-values

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Testing between maximum entropy models can be done with a closed-form e-variable

desk verdict The exact microcanonical GRO e-variable is a clean, genuinely new result; the canonical near-optimality claim is honestly labeled heuristic but not proven, and the paper deserves a serious referee. read the letter →

arxiv 2509.01064 v1 pith:E7AYYL45 submitted 2025-09-01 stat.ME cond-mat.stat-mechphysics.data-an

classification stat.MEcond-mat.stat-mechphysics.data-an MSC 62B1062F0362H17
keywords e-valuesgrowth-rateoptimalitymaximumentropymodelsmicrocanonicalapproximationcanonicaltestscontingencytablesexponentialfamiliesminimumdescriptionlength
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper derives an exact, closed-form expression for the growth-rate optimal e-variable when both hypotheses are microcanonical maximum entropy models (hard constraints). The same e-variable remains valid when the models are canonical (soft constraints), a setting where the true optimal e-variable is usually intractable. This makes it practical to test whether a dataset's sufficient statistics are explained by a simpler or a richer constraint set, such as whether binary data from several groups are generated by one shared probability or by group-specific probabilities. The construction is applied to 2×k contingency tables, including cases where the number of groups k grows with the sample size, a regime relevant to network and time-series models.

What carries the argument

The central object is the microcanonical GRO e-variable, constructed as a Bayes factor between a Bayesian mixture on the alternative and a Bayesian mixture on the null, where the null prior W*_0 is chosen to be the c0-marginal of the alternative mixture. The identity S(x) = (Omega0(c0)/W*_0(c0)) * (W1(c1)/Omega1(c1)) decomposes the e-variable into a combinatorial degeneracy ratio and a prior ratio, which is what makes exact computation possible. The duality fact that every microcanonical e-variable is also a canonical e-variable, because the canonical distribution averages the microcanonical conditional distribution, is what transfers the exact result to the canonical setting.

What would settle it

Run a 2×2 canonical test with independent beta priors having shape γ > 1, and fit the worst-case regret of the microcanonical approximation for m = 100 to 10,000 at interior parameter points away from the boundaries. If the fitted slope exceeds 1/2 by more than the O(1) remainder, or if the gap r fails to decay to zero, the claimed asymptotic optimality of the microcanonical approximation fails.

Watch

Extended reading notes

Core claim

For a microcanonical test with null statistic c0 and alternative statistic c1, the growth-rate optimal e-variable is S(x) = Omega0(c0(x))/Omega1(c1(x)) * W1(c1(x))/W*_0(c0(x)), where Omega_i counts configurations realizing a given constraint value, W1 is the prior on the alternative constraints, and W*_0 is the marginal distribution of c0 induced by the alternative model. The paper proves this form exactly and shows that the same variable is a valid e-variable for the corresponding canonical test, even though it is not generally the canonical GRO optimum. For canonical tests it proposes a microcanonical approximation, sandwiched between upper and lower bounds given by a pseudo approximation,

Load-bearing premise

The load-bearing premise is that the microcanonical approximation gap r shrinks to zero fast enough that the easily computed e-variable is nearly as powerful as the intractable canonical optimum; the paper supports this with heuristic reasoning and numerical experiments, not a complete proof.

Editorial extensions

If this is right

  • For 2×k contingency tables, the microcanonical GRO e-variable has an explicit or easily computed form, and when k is large the optimal null prior W*_0 is well approximated by a discrete Gaussian.
  • Because the microcanonical e-variable is a valid canonical e-variable, it can be used directly in canonical tests, with its e-power lying between two computable bounds given by the pseudo approximation.
  • Under regularity conditions, both the canonical GRO e-variable and its microcanonical approximation achieve worst-case regret (d1-d0)/2 log m + O(1), so the approximation is asymptotically near-optimal in the canonical problem.
  • The framework covers non-Bayesian universal distributions such as NML, connecting e-values to Minimum Description Length model comparison and giving code-length differences a frequentist Type-I error guarantee.
  • The same test applies to network models such as Erdős–Rényi versus stochastic block models and homogeneity tests for degree sequences, by mapping them to 2×k contingency tables.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implicit consequence is that the exact closed form extends beyond contingency tables to any pair of maximum entropy models where the alternative sufficient statistics determine the null statistic (Condition A), so a broad class of network and time-series tests inherit the same formula.
  • The paper's heuristic justification suggests a testable strengthening: if the concentration theorem could be proved for expectations of log densities rather than probabilities of sets, the microcanonical approximation would be formally asymptotically optimal; the current proof only establishes boundary-layer errors of order sqrt(log m/m) while the theorem states O(log m/m).
  • Since the pseudo approximation upper-bounds the canonical GRO e-power, iterating the high-resolution-limit construction could yield a numerical scheme to approximate the canonical optimal prior itself, beyond the two proposed bounds.
  • The numerical finding that Jeffreys-type priors (γ < 1) degrade the regret rate suggests that default priors chosen for minimax redundancy are not automatically good for e-value regret, pointing to a separate design criterion for default e-value priors.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper develops optimal e-variables (GRO e-variables) for hypothesis testing between maximum entropy models. For tests where both null and alternative are microcanonical MEMs with different sufficient statistics, it derives an exact closed-form GRO e-variable, Eq. (32), and proves directly that its null expectation is one, Eq. (33). It then shows that this microcanonical e-variable remains a valid e-variable for the corresponding canonical MEM test, Fact (40), and proposes it as an approximation to the usually intractable canonical GRO e-variable. The approximation is assessed through an interval width r between a 'pseudo' upper bound and the microcanonical lower bound, Eq. (48), with numerical evidence for 2×2 and 2×k contingency tables, and links to network and time-series models. The paper is honest in several places that the asymptotic justification of the canonical approximation is heuristic.

Significance. If the results hold, the exact microcanonical GRO formula is a genuine and useful contribution to e-value methodology for exponential-family-like models: it is explicit, simple, and directly checkable. The observation that every microcanonical e-variable is automatically a canonical e-variable is elegant and appears sound. The proposed approximation for canonical MEM tests addresses a real computational bottleneck and is supported by promising simulations. However, the advertised near-optimality of the microcanonical approximation in the canonical setting is not proven; the asymptotic argument rests on a theorem whose stated rate is not established by the supplied proof and, more fundamentally, on a mode of convergence too weak to control the quantity r. This gap affects the paper's central practical claim, though not the exact microcanonical derivation.

major comments (3)
  1. [SM S4, Theorem 1 (Eq. S27–S31)] The proof of Theorem 1 does not establish the claimed O(log m/m) rate. The boundary terms in (S29) and (S31) are bounded by exp(-m * (1/2) c k a log m / m) = m^{-c k a/2}; setting a = 1 as instructed yields m^{-c k/2}, a polynomial rate, not O(log m/m). Thus Eq. (58) is unsupported by the furnished derivation. This is load-bearing because Example C and Section III.C invoke Theorem 1 to argue that r in Eq. (48) vanishes. The theorem should either be proved at the stated rate, or restated with the rate actually established and all downstream claims adjusted accordingly.
  2. [III.C, Eqs. (48)–(49), Example C] The theoretical support for the central claim that SGRO_mic is near-optimal for canonical tests is not merely missing a rate; it is missing the right mode of convergence. Theorem 1 controls probabilities of sets of normalized sufficient statistics, whereas r in Eq. (49) is an expectation of logarithms of probability masses. Setwise convergence does not imply convergence of such log-expectations, and the manuscript explicitly concedes in Example C that 'the convergence in (58) is too weak to formally imply r → 0'. Yet the abstract, introduction, and conclusion describe the approximation as 'excellent' and 'asymptotically exact' on the basis of 'theoretical arguments'. I request that either a stronger formal result be proved under explicit regularity conditions, or the asymptotic near-optimality be explicitly presented as a conjecture supported by simulations throughout the paper, includin
  3. [S6, Eq. (S33) and Section IV.A, Figure 5] The regret bound (87), REG1 = ((d1-d0)/2) log m + O(1), is derived under the assumption that r' vanishes and that wpseudo,0 is a regular density. The paper itself notes that for beta priors with γ < 1 the convolution wpseudo is non-differentiable and that the bound fails, with Figure 5 showing fitted slopes exceeding 1/2 even on INECCSI sets. This is an important limitation of the proposed approximation for a practically relevant class of default priors, and it should be stated alongside the central claims rather than only in the discussion of the experiments.
minor comments (4)
  1. [Eq. (49)] The notation Wpseudo,0(c0(x)) is confusing: wpseudo,0 is introduced as a density on the mean-value space, while Wpseudo,0 appears to be the induced distribution on the sufficient statistic via (39). Please define the two objects explicitly and use distinct symbols throughout.
  2. [SM S4, around Eq. (S27)] The displayed chain of inequalities in (S27) appears to contain a typographical corruption: the term Qw(B)+O(log m/m) appears inside a lower bound without clear justification. This should be carefully rewritten, since the proof is otherwise hard to follow.
  3. [Section IV.B, Eq. (95)] The Gaussian approximation writes σ2_k for the variance but then uses σ_k in the exponent. Please ensure the notation is consistent, e.g. define σ_k = sqrt(sum_i Var_{W^i_1}(n^i_1)).
  4. [Figure S2] The caption says 'e-power difference' but the plotted quantity is not defined in the caption or surrounding text. Please state whether the difference is E[log SGRO_can] - E[log Sapprox] and whether it is absolute.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the exact microcanonical GRO derivation is self-contained; the canonical near-optimality gap is a heuristic rate issue, not a circular reduction.

full rationale

Score 2: no significant circularity. The exact microcanonical derivation is self-contained: W*0 in Eq. (29) is the explicit argmin of the KL objective solved via Gibbs inequality, and the closed form (32) is algebraic; the e-variable property is proved directly (Eq. 33, SM S2.B), and canonical validity follows from total expectation (SM S3.A). The canonical near-optimality claim rests on Theorem 1, but the paper itself limits this: 'the convergence in (58) is too weak to formally imply r → 0' and 'All reasoning based on Theorem 1 should thus be understood as heuristic rather than fully formal' (Section III C). The SM proof also bounds the key sup term by m^{-ck/2} (S29/S30), which does not establish the stated O(log m/m) rate for all INECCSI sets; this is a correctness/rigor gap, not a circular reduction, since the approximation is tested numerically (Figs. S2-S3, S6) and is not assumed by construction. The only self-citation of note is [36] for the NML/microcanonical-uniform equivalence ('As shown in [36], the NML microcanonical distribution is equivalent to a Bayesian distribution with a uniform prior over the sufficient statistics'), used in examples only; it does not feed the main theorem and is not load-bearing. The GRO framework from [7] is an external peer-reviewed theorem. No equation reduces to its own input.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters are fitted to data; the priors (uniform, beta, NML) are user-specified inputs to the testing problem. No new entities are postulated beyond the mathematical framework.

assumptions (5)
  • standard math Canonical MEMs are exponential families with uniform carrier; the mean-value parameterization is a smooth bijection for finite support C (steepness condition).
    Used in Section III to identify canonical models with exponential families and justify working in mean-value parameters.
  • domain assumption NML for a microcanonical model equals a Bayesian mixture with a uniform prior over the sufficient statistics.
    Invoked in Example 1 (Section IV A) citing [36]; needed to reduce NML tests to the uniform-prior case.
  • standard math Theorem 1: under a regular prior, the normalized sufficient statistic under the Bayesian marginal converges in distribution to the prior at rate O(log m/m).
    Stated as Theorem 1 and 'proved' in SM S4, but the proof as written bounds shell errors by O(sqrt(log m/m)); the O(log m/m) rate is not derived. The paper explicitly labels applications of this theorem as heuristic.
  • domain assumption The priors on alternative parameters are regular and, for the multi-group results, factorize into independent component priors; asymptotics are restricted to INECCSI subsets of the parameter space.
    Assumed throughout for redundancy, regret, and the r->0 heuristic; excludes boundaries where Jeffreys-type priors misbehave.
  • ad hoc to paper If the gap r (Eq. 48) tends to zero, the microcanonical approximation is asymptotically as powerful as the canonical GRO e-variable.
    This is the load-bearing heuristic of Sections III C and IV; not proven formally. The paper notes that for gamma<1 the associated pseudo density is not regular and the regret rate fails, so the heuristic has limited scope.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Testing maximum entropy models with e-values." pith.science (2026). https://pith.science/paper/E7AYYL45

@misc{pith2026250901064,
  author       = {Pith},
  title        = {Pith review of: Testing maximum entropy models with e-values},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E7AYYL45}},
  note         = {Machine review of arXiv:2509.01064}
}
abstract

E-values have recently emerged as a robust and flexible alternative to p-values for hypothesis testing, especially under optional continuation, i.e., when additional data from further experiments are collected. In this work, we define optimal e-values for testing between maximum entropy models, both in the microcanonical (hard constraints) and canonical (soft constraints) settings. We show that, when testing between two hypotheses that are both microcanonical, the so-called growth-rate optimal e-variable admits an exact analytical expression, which also serves as a valid e-variable in the canonical case. For canonical tests, where exact solutions are typically unavailable, we introduce a microcanonical approximation and verify its excellent performance via both theoretical arguments and numerical simulations. We then consider constrained binary models, focusing on $2 \times k$ contingency tables -- an essential framework in statistics and a natural representation for various models of complex systems. Our microcanonical optimal e-variable performs well in both settings, constituting a new tool that remains effective even in the challenging case when the number $k$ of groups grows with the sample size, as in models with growing features used for the analysis of real-world heterogeneous networks and time-series.

Figures

Figures reproduced from arXiv: 2509.01064 by the authors.

Figure 1
Figure 1. FIG. 1. The GRO e-variable [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. In the microcanonical Example A, when the prior [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 4
Figure 4. FIG. 4. Procedures to compute the microcanonical [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: FIG. 5. Fitted slope of the logarithmic growth [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6. The microcanonical GRO-optimal prior on the null [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

88 extracted references · 76 canonical work pages

  1. [1]

    For the sake of this example, we put independent, dis- crete uniform priors on the alternative parametersna 1 and nb 1: W1(na 1, nb

    = na na 1 nb nb 1 is the number of permuta- tions of x preserving the total number of 1s in each group. For the sake of this example, we put independent, dis- crete uniform priors on the alternative parametersna 1 and nb 1: W1(na 1, nb

  2. [2]

    (34) In this case, Condition A (30) holds, as the null sufficient statistics can be written as a function of the alternative one: n1 = na 1 + nb

    = 1 na + 1 1 nb + 1. (34) In this case, Condition A (30) holds, as the null sufficient statistics can be written as a function of the alternative one: n1 = na 1 + nb

  3. [3]

    Thus, the optimal prior on the null W ∗ 0 is the distribution of n1 induced by W1. In this case, that is simply the convolution of Ua and Ub, which is a triangular discrete function: W ∗ 0 (n1) =    n1+1 (na+1)(nb+1) , if 0≤n1≤min(na,nb), min(na,nb)+1 (na+1)(nb+1) , if min(na,nb)<n1≤max(na,nb), na+nb+1−n1 (na+1)(nb+1) , if max(na,nb)<n1≤na+nb...

  4. [4]

    Indeed, from Theorem 1 of [7], SGRO can is the only e-variable of that form. Moreover, from (8), it holds: DKL( ¯Pcan,1∥P w∗ 0 can,0) ≤ DKL( ¯Pcan,1∥P wpseudo,0 can,0 ) (45) or, equivalently: E ¯Pcan,1 log SGRO can ≤ E ¯Pcan,1 [log Spseudo] (46) Consequently, one has E ¯Pcan,1 log SGRO mic ≤ E ¯Pcan,1 log SGRO can ≤ E ¯Pcan,1 [log Spseudo] , (47) i.e., th...

  5. [5]

    To build it, we first transform the canonical universal distribution into a microcanonical one, by finding Wcan,1 as in (39)

    We build the corresponding microcanonical approximation, knowing that it is a valid candi- date e-variable. To build it, we first transform the canonical universal distribution into a microcanonical one, by finding Wcan,1 as in (39). Then, we compute SGRO mic according to the formulas expressed in the previous section (Equations (29) and (32))

  6. [6]

    (49) where Wpseudo,0(c0(x)) is defined as in (39)

    The goodness of the microcanonical approximation can be evaluated by looking at the width of the interval r = E ¯Pcan,1 [log Spseudo] − E ¯Pcan,1 log SGRO mic ≥ 0 (48) where, for future reference, it is useful to note that, using definitions (44) and (41) and sufficiency, we can rewrite r = E ¯Pcan,1 h log P W ∗ 0 mic,0(x) − log P wpseudo,0 can,0 (x) i = ...

  7. [7]

    Thus, we can compute the microcanonical approximation SGRO mic by us- ing the results of Example A

    Ub(nb 1). Thus, we can compute the microcanonical approximation SGRO mic by us- ing the results of Example A. As a second step, we check whether this microcanonical e-variable is a good approximation by studying the behavior of the inter- val width r as the total size n increases. To evalu- ate Spseudo, we compute the prior wpseudo,0 as described above (d...

  8. [8]

    suggest” rather than “prove

    Applying Theorem 1 to M0 with this density shows that the induced distribution V 2(m) pseudo,0 on s(m) 0 /m converges to a discretized version of w2 pseudo,0. At the same time, V ∗(m) 0 , being the convolution of a discretized w1, converges to the discretized convolution of w1. Thus, for large m, w1 pseudo,0 and w2 pseudo,0 become indistinguish- able. In ...

Show all 88 references
  1. [9]

    and zeros ( na 0 and nb

  2. [10]

    The key question is whether the probability of observing x = 1 differs between the two groups

    in each group, along with their totals, n1 and n0. The key question is whether the probability of observing x = 1 differs between the two groups. This problem translates into a hypothesis testing problem, where: • In the alternative hypothesis, the two groups are distinct, mea...

  3. [11]

    reads Pmic, 1(x; na 1, nb

  4. [12]

    =    1 Ω1(na 1 ,nb

  5. [13]

    , if (na 1(x), nb 1(x)) = (na 1, nb 1); 0, else, (64) where Ω1(na 1, nb

  6. [14]

    For any given prior W1 on the alternative sufficient statistics, SGRO mic is found exactly by computing W ∗ 0 and applying (32)

    = na na 1 nb nb 1 (65) is the number of permutations of x preserving the total number of 1s in each group. For any given prior W1 on the alternative sufficient statistics, SGRO mic is found exactly by computing W ∗ 0 and applying (32). In this case, Condition A (30) is satisfi...

  7. [15]

    If na 1 and nb 1 are independently distributed: W1(na 1, nb

    (66) Thus, following (31), the optimal prior on the null is the distribution of n0 induced by W1(na 1, nb 1). If na 1 and nb 1 are independently distributed: W1(na 1, nb

  8. [16]

    · W b 1 (nb 1), (67) then W ∗ 0 is simply the convolution of W a 1 and W b 1 : W ∗ 0 = W a 1 ∗ W b 1 , (68) where f ∗ g represents the convolution between functions f and g. Example 1: Microcanonical test with NML In the microcanonical case, resorting to the Normalized Maximum...

  9. [17]

    induced by P i, NML can,1 , which is Wcan,1(na 1, nb

  10. [18]

    (77) with W a can,1(na

  11. [19]

    · na 1 (x) na na 1 (x) 1 − na 1 (x) na na−na 1 (x) ena Γ(na,na) (na)na −1 + 1 (78) and W b can,1(nb

  12. [20]

    (79) Given that na 1 and nb 1 are independently distributed, W ∗ 0 (n1) is the convolution of W a can,1(na

    · nb 1(x) nb nb 1(x) 1 − nb 1(x) nb nb−nb 1(x) enb Γ(nb,nb) (nb)nb −1 + 1 . (79) Given that na 1 and nb 1 are independently distributed, W ∗ 0 (n1) is the convolution of W a can,1(na

  13. [21]

    Example 3: Independent beta priors

    and W b can,1(nb 1), which can be computed numerically. Example 3: Independent beta priors. The beta probability distribution reads: Beta(y; α, β) = yα−1(1 − y)β−1 B(α, β) , for y ∈ (0, 1), (80) where B(α, β) is the beta function, defined as B(α, β) = Z 1 0 tα−1(1 − t)β−1dt. (...

  14. [22]

    (83) with W a can,1(na

  15. [23]

    · B( ¯αa, ¯βa) B(αa, βa) (84) and W b can,1(nb

  16. [24]

    (85) With this choice, W a 1 and W b 1 are beta-binomial distributions

    · B( ¯αb, ¯βb) B(αb, βb) . (85) With this choice, W a 1 and W b 1 are beta-binomial distributions. Given that na 1 and nb 1 are independently distributed, as we expected because we put indepen- dent priors on pa and pb, W ∗ 0 (n1) is the convolution of W a can,1(na

  17. [25]

    Whether this expression can 15 be written in closed form depends on the specific values of the beta parameters chosen

    and W b can,1(nb 1). Whether this expression can 15 be written in closed form depends on the specific values of the beta parameters chosen. For example, if all beta parameters are equal to 1, W ∗ 0 reduces to the convolution between two discrete uniform distributions (35). Whe...

  18. [26]

    The key question remains whether the probability of observing x = 1 differs between groups

    in each group, along with their respective totals, n1 and n0. The key question remains whether the probability of observing x = 1 differs between groups. This problem again translates into a hypothesis testing problem, where: FIG. 5. Fitted slope of the logarithmic growth a lo...

  19. [27]

    (93) then W ∗ 0 is simply the convolution of the individual alternative priors: W ∗ 0 = W 1 1 ∗ ... ∗ W k 1 . (94) Interestingly, when the number of groups k is large, and the priors are regular enough, a Central Limit Theorem holds; thus, W ∗ 0 is well approximated by a discr...

  20. [28]

    k} (−1)|S| n1 + k − 1 − P j∈S(nj − n) n1 − 1 × " kY i=1 ni + 1 #−1 , (97) where the sum runs over all possible subsets of {1,

    = kY i=1 1 ni + 1, (96) the GRO null prior is again the convolution of all the in- dividual priors, i.e., the convolution of k discrete uniform distributions, which reads [41]: W ∗ 0 (n1) = = X S⊆{1, ... k} (−1)|S| n1 + k − 1 − P j∈S(nj − n) n1 − 1 × " kY i=1 ni + 1 #−1 , (97)...

  21. [29]

    (103) with W i can,1(ni

  22. [30]

    (104) W ∗ 0 is then the convolution of all W i can,1(ni 1), which again can be computed numerically or by resorting to the Gaussian approximation (95)

    · ni 1(x) ni ni 1(x) 1 − ni 1(x) ni ni−ni 1(x) eni Γ(ni,ni) (ni)ni −1 + 1 . (104) W ∗ 0 is then the convolution of all W i can,1(ni 1), which again can be computed numerically or by resorting to the Gaussian approximation (95). Example 6: Independent beta priors. Here, we ex- ...

  23. [31]

    (106) with W i can,1(ni

  24. [32]

    · B( ¯αi, ¯βi) B(αi, βi) for each i ∈ {1, . . . , k}. (107) W ∗ 0 (n1) is, then, the convolution of all W i can,1(ni 1). If all beta parameters are equal to 1, W ∗ 0 reduces to the con- volution between k discrete uniform distributions (97). When the beta parameters are such t...

  25. [33]

    r converges quickly to 0 for fixed k as the m in- creases (or, equivalently, the total sample size in- creases)

  26. [34]

    r grows slowly for n fixed and k getting bigger

  27. [35]

    thermodynamic limit

    r converges quickly to 0 whenever k and m grow together according to different power laws. The only case where r does not converge to 0 corresponds to a decreasing m as O(1/k). Our conclusion is that our microcanonical approximation SGRO mic is an optimal candidate as long as ...

  28. [36]

    J. P. A. Ioannidis. Why most published research findings are false. PLoS medicine, 2(8):e124, 2005

  29. [37]

    D. J. Benjamin et al. Redefine statistical significance. Nature Human Behaviour , 2(1):6–10, 2017

  30. [38]

    B. B. McShane, D. Gal, A. Gelman, C. Robert, and J. L. Tackett. Abandon statistical significance. The American Statistician, 73(sup1):235–245, 2019

  31. [39]

    Game-theoretic statistics and safe anytime-valid inference

    Aaditya Ramdas, Peter Gr¨ unwald, Vladimir Vovk, and Glenn Shafer. Game-theoretic statistics and safe anytime-valid inference. Statist. Sci. , 38(4):576–601, 2023

  32. [40]

    Hypothesis test- ing with e-values

    Aaditya Ramdas and Ruodu Wang. Hypothesis test- ing with e-values. Foundations and Trends in Statistics ,

  33. [41]

    Asymptotically optimal data analysis for rejecting local realism

    Yanbao Zhang, Scott Glancy, and Emanuel Knill. Asymptotically optimal data analysis for rejecting local realism. Physical Review A , 84(6):062118, 2011

  34. [42]

    Gr¨ unwald, R

    P. Gr¨ unwald, R. de Heide, and W. Koolen. Safe test- ing. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86(5):1091–1128, 2024

  35. [43]

    Wasserman, A

    L. Wasserman, A. Ramdas, and S. Balakrishnan. Uni- versal inference. Proceedings of the National Academy of Sciences, 117(29):16880–16890, 2020

  36. [44]

    Vovk and R

    V. Vovk and R. Wang. E-values: Calibration, combina- tion and applications. The Annals of Statistics , 49(3), 2021

  37. [45]

    G. Shafer. Testing by betting: A strategy for statistical and scientific communication. Journal of the Royal Statistical Society Series A: Statistics in Society , 184(2):407–431, 2021

  38. [46]

    Re- verse information projections and optimal e-statistics

    Tyron Lardy, Peter Gr¨ unwald, and Peter Harremo¨ es. Re- verse information projections and optimal e-statistics. IEEE Transactions on Information Theory, 70(11):7616– 7631, 2024

  39. [47]

    The numeraire e-variable and reverse information pro- jection

    Martin Larsson, Aaditya Ramdas, and Johannes Ruf. The numeraire e-variable and reverse information pro- jection. Annals of Statistics , 2025

  40. [48]

    J. Gibbs. Elementary Principles in Statistical Mechanics: Developed with Especial Reference to the Rational Foun- dation of Thermodynamics . Cambridge Library Collec- tion - Mathematics. Cambridge University Press, 2010

  41. [49]

    E. T. Jaynes. Information theory and statistical mechan- ics. Phys. Rev., 106:620–630, 1957

  42. [50]

    L.D. Brown. Fundamentals of Statistical Exponential Families, with Applications in Statistical Decision The- ory. Institute of Mathematical Statistics, Hayward, CA, 21 1986

  43. [51]

    Squartini and D

    T. Squartini and D. Garlaschelli. Maximum-Entropy Net- works: Pattern Detection, Network Reconstruction and Graph Combinatorics. Springer Cham, 2017

  44. [52]

    Squartini, G

    T. Squartini, G. Caldarelli, G. Cimini, A. Gabrielli, and D. Garlaschelli. Reconstruction methods for networks: The case of economic and financial systems. Physics Re- ports, 757, 2018

  45. [53]

    Cimini, T

    G. Cimini, T. Squartini, F. Saracco, D. Garlaschelli, A. Gabrielli, and G. Caldarelli. The statistical physics of real-world networks. Nature Reviews Physics, 1, 2019

  46. [54]

    Marcaccioli and G

    R. Marcaccioli and G. Livan. Correspondence be- tween temporal correlations in time series, inverse prob- lems, and the spherical model. Physical Review E , 102(1):012112, 2020

  47. [55]

    Marcaccioli and G

    R. Marcaccioli and G. Livan. Maximum entropy ap- proach to multivariate time series randomization. Sci- entific Reports, 10(1):10656, 2020

  48. [56]

    Anytime-valid confidence intervals for contingency tables and beyond

    Rosanne Turner and Peter Gr¨ unwald. Anytime-valid confidence intervals for contingency tables and beyond. Statistics and Probability Letters , 2023

  49. [57]

    Generic e-variables for exact sequential k-sample tests that allow for optional stopping

    Rosanne Turner, Alexander Ly, and Peter Gr¨ unwald. Generic e-variables for exact sequential k-sample tests that allow for optional stopping. Statistical Planning and Inference, 230:106116, 2024

  50. [58]

    E-values for exponential families: the general case

    Yunda Hao and Peter Gr¨ unwald. E-values for exponential families: the general case. arXiv: 2409.11134 , 2024

  51. [59]

    Bar Lev, and Martijn de Jong

    Peter Gr¨ unwald, Tyron Lardy, Yunda Hao, Shaul K. Bar Lev, and Martijn de Jong. Optimal e-values for exponen- tial families: the simple case. arXiv: 2404.19465 , 2024

  52. [60]

    Gr¨ unwald

    Peter D. Gr¨ unwald. Beyond neyman–pearson: E- values enable hypothesis testing with a data-driven al- pha. Proceedings of the National Academy of Sciences , 121(39):e2302098121, 2024

  53. [61]

    Zhang, A

    Z. Zhang, A. Ramdas, and R. Wang. On the existence of powerful p-values and e-values for composite hypotheses. arXiv: 2305.16539 , 2024

  54. [62]

    Q. Wang, R. Wang, and J. Ziegel. E-backtesting. arXiv: 2209.00991, 2024

  55. [63]

    Vovk and R

    V. Vovk and R. Wang. Efficiency of nonparametric e- tests. arXiv: 2208.08925 , 2024

  56. [64]

    The philosophy of Bayes factors and the quan- tification of statistical evidence

    Richard D Morey, Jan-Willem Romeijn, and Jeffrey N Rouder. The philosophy of Bayes factors and the quan- tification of statistical evidence. Journal of Mathematical Psychology, 72(6-18):36, 2016

  57. [65]

    Gr¨ unwald.The minimum description length principle

    P. Gr¨ unwald.The minimum description length principle . MIT press, 2007

  58. [66]

    Gr¨ unwald and T

    P. Gr¨ unwald and T. Roos. Minimum description length revisited. International journal of mathematics for in- dustry, 11(01):1930001, 2019

  59. [67]

    Yamanishi

    K. Yamanishi. Learning with the Minimum Description Length Principle. Springer Nature Singapore, 2023

  60. [68]

    Li and P

    M. Li and P. Vit´ anyi. An Introduction to Kolmogorov Complexity and Its Applications . Springer, New York, 3rd edition, 2008

  61. [69]

    Tighter pac-bayes bounds through coin-betting

    Kyoungseok Jang, Kwang-Sung Jun, Ilja Kuzborskij, and Francesco Orabona. Tighter pac-bayes bounds through coin-betting. In Proceedings COLT 2023, 2023

  62. [70]

    Barndorff-Nielsen

    O.E. Barndorff-Nielsen. Information and Exponential Families in Statistical Theory . Wiley, Chichester, UK, 1978

  63. [71]

    Giuffrida, T

    F. Giuffrida, T. Squartini, P. Gr¨ unwald, and D. Gar- laschelli. Description length of canonical and microcanonical models. Phys. Rev. Res. , 2025

  64. [72]

    Commu- nity detection in interval-weighted networks

    H´ elder Alves, Paula Brito, and Pedro Campos. Commu- nity detection in interval-weighted networks. Data Min- ing and Knowledge Discovery , 38(2):653–698, 2024

  65. [73]

    Max Jerdee, Alec Kirkley, and Mark E. J. Newman. Mu- tual information and the encoding of contingency tables. Physical review. E , 110 6-1:064306, 2024

  66. [74]

    P. P. A. Staniczenko, M.J. Smith, and S. Allesina. Se- lecting food web models using normalized maximum like- lihood. Methods in Ecology and Evolution , 5(6):551–562, 2014

  67. [75]

    Clarke and A.R

    B.S. Clarke and A.R. Barron. Jeffreys’ prior is asymp- totically least favorable under entropy risk. Journal of Statistical Planning and Inference , 41:37–60, 1994

  68. [76]

    Earnest (user)

    M. Earnest (user). Extended stars-and-bars problem(where the upper limit of the variable is bounded). Mathematics Stack Exchange. URL:https://math.stackexchange.com/q/3182858 (version: 2019-04-14)

  69. [77]

    Maximum Entropy Econometrics: Robust Estimation with Limited Data

    A Golan, G Judge, and D Miller. Maximum Entropy Econometrics: Robust Estimation with Limited Data . John Wiley), Chichester, UK, 1996

  70. [78]

    P. W. Holland and S. Leinhardt. An exponential family of probability distributions for directed graphs. Journal of the American Statistical Association , 76(373):33–50, 1981

  71. [79]

    D. R. Hunter and M. S. Handcock. Inference in curved exponential family models for networks. Journal of Com- putational and Graphical Statistics , 15(3):565–583, 2006

  72. [80]

    T. P. Peixoto. Nonparametric bayesian inference of the microcanonical stochastic block model. Phys. Rev. E , 95:012317, 2017

  73. [81]

    Tiago P. Peixoto. Network reconstruction via the minimum description length principle. Phys. Rev. X , 15:011065, Mar 2025

  74. [82]

    Vall` es-Catal` a, T

    T. Vall` es-Catal` a, T. P. Peixoto, M. Sales-Pardo, and R. Guimer` a. Consistencies and inconsistencies between model selection and link prediction in networks. Phys. Rev. E, 97:062316, 2018

  75. [83]

    Strong ensemble nonequivalence in systems with local constraints

    Qi Zhang and Diego Garlaschelli. Strong ensemble nonequivalence in systems with local constraints. New Journal of Physics , 24(4):043011, 2022

  76. [84]

    Breaking of ensemble equivalence in networks

    Tiziano Squartini, Joey de Mol, Frank den Hollander, and Diego Garlaschelli. Breaking of ensemble equivalence in networks. Physical review letters , 115(26):268701, 2015

  77. [85]

    Csisz´ ar

    I. Csisz´ ar. Sanov property, generalized I-projection and a conditional limit theorem. The Annals of Probability , pages 768–793, 1984

  78. [86]

    log SGRO(θ1) − log ¯P1(x) P w∗ 0 0 (x) # = max θ1∈Θ′ 1 min w′ 0∈Wθ0 Eθ1

    Peter Gr¨ unwald, Yunda Hao, and Akshay Balsubra- mani. Growth-optimal e-variables and an extension to the multivariate Csisz´ ar-Sanov-Chernoff theorem.arXiv: 2412.17554, 2024. 1 Supplementary Materials S1. REDUNDANCY AND REGRET Here we show that, given the null and alternati...

  79. [87]

    Given a canonical universal distribution ¯Pcan relative to a sufficient statistic c, we can always define a prior distribution Wcan(c) on the sufficient statistic such that ¯Pcan = ¯P Wcan mic . (S21)

  80. [88]

    log P wpseudo,0 0 (x) P ˜w′ 0 0 (x) # + RED1(θ1; P w1 1 ) = Eθ1

    The opposite is not true: for some ¯P W mic, there is no choice of prior density w(θ) such that ¯P W mic = ¯P w can. 4 Proof 1 By construction, a canonical universal distribution ¯Pcan(x) relative to the sufficient statistics c always assigns the same probability mass to confi...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.