{"id":"2dad762b-818a-4a5a-a987-257dbe80bdef","arxiv_id":"2505.03582","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors derive a fixed-point iteration for MLE in λ-exponential families and prove the log-likelihood increases monotonically along its iterates.","lead":"This paper presents an iterative recipe for maximum likelihood estimation in λ-exponential families, a broad class that includes q-Gaussian and Dirichlet perturbation distributions. Each step of the recipe is guaranteed to improve the fit, which makes the method attractive for practitioners in information geometry and compositional data analysis.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The key inequality (21) in Theorem 1's proof has the wrong direction, so the central monotonicity claim is not established as printed.","rationale":"The reader's verdict is CONDITIONAL, and I agree that the paper should not be accepted without fixes. My most load-bearing concern is the sign error in inequality (21), which is the pivotal step of the proof of Theorem 1. The reader's rationale explicitly flags this sign error, but the reader's stated weakest assumption is Assumption 1 rather than the proof's key inequality. The sign error is not merely cosmetic: the proof first claims A < B and then, in the following displayed inequality, effectively uses B < A; the AM-GM step only works with the reversed direction. Because strict concavity gives the opposite inequality from what is printed, the proof as written is invalid, although a simple correction likely restores the intended argument. This warrants the same CONDITIONAL verdict: the central claim is plausible and likely repairable, but the printed proof does not yet establish it. Assumption 1 also has a genuine type problem in its containment condition, but it does not by itself invalidate the main theorem if interpreted as F(S) ⊂ Ξ. Proposition 2 is explicitly unproved, but it is an illustration and not the central claim. Overall, my read does not change the reader's verdict; it sharpens the reason: the decisive defect is the reversed inequality in the proof of Theorem 1.","tokens_in":6829,"tokens_out":17353,"duration_ms":162736,"concrete_test":"Re-derive (20)–(21) without invoking the printed direction: compute A = Σ_i w_i(θ^(k)) κ_i(η^(k)) and B = Σ_i w_i(θ^(k)) κ_i(η^(k+1)) directly from (20), and check which of A < B or B < A follows from strict concavity of Ψ. Then verify whether the AM-GM conclusion requires the average ratio (1/n)Σ κ_i(η^(k+1))/κ_i(η^(k)) to be less than 1 or greater than 1, and whether the corrected direction still implies Π κ_i(η^(k+1)) < Π κ_i(η^(k)). If the corrected sign yields the desired product inequality, the central argument survives with a typographical fix; if not, Theorem 1 is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1 is the paper's central claim, and its proof hinges on inequality (21). From (20), the two sides are A = Σ_i w_i(θ^(k)) κ_i(η^(k)) and B = Σ_i w_i(θ^(k)) κ_i(η^(k+1)). By (20), A = Ψ(η^(k)) + ∇Ψ(η^(k))·(η^(k+1)−η^(k)), while B = Ψ(η^(k+1)). Since Ψ = e^{λψ} is strictly concave on Ξ when λ<0, the tangent line at η^(k) lies above the graph, so Ψ(η^(k+1)) ≤ Ψ(η^(k)) + ∇Ψ(η^(k))·(η^(k+1)−η^(k)), i.e. B ≤ A. This is the reverse of the strict inequality claimed in (21). Interestingly, the next displayed inequality in the proof, and the AM-GM step that follows, use the reversed direction B < A (equivalently (1/n)Σ κ_i(η^(k+1))/κ_i(η^(k)) < 1). Thus the argument is internally inconsistent as printed: it asserts A < B, then derives the consequence of B < A. The theorem may be repairable by changing (21) to the correct direction and adjusting the comparison, but the printed proof does not validly establish monotonicity. This is a load-bearing defect in the central claim. Separately, the paper explicitly leaves Proposition 2 unproved, but that concerns an illustrative example rather than the main theorem. Assumption 1 also states that Ξ contains the common support S; as written S is a subset of the sample space X, while Ξ is a subset of the dual parameter space, so the assumption is ill-typed for the q-Gaussian example and should presumably read F(S) ⊂ Ξ (or its closure).","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a fixed-point algorithm for maximum likelihood estimation in the λ-exponential family, using the dual parameterization provided by λ-duality. The main result, Theorem 1, asserts that under Assumption 1 and for λ<0, the log-likelihood strictly increases at every iterate unless the current parameter already satisfies the first-order fixed-point condition. The method is illustrated on the q-Gaussian distribution and the Dirichlet perturbation model, and Proposition 2 states that the Dirichlet perturbation update is independent of λ.","tokens_in":7170,"tokens_out":24935,"duration_ms":217103,"significance":"If the result is correct, the paper contributes a computationally useful and theoretically interesting MM-type algorithm for the λ-exponential family, and it demonstrates a nontrivial application of λ-duality to estimation. The core idea, using strict concavity of Ψ=e^{λψ} to obtain a monotonicity argument, is elegant and is likely repairable. However, as printed, the proof of the main theorem contains an inconsistent inequality and several auxiliary statements are unproved or contain sign errors. These issues are substantial but appear to be local rather than fatal to the intended result.","major_comments":[{"comment":"In Eq. (21) the strict inequality has the wrong direction. With A=Σ_i w_i(θ^{(k)})κ_i(η^{(k)}) and B=Σ_i w_i(θ^{(k)})κ_i(η^{(k+1)}), Eq. (20) gives A=Ψ(η^{(k)})+∇Ψ(η^{(k)})·(η^{(k+1)}−η^{(k)}) and B=Ψ(η^{(k+1)}). Since Ψ=e^{λψ} is strictly concave on Ξ, the tangent line at η^{(k)} lies above the graph, so B<A. The paper asserts A<B, while the following displayed inequality, namely (1/n)Σ_i κ_i(η^{(k+1)})/κ_i(η^{(k)})<1, is the consequence of B<A. The printed proof is therefore internally inconsistent; the argument can be repaired by reversing the inequality in (21) and then applying the AM-GM step as written.","section":"§2, Theorem 1 proof, Eq. (21)"},{"comment":"The passage from the definition of κ_i to κ_i=p(x_i;θ)^λ is not justified as printed. From (15) and (16) one obtains 1+λθ·y_i = κ_i(η)(1+λθ·η)/Ψ(η), not κ_i(η)/(Ψ(η)(1+λθ·η)) as in the first fraction of (17). The final identity κ_i=p_i^λ becomes correct only if the Fenchel–Young relation used is φ(θ)+ψ(η)=(1/λ)log(1+λθ·η), but (7) states a relation involving φ_λ and ψ_λ, which are never defined. If φ_λ and ψ_λ are intended to mean (e^{λφ}−1)/λ and (e^{λψ}−1)/λ, the identity does not match the c-conjugate defined in (6), and the derivation of (18) does not follow as written. Please define the notation and correct (17) and (7).","section":"§2, Eqs. (7), (17), (18)"},{"comment":"Assumption 1 is ill-typed: S is a subset of the sample space X, while Ξ is a subset of R^d, so the statement that Ξ contains S cannot hold for the examples. The intended condition is presumably F(S)⊂Ξ, possibly with closure or with the convex hull of F(S) contained in Ξ; this is needed so that the iterates η^{(k+1)} remain in Ξ and the concavity of Ψ can be applied. Moreover, Example 1 contains sign errors: with φ(θ)=−1/2 log(−θ), definition (5) gives η=−1/[(2+λ)θ]∈(0,∞) and θ=−1/[(2+λ)η], not η=1/(2+λ)·1/θ, θ=1/(2+λ)·1/η, and Ξ=(0,∞), not (−∞,0). These errors affect the correctness of Figure 1 and the claimed verification of Assumption 1.","section":"§2, Assumption 1 and Example 1"},{"comment":"Proposition 2 is stated without proof; the proof is dismissed with 'details are omitted due to space constraints.' This is a substantive claim about the algorithm's update being independent of λ, and it is advertised as an analogue of the normal location model. The proof should be included, or the statement should be explicitly labeled as a conjecture or empirical observation, so that the result can be verified.","section":"§3, Proposition 2"}],"minor_comments":[{"comment":"The notation φ_λ and ψ_λ in Eq. (7) is inconsistent with the notation φ^{(λ)} used in (3) and (6); please define these symbols or use φ and ψ throughout.","section":"§2, Eq. (7)"},{"comment":"The indexing in the sentence defining (η^{(k)})_{k≥1} and (θ^{(0)})_{k≥0} is garbled; it should be (η^{(k)})_{k≥0} and (θ^{(k)})_{k≥0} after assigning an initial value.","section":"§2, Algorithm 1"},{"comment":"The theorem statement uses θ^{(k)} with a hat in some places and without in others; please unify the notation.","section":"§2, Theorem 1 statement"},{"comment":"The statement that the regularity conditions 'can be verified' for the examples is unsupported; please provide at least a sketch of the verification, especially for the q-Gaussian example given the sign issues in Example 1.","section":"§2, Assumption 1"},{"comment":"If the corrected formulas imply η>0, the horizontal axis of Figure 1, which runs from −5 to 0, is inconsistent with those formulas; the figure should be regenerated or the parametrization clarified.","section":"§2, Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The central algorithmic idea is sound, and a corrected version of the proof of Theorem 1 appears to be short and valid. However, the manuscript contains an unusual number of sign, notation, and algebra errors in equations that are central to the exposition, including the main theorem's proof and the q-Gaussian example. I recommend that the editor ask for a careful revision that addresses the specific errors listed in the report before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The fixed-point iteration in Algorithm 1 is a natural and apparently new way to compute MLEs for λ-exponential families, and the monotonicity theorem is probably true. But the proof as printed is not valid: the key inequality (21) has the wrong direction, and the examples do not verify the standing Assumption 1. Both are fixable, but a reader cannot currently trust the central claim as written.\n\nWhat's genuinely new: the update rule follows directly from the first-order condition, and the proof strategy via λ-duality is distinct from the bound-optimization approach in [9]. The observation that the Dirichlet perturbation update is λ-independent (Prop. 2) is nice and potentially useful. The paper is clearly structured and the relation between the weights and the dual potential is worked out in a way that should make the proof repairable.\n\nThe soft spots, in proportion. The sign error is the main one. From (20), strict concavity of Ψ gives Ψ(η^{k+1}) < Ψ(η^k) + ∇Ψ(η^k)·(η^{k+1}−η^k), so the weighted sum is Σ w_i κ_i(η^{k+1}) < Σ w_i κ_i(η^k), the reverse of (21). The next displayed inequality and the AM-GM step use this reversed direction, so the intended proof does go through if (21) is flipped. But as printed, the argument asserts A < B and then derives consequences of B < A; that is an internal inconsistency in the central theorem. It's a one-line fix, but it is a load-bearing line.\n\nSecond, Assumption 1 says Ξ contains the common support S. That assumption is not verified for the examples, and in the q-Gaussian case it looks ill-typed: S is a subset of the sample space X=R, while Ξ is a subset of the dual parameter space. The authors also seem to get the sign of Ξ wrong in Example 1—the computation I did gives η > 0 for their parametrization, not η < 0, which matters because the data statistics are nonnegative. The assumption presumably should involve F(S) ⊂ Ξ (or its convex hull), and it needs to be checked.\n\nThird, Proposition 2 is explicitly left without proof; minor, but it shows the paper is not fully self-contained.\n\nBottom line: this is a paper for the information-geometry / deformed exponential family community. The central idea is solid and likely correct, but the printed proof fails at one point. A serious referee should engage. I'd accept it conditionally and ask for the fix.","headline":"A genuinely new fixed-point MLE method, but the central monotonicity proof has a sign error and the examples don't verify the standing assumption—both fixable, but not as printed.","tokens_in":7725,"tokens_out":6097,"would_cite":false,"duration_ms":54836,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F10","62B10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves a fixed-point iteration for maximum likelihood in λ-exponential families that strictly increases the log-likelihood at every non-fixed step when λ<0, yielding MLEs through escort-expectation weighted averages of the…","keywords":["λ-exponential family","maximum likelihood estimation","fixed-point iteration","λ-duality","escort expectation","q-Gaussian distribution","Dirichlet perturbation"],"falsifier":"With a simulated $q$-Gaussian sample of size $n=500$ and $\\lambda=-1.2$, run Algorithm 1 from a dense grid of initial $\\theta^{(0)}$ in $\\Theta=(-\\infty,0)$; if any run has $\\ell(\\theta^{(k+1)})\\leq\\ell(\\theta^{(k)})$ while $\\theta^{(k)}$ is not a fixed point of (9), Theorem 1 fails. A second check is to compute $F(S)$ and $\\Xi$ explicitly for the $q$-Gaussian example under the paper's sign conventions and test the containment $S\\subset\\Xi$ required by Assumption 1.","tokens_in":6608,"feed_emoji":"📈","tokens_out":21300,"duration_ms":163066,"temperature":0.7,"pith_summary":"The paper proposes and analyzes a fixed-point algorithm for maximum likelihood estimation in the $\\lambda$-exponential family, a constant-curvature generalization of the exponential family built on a deformed duality from optimal transport. The core claim is that, for $\\lambda<0$ and under a set of regularity conditions, the log-likelihood strictly increases at every non-fixed iteration of the algorithm, so the iterates can be driven to the first-order conditions of the MLE. The scheme replaces the classical condition $\\nabla\\phi(\\hat{\\theta})=\\frac{1}{n}\\sum_i F(x_i)$ with an escort-expectation fixed point $\\hat{\\eta}=\\sum_i w_i(\\hat{\\theta})F(x_i)$, and the update is $\\eta^{(k+1)}=\\sum_i w_i(\\theta^{(k)})F(x_i)$. The paper illustrates the algorithm on a simulated $q$-Gaussian distribution and on the Dirichlet perturbation model, where the update becomes a simplex operation that does not depend on the noise parameter $\\lambda$.","feed_headline":"A fixed-point iteration provably raises likelihood at every step","feed_subtitle":"Maximum likelihood becomes a convex-combination update of escort expectations; q-Gaussian and Dirichlet examples converge.","key_machinery":"The load-bearing object is the $\\lambda$-duality between the primal potential $\\varphi$ and the dual potential $\\psi=\\varphi^{(\\lambda)}$, defined by $\\psi(\\eta)=\\sup_{\\theta\\in\\Theta}\\left[\\frac{1}{\\lambda}\\log(1+\\lambda\\theta\\cdot\\eta)-\\varphi(\\theta)\\right]$. The $\\lambda$-gradient $\\nabla^{(\\lambda)}\\varphi(\\theta)=\\nabla\\varphi(\\theta)/(1-\\lambda\\nabla\\varphi(\\theta)\\cdot\\theta)$ gives the escort expectation parameter $\\eta$, and its inverse is $\\nabla^{(\\lambda)}\\psi$. The algorithm is the fixed-point map $\\eta\\mapsto\\sum_i w_i(\\eta)F(x_i)$ with weights $w_i\\propto 1/(1+\\lambda\\theta\\cdot y_i)$. The monotonicity proof uses $\\Psi(\\eta)=e^{\\lambda\\psi(\\eta)}$, which is strictly concave when $\\lambda<0$; strict concavity makes the weighted average of the tangent terms $\\kappa_i(\\eta)=\\Psi(\\eta)+\\nabla\\Psi(\\eta)\\cdot(y_i-\\eta)$ move in the right direction between iterates, and the AM-GM inequality turns that directional movement into a log-likelihood increase.","core_discovery":"For a regular exponential family, the MLE satisfies $\\nabla\\phi(\\hat{\\theta})=\\frac{1}{n}\\sum_i F(x_i)$. The paper's first-order condition (9) replaces this with $\\hat{\\eta}=\\sum_i w_i(\\hat{\\theta})y_i$, where $w_i(\\theta)$ are normalized weights proportional to $1/(1+\\lambda\\theta\\cdot y_i)$ and $\\eta$ is the $\\lambda$-gradient escort parameter. Algorithm 1 iterates this condition: $\\eta^{(k+1)}=\\sum_i w_i(\\theta^{(k)})y_i$, $\\theta^{(k+1)}=\\nabla^{(\\lambda)}\\psi(\\eta^{(k+1)})$. The central discovery is Theorem 1: under Assumption 1, unless $\\theta^{(k)}$ already solves (9), the next iterate has strictly larger log-likelihood. The proof writes the likelihood as a product of tangent terms $\\kappa_i(\\eta)=\\Psi(\\eta)+\\nabla\\Psi(\\eta)\\cdot(y_i-\\eta)$ with $\\Psi=e^{\\lambda\\psi}$; the strict concavity of $\\Psi$ for $\\lambda<0$, combined with an AM-GM argument, converts a one-step inequality into monotone likelihood improvement. The paper also reports that the Dirichlet perturbation update is $p^{(k+1)}=p^{(k)}\\oplus\\frac{1}{n}\\sum_i(q_i\\ominus p^{(k)})$, independent of $\\lambda$, analogous to the sample mean in a normal location model.","pith_inferences":["Monotonicity alone does not establish a convergence rate; the paper also leaves uniqueness of the fixed point open, and Proposition 2's proof is said to be omitted for space, so the claimed $\\lambda$-independence of the Dirichlet update currently rests on a sketch. A natural next step is to analyze the Jacobian of the map $T$ at the fixed point and prove linear convergence with an explicit constan","Assumption 1 requires the observed statistics $F(x_i)$ to lie in the convex dual set $\\Xi$, but the paper only states that this can be verified rather than carrying out the verification; checking this containment should be a routine diagnostic before applying the algorithm to a new model.","Because the proof uses only the strict concavity of $e^{\\lambda\\psi}$, the same argument is likely to transfer to other $c$-dualities from optimal transport whose deformed potentials are strictly concave, giving a general template for monotone maximum-likelihood algorithms.","The Dirichlet update's independence from $\\lambda$ suggests a testable hypothesis: in models where $\\lambda$ is a nuisance scale parameter, the iteration may still estimate the structural parameter consistently under misspecification of $\\lambda$."],"forward_implications":["Any fixed point of Algorithm 1 satisfies the MLE first-order condition (9), so in the $\\lambda<0$ regime the maximum likelihood estimator must be a fixed point of the escort-expectation map.","The log-likelihood never decreases along an orbit of Algorithm 1 when Assumption 1 holds, so the iteration can be started from any valid parameter without risking a worse estimate.","In the $\\lambda\\to0$ limit the fixed-point condition reduces to the classical equation $\\nabla\\phi(\\hat{\\theta})=\\frac{1}{n}\\sum_i F(x_i)$, making the algorithm a continuous bridge between standard and deformed exponential-family estimation.","For the Dirichlet perturbation model, the update $p^{(k+1)}=p^{(k)}\\oplus\\frac{1}{n}\\sum_i(q_i\\ominus p^{(k)})$ does not involve $\\lambda$, so the composition parameter can be estimated without knowing the fixed noise size $\\sigma$.","The simulated $q$-Gaussian case with $\\lambda=-1.2$ and $n=500$ and the Dirichlet case with $d=2$ and $n=100$ converge quickly from different initial values, supporting the practical use of the algorithm."],"supporting_citations":[{"why":"Supplies the $\\lambda$-exponential family definition, the $\\lambda$-duality, Condition III.10, and the q-Gaussian and Dirichlet perturbation examples.","marker":"[17]"},{"why":"Gives the classical exponential-family MLE characterization $\\nabla\\phi(\\hat{\\theta})=\\frac{1}{n}\\sum_i F(x_i)$ that the fixed-point condition reduces to as $\\lambda\\to0$.","marker":"[2]"},{"why":"Proposes the alternative bound-based MLE algorithm for the $\\lambda$-exponential family; the new iteration is offered as a different route.","marker":"[9]"},{"why":"Defines the Aitchison perturbation and difference operations used to express the Dirichlet perturbation update.","marker":"[7]"},{"why":"Introduces the Dirichlet perturbation model that serves as the paper's second example.","marker":"[13]"},{"why":"Provides the c-duality in optimal transport of which the $\\lambda$-conjugate is a special case.","marker":"[15]"}],"fun_headline_variants":["Provably monotone likelihood for λ-exponential MLE","Fixed-point MLE: likelihood strictly rises each iteration","λ-exponential MLE via convex combination with proof","Constant-curvature MLE: iterate never lowers likelihood","New MLE for λ-family: monotone likelihood confirmed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is Assumption 1: for $\\lambda<0$, the dual parameter set $\\Xi=\\nabla^{(\\lambda)}\\varphi(\\Theta)$ is convex and contains every statistic $F(x_i)$ from the common support, so the next dual iterate stays in $\\Xi$ and the strict concavity of $e^{\\lambda\\psi}$ can be used.","fun_headline_variants_meta":{"raw":{"variants":["Provably monotone likelihood for λ-exponential MLE","Fixed-point MLE: likelihood strictly rises each iteration","λ-exponential MLE via convex combination with proof","Constant-curvature MLE: iterate never lowers likelihood","New MLE for λ-family: monotone likelihood confirmed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1585,"prompt_tokens":941,"completion_tokens":644,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":565}},"tokens_in":557,"tokens_out":644,"duration_ms":6339,"temperature":1.0,"reasoning_tokens":565,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:49:13.827526+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"With a simulated $q$-Gaussian sample of size $n=500$ and $\\lambda=-1.2$, run Algorithm 1 from a dense grid of initial $\\theta^{(0)}$ in $\\Theta=(-\\infty,0)$; if any run has $\\ell(\\theta^{(k+1)})\\leq\\ell(\\theta^{(k)})$ while $\\theta^{(k)}$ is not a fixed point of (9), Theorem 1 fails. A second check is to compute $F(S)$ and $\\Xi$ explicitly for the $q$-Gaussian example under the paper's sign conventions and test the containment $S\\subset\\Xi$ required by Assumption 1.","supporting_citations":[{"cited_title":"IEEE Transactions on Information Theory68(8), 5353–5373 (2022)","cited_arxiv_id":null,"evidence_quote":"Supplies the $\\lambda$-exponential family definition, the $\\lambda$-duality, Condition III.10, and the q-Gaussian and Dirichlet perturbation examples."},{"cited_title":"Springer (2016)","cited_arxiv_id":null,"evidence_quote":"Gives the classical exponential-family MLE characterization $\\nabla\\phi(\\hat{\\theta})=\\frac{1}{n}\\sum_i F(x_i)$ that the fixed-point condition reduces to as $\\lambda\\to0$."},{"cited_title":"Foundations of Data Science 6(1), 85–123 (2024)","cited_arxiv_id":null,"evidence_quote":"Proposes the alternative bound-based MLE algorithm for the $\\lambda$-exponential family; the new iteration is offered as a different route."},{"cited_title":"Mathematical Ge- ology 35(3), 279–300 (2003)","cited_arxiv_id":null,"evidence_quote":"Defines the Aitchison perturbation and difference operations used to express the Dirichlet perturbation update."},{"cited_title":"Probability Theory and Related Fields178(1), 613–654 (2020)","cited_arxiv_id":null,"evidence_quote":"Introduces the Dirichlet perturbation model that serves as the paper's second example."},{"cited_title":"American Mathematical Society (2003)","cited_arxiv_id":null,"evidence_quote":"Provides the c-duality in optimal transport of which the $\\lambda$-conjugate is a special case."}],"review_version":1}