{"id":"c90a4db0-9949-4a21-8654-2db05398a8a9","arxiv_id":"2506.10110","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The KL exponent of a squared-variable objective is deduced from the original: max{alpha,1/2} under strict complementarity and (1+beta)/2 with beta=1-gamma(1-alpha) under a convex error-bound condition.","lead":"This paper analyzes what happens to the Kurdyka-Lojasiewicz (KL) exponent, a quantity that controls how fast optimization algorithms converge, when an optimization problem is rewritten using squared variables. It proves formulas that transfer the KL exponent from the original problem to the squared problem, under strict complementarity or under a convex error-bound assumption.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 6.2's proof does not follow as printed: its error-bound set Sρ is defined on I, while Lemma 6.1, which the proof invokes unchanged, requires Sρ on I^c; the non-strict-complementarity exponent formula therefore lacks a valid hypothesis as stated.","rationale":"Good-faith reading: the paper's contribution is a calculus for KL exponents under square lifting. The strict-complementarity transfer (Theorem 5.12) is supported by a long, detailed chain involving Lemmas 5.8–5.10, and my reading found no gap there. The non-strict-complementarity result (Theorem 6.2) is the advertised extension, and it is exactly where the argument is least secure: it depends on Lemma 6.1, whose proof is omitted and whose Sρ set does not match the theorem's stated Sρ set. The reader's weakest_assumption identifies this mismatch, and I agree with that diagnosis. Because the flaw is localized and likely correctable rather than a demonstrated counterexample, the existing CONDITIONAL verdict remains appropriate instead of ACCEPT or REJECT. A written proof of Lemma 6.1 with a corrected Sρ definition would settle the issue.","tokens_in":24903,"tokens_out":14851,"duration_ms":178801,"concrete_test":"Write out the missing proof of Lemma 6.1 for nonsmooth (polyhedral) g, tracking the set Sρ through every inequality. If the admissible Sρ is {x : |xi − xbar_i| ≤ ρ_i ∀ i ∈ I^c}, then re-run Theorem 6.2's proof with the corrected set and verify the corrected hypothesis on a canonical example such as f(x)=||Ax−b||^2 with xbar a minimizer having zero coordinates; if the printed (6.2) with Sρ defined on I were used at ρ=0, that example would violate the resulting slice condition. If the lemma's proof instead goes through with Sρ defined on I, or if Theorem 6.2 can be proved without invoking Lemma 6.1, then the concern fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's advertised non-strict-complementarity transfer is Theorem 6.2, and its proof is a direct invocation of Lemma 6.1. The two statements do not line up. Lemma 6.1 fixes an index set I and assumes (6.1) with Sρ := {x : |xi − xbar_i| ≤ ρ_i for i ∈ I^c}; its conclusion is the bound sum_{i∈I}|v_i|^2 + sum_{i∈I^c}|x_i − xbar_i||v_i|^2 ≥ σ(φ(x)−φ(xbar))^{1+β}. Theorem 6.2 takes I := {i : xbar_i ≠ 0} but states its error-bound hypothesis (6.2) with Sρ := {x : |xi − xbar_i| ≤ ρ_i for i ∈ I}. Since I^c is the zero-coordinate block, the two definitions are not interchangeable. On a literal reading, (6.2) controls only the nonzero block I, whereas the needed estimate (6.3) has |x_i||v_i|^2 terms on I^c. The defect is not cosmetic: if the paper's convention R_+ includes 0, taking ρ = 0 in the printed (6.2) forces every point of U ∩ {x_I = xbar_I} to be a global minimizer, a condition that the intended convex-quadratic examples do not satisfy. Moreover, Lemma 6.1 itself is stated without proof; the text only says it follows by the same argument as [22, Lemma 3.10], but that argument concerns a different, separable smooth setting. The proof of Theorem 6.2 gives no independent derivation of (6.3), so the fractional exponent (1+β)/2 is not established as printed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the square (Hadamard) reparametrization Φ(y)=f(y^2)+g(y^2) of the composite objective ϕ(x)=f(x)+g(x), where f∈C^2 and g is a proper polyhedral function with dom g⊆R_+^n. The main contributions are: a formula for the distance from 0 to the subdifferential of Φ (Prop. 3.2); a computation of the second subderivative of g∘T on a linear subspace and a characterization of those stationary points of Φ that correspond to stationary points of ϕ (Prop. 4.2 and 4.3); and two KL-exponent transfer theorems. Theorem 5.12 states that under the strict complementarity condition 0∈ri(∂ϕ(\\bar{x})), if ϕ has KL exponent α at \\bar{x}, then Φ has KL exponent max{α,1/2} at \\bar{y} with \\bar{x}=\\bar{y}^2. Theorem 6.2 claims that under convexity and a Hölderian error-bound condition (6.2), the exponent becomes (1+β)/2 with β=1−γ(1−α), without strict complementarity. The paper is written as a sequence of theorems with detailed proofs for the strict-complementarity part, but the non-strict part depends on an unproved lemma and a hypothesis that is stated on the wrong index block.","tokens_in":25282,"tokens_out":13052,"duration_ms":139622,"significance":"If the main claims are correct, the KL-exponent transfer is a valuable parameter-free result: it gives a principled way to compute the KL exponent of a frequently used reparametrization without fitting constants, and it extends the Hadamard-parametrization analysis of [22] to the nonsmooth case. The first-order distance formula and the face-based second-subderivative computation are substantive original tools, and the paper is honest in stating its assumptions. However, the advertised non-strict-complementarity transfer is not established as printed: the error-bound hypothesis in Theorem 6.2 is defined on the wrong coordinate block relative to Lemma 6.1, and Lemma 6.1 itself is asserted by reference rather than proved. The strict-complementarity theorem (Section 5) appears internally consistent and is developed in detail, which supports a major-revision verdict rather than rejection, but the paper's full advertised scope requires repair of Section 6.","major_comments":[{"comment":"Theorem 6.2's error-bound hypothesis does not match the lemma it invokes. Lemma 6.1 assumes (6.1) with S_ρ = {x : |x_i − \\bar{x}_i| ≤ ρ_i for i ∈ I^c}, while Theorem 6.2 states (6.2) with S_ρ = {x : |x_i − \\bar{x}_i| ≤ ρ_i for i ∈ I}. For I = {i : \\bar{x}_i ≠ 0}, these are different sets and neither contains the other. The proof of Theorem 6.2 then says “By Lemma 6.1” to obtain (6.3); on a literal reading, (6.2) does not imply the hypothesis of Lemma 6.1. Since (6.3) is exactly the inequality that produces the claimed exponent (1+β)/2, the non-strict-complementarity theorem lacks a valid hypothesis as printed. The fix is not a one-character typo: the geometry of Lemma 6.1 is tailored to the complementary (zero) block, so either (6.2) must be restated with I^c and the proof re-checked, or a genuinely different argument must be supplied.","section":"6 (Eqs. (6.1)-(6.3))"},{"comment":"Lemma 6.1 is load-bearing but is not proved in the manuscript. The text says it follows by essentially the same argument as [22, Lemma 3.10], with two listed differences: g is not assumed differentiable, and the separable structure encoded by J_1 in [22] is absent. These are substantive differences, not cosmetic ones: the proof in [22] relies on separable smooth calculus, while the present setting is nonsmooth and nonseparable. The manuscript does not reproduce the modified argument, state which steps change, or indicate how the constants are obtained. Because Theorem 6.2's only route to (6.3) is Lemma 6.1, the advertised non-strict exponent transfer is not verifiable from the manuscript as written. A full proof of Lemma 6.1, or a precise step-by-step reduction to [22, Lemma 3.10] with all adaptations, is required.","section":"6 (Lemma 6.1)"}],"minor_comments":[{"comment":"In the proof of Proposition 3.2, the text invokes “Proposition 3.2” to obtain w ∈ ∂g(y^2) from v ∈ ∂(g∘T)(y); this should refer to Proposition 3.1(ii), since Proposition 3.2 is the statement being proved.","section":"3 (Prop. 3.2)"},{"comment":"In the proof of Proposition 4.2, several occurrences of ∂f(\\bar{x}) in Equations (4.10)(d), (4.11), and the surrounding display should be ∂g(\\bar{x}); f is not polyhedral in the statement and the result concerns g.","section":"4 (Prop. 4.2)"},{"comment":"In the proof of Proposition 4.3, the phrase “d²Φ(y)(w) ≥ 0 for all w ∈ S_{I^c}” should read w ∈ S_I, to match hypothesis (i) and the subsequent display that quantifies over w ∈ S_I.","section":"4 (Prop. 4.3)"},{"comment":"Definition 2.2(b) is circular as printed: it defines v ∈ ∂f(\\bar{x}) in terms of sequences v^k ∈ ∂f(x^k). The definition should use regular subgradients, v^k ∈ \\hat∂ f(x^k), as in [26, Definition 8.3].","section":"2 (Def. 2.2)"},{"comment":"In Definition 2.3, the phrase “the φ(s) in (2.5)” should refer to Equation (2.1), which is the sharpened KL inequality; there is no Equation (2.5) in the paper.","section":"2 (Def. 2.3)"},{"comment":"In the proof of Lemma 5.8, the sentence “Let V be a neighborhood of s such that (5.4) holds by using Lemma 5.7” appears to refer to Equation (5.3), not to the inequality (5.4) that the lemma is proving.","section":"5 (Lemma 5.8)"}],"recommendation":"major_revision","confidential_remarks":"The index-block mismatch in Theorem 6.2 looks like a repairable typo: if (6.2) is restated with I^c, then Lemma 6.1 would apply as intended. The larger obstruction is that Lemma 6.1 is not proved despite being adapted to a nonsmooth, nonseparable setting, and the non-strict transfer is one of the paper's advertised contributions. Section 5 is detailed and appears internally consistent, and the first-order and second-order tools have independent value. I therefore recommend major revision rather than rejection. The editor may wish to verify that the author supplies a full proof of Lemma 6.1 and a corrected Theorem 6.2 in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my take. The paper's real contribution is a transfer principle for KL exponents under the square reparameterization, and the strict-complementarity half is solid. I would not desk-reject it.\n\nWhat is new: Proposition 3.2 (distance to the subdifferential of the squared map), Proposition 4.2 (second subderivative on a subspace), Proposition 4.3 (stationary-point characterization via second subderivatives), and Theorem 5.12 (KL exponent max{alpha, 1/2} when 0 is in the relative interior of the subdifferential). The proof of Theorem 5.12 is detailed and internally consistent; I found no hidden circularity or fitted constants. The use of [22] for Lemma 6.1 is honest, and the parameter-free derivation is a plus.\n\nThe soft spot is Section 6. Lemma 6.1 is stated without proof, with only a pointer to [22, Lemma 3.10], despite the setting here being genuinely different: g is not smooth and there is no explicit separable structure. More importantly, Theorem 6.2's hypothesis (6.2) defines S_rho by constraining the nonzero block I, while Lemma 6.1 requires S_rho on I^c. The proof invokes Lemma 6.1 unchanged, so the advertised fractional exponent (1+beta)/2 is not established as printed. The sharper version of this objection is worth stating: with the printed S_rho, taking rho = 0 forces every point of U cap {x_I = xbar_I} to be a global minimizer, which the intended convex-quadratic examples do not satisfy.\n\nThis is a load-bearing flaw in Section 6, but it looks correctable: likely S_rho in (6.2) should be on I^c, and Lemma 6.1 needs a real proof. The flaw does not taint Sections 3-5. The first-order distance formula, the second-subderivative calculus, and the strict-complementarity transfer stand on their own.\n\nWho this is for: researchers working on squared-variable/Hadamard reformulations, KL exponents of nonsmooth lifted problems, basis pursuit, and simplex-constrained optimization. They should read Sections 3-5 now and treat Section 6 with caution until corrected. My recommendation: send to a serious referee. The referee should require a corrected Section 6 with Lemma 6.1 proved; after that, this is a useful contribution worth citing.","headline":"A genuinely strong strict-complementarity KL transfer with a broken non-strict-complementarity section that needs a rewrite before this is citable in full.","tokens_in":25805,"tokens_out":2548,"would_cite":true,"duration_ms":29181,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C25","90C26","68Q25"],"pacs":[],"model":"deepseek-v4-flash","headline":"Squaring a constrained composite objective yields an explicit Kurdyka–Łojasiewicz exponent for the unconstrained lift, deduced from the original problem's data.","keywords":["Kurdyka–Łojasiewicz exponent","square transformation","Hadamard parameterization","polyhedral function","second subderivative","subdifferential distance","error bound","strict complementarity"],"falsifier":"Take a small convex instance with polyhedral $g$ and a degenerate stationary point where strict complementarity fails, and check whether $\\mathrm{dist}^2(0, \\partial\\Phi(y)) \\ge c(\\Phi(y) - \\Phi(\\bar{y}))^{1+\\beta}$ holds with the stated $\\beta$ for all nearby $y$; any failing sequence with no positive $c$ refutes the transfer. Separately, inspect the step of Theorem 6.2 that invokes Lemma 6.1: replacing the printed condition (6.2) by the lemma's $I^c$ version and re-running the proof would show whether the index-set mismatch is a typo or a substantive gap.","tokens_in":24680,"feed_emoji":"📐","tokens_out":9822,"duration_ms":89061,"temperature":0.7,"pith_summary":"This paper studies the oldest trick in constrained optimization: replace $x \\ge 0$ by $x = y^2$, turning a constrained composite problem into an unconstrained one. The central claim is that the Kurdyka--\\L ojasiewicz (KL) exponent of the squared problem---the number that controls how quickly first-order methods approach a stationary point---can be deduced from the original problem instead of being re-established from scratch. Under a strict-complementarity condition the transfer is exact: the new exponent is $\\max\\{\\alpha, 1/2\\}$ when the old one is $\\alpha$. In a convex degenerate case satisfying an error-bound condition, the paper derives the exponent $(1+\\beta)/2$ with $\\beta = 1 - \\gamma(1-\\alpha)$, where $\\gamma$ is the error-bound exponent. The proof matters because the square map has zero or low-rank Jacobian at boundary points, so the usual chain rules for subdifferentials and second subderivatives are not available.","feed_headline":"Squaring variables: KL exponent is max(α, 1/2)","feed_subtitle":"The exponent controlling first-order convergence is deduced, not fitted, for polyhedral composite problems.","key_machinery":"The machine that carries the argument is the exact distance formula of Proposition 3.2, $\\mathrm{dist}(0, \\partial\\Phi(y)) = 2\\,\\mathrm{dist}(-y \\circ \\nabla f(y^2),\\, y \\circ \\partial g(y^2))$, derived by exploiting the evenness of $H(y) = g(y^2)$ to avoid the failed chain rule for $T: y \\mapsto y^2$. Alongside it sits the second-subderivative formula of Proposition 4.2: on the subspace $S_I = \\{w : w_I = 0\\}$, $d^2(g \\circ T)(\\bar{y} \\mid 2\\bar{y} \\circ v)(w)$ equals $2 \\sup_{p \\in S(I,v) \\cap \\partial g(\\bar{x})} \\langle w_{I^c}^2,\\, p_{I^c} \\rangle$, where $S(I,v)$ is the set of vectors agreeing with $v$ on $I$. This formula is what allows the paper to single out the stationary points of $\\Phi$ that correspond to genuine stationary points of $\\phi$. The KL transfer in Section 5 then combines these formulas with finite face decompositions of polyhedral subdifferentials, minimal exposed faces, and projection estimates on relative interiors.","core_discovery":"On the paper's own terms, the discovery is two precise transfer theorems. Theorem 5.12: if $\\phi = f + g$ with $f \\in C^2$ and $g$ proper polyhedral, $0$ lies in the relative interior of $\\partial\\phi(\\bar{x})$, and $\\phi$ has the KL property at $\\bar{x}$ with exponent $\\alpha$, then $\\Phi(y) = f(y^2) + g(y^2)$ has the KL property at $\\bar{y}$ (with $\\bar{x} = \\bar{y}^2$) with exponent $\\max\\{\\alpha, 1/2\\}$. Theorem 6.2: if $\\phi$ is convex, has KL exponent $\\alpha$ at a stationary point $\\bar{x}$, and satisfies the H\\\"older error-bound condition (6.2) with exponent $\\gamma$, then $\\Phi$ has KL exponent $(1+\\beta)/2$, where $\\beta = 1 - \\gamma(1-\\alpha)$. Both theorems rest on the identity $\\mathrm{dist}(0, \\partial\\Phi(y)) = 2\\,\\mathrm{dist}(-y \\circ \\nabla f(y^2),\\, y \\circ \\partial g(y^2))$, which makes the subdifferential distance in the lifted problem a weighted subdifferential distance in the original variables, and on a second-subderivative formula for $g \\circ T$ along the subspace where the nonzero coordinates are frozen. The paper's contribution is that the exponent is deduced rather than fitted: no new KL analysis of the lifted problem is needed once the original data are known.","pith_inferences":["If Theorem 5.12 holds, local rate analyses for algorithms on the squared variables can be imported from the original problem, which would save re-deriving KL exponents for each reparameterized model.","The face-based proof suggests a testable generalization: replace polyhedral $g$ by a definable (e.g., semi-algebraic) function with known KL exponent; the squared lift may still admit an explicit exponent, though the face machinery would need a tame-geometry substitute.","A literal reading of Theorem 6.2 defines $S_\\rho = \\{x : |x_i - \\bar{x}_i| \\le \\rho_i \\text{ for } i \\in I\\}$ while Lemma 6.1, which the proof invokes unchanged, requires the complementary index set $I^c$; if that is a typo, the theorem should be read with $I^c$, and the numerical predictions of Remark 6.4 still stand."],"forward_implications":["If $\\phi$ has KL exponent $\\alpha \\le 1/2$ at a strict-complementarity stationary point, the squared problem inherits exactly $\\alpha$; if $\\alpha > 1/2$, the squared problem improves to $1/2$.","In the convex degenerate case, the formula $(1+\\beta)/2$ converts known original-problem data (KL exponent $\\alpha$ and error-bound exponent $\\gamma$) into a specific convergence exponent for the lifted problem, so no additional KL computation is required.","The identity for $\\mathrm{dist}(0, \\partial\\Phi(y))$ gives a computable stationarity measure for the lifted problem, a practical stopping or certificate quantity for algorithms run directly on $y$.","The second-order characterization recovers, as special cases, earlier results for simplex-to-sphere Hadamard parameterizations, basis-pursuit reformulations, and polyhedral Hadamard parameterizations."],"supporting_citations":[{"why":"supplies the definitions and calculus of subderivatives and subdifferentials whose failure under the square map motivates the direct arguments","marker":"[26]"},{"why":"predecessor for the Hadamard-parametrization KL exponent transfer and the source of the error-bound lemma used in the degenerate case","marker":"[22]"},{"why":"provides the H\\\"older error-bound inequality and polyhedral geometry facts used in Lemmas 5.6 and 6.1","marker":"[11]"},{"why":"gives the polyhedral convex analysis facts: finite faces, projections, and support functions used throughout Sections 4 and 5","marker":"[25]"},{"why":"supplies the theory of minimal exposed faces and finite face decompositions used in Lemmas 5.8 through 5.10","marker":"[28]"},{"why":"provides the polyhedral subdifferential outer-semicontinuity result used in Proposition 3.1","marker":"[21]"},{"why":"motivating application and special case: simplex-to-sphere Hadamard parametrization recovered by the second-order characterization","marker":"[17]"},{"why":"motivating application and special case: polyhedral Hadamard parametrizations whose stationary-point characterization is recovered","marker":"[29]"},{"why":"supplies the known KL exponent $1/2$ for convex piecewise linear-quadratic objective forms used in Remark 6.4","marker":"[15]"}],"fun_headline_variants":["KL exponent after squaring: deduced, not fitted","Square transform: KL exponent transfers exactly","Squaring variables: KL exponent max(α, 1/2)","Exact KL exponent for reparameterized objectives","Squaring maps KL exponent to max(α, 1/2)"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the H\\\"older error-bound condition (6.2) in Theorem 6.2: it is assumed for every $\\rho$ rather than proved for the general polyhedral problem, and as printed it constrains the nonzero-coordinate set $I$ while the lemma cited in the proof requires the complementary set $I^c$.","fun_headline_variants_meta":{"raw":{"variants":["KL exponent after squaring: deduced, not fitted","Square transform: KL exponent transfers exactly","Squaring variables: KL exponent max(α, 1/2)","Exact KL exponent for reparameterized objectives","Squaring maps KL exponent to max(α, 1/2)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1578,"prompt_tokens":1012,"completion_tokens":566,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":484}},"tokens_in":628,"tokens_out":566,"duration_ms":6015,"temperature":1.0,"reasoning_tokens":484,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:35:11.574171+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a small convex instance with polyhedral $g$ and a degenerate stationary point where strict complementarity fails, and check whether $\\mathrm{dist}^2(0, \\partial\\Phi(y)) \\ge c(\\Phi(y) - \\Phi(\\bar{y}))^{1+\\beta}$ holds with the stated $\\beta$ for all nearby $y$; any failing sequence with no positive $c$ refutes the transfer. Separately, inspect the step of Theorem 6.2 that invokes Lemma 6.1: replacing the printed condition (6.2) by the lemma's $I^c$ version and re-running the proof would show whether the index-set mismatch is a typo or a substantive gap.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the definitions and calculus of subderivatives and subdifferentials whose failure under the square map motivates the direct arguments"},{"cited_title":"F acchinei and J.-S","cited_arxiv_id":null,"evidence_quote":"provides the H\\\"older error-bound inequality and polyhedral geometry facts used in Lemmas 5.6 and 6.1"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"gives the polyhedral convex analysis facts: finite faces, projections, and support functions used throughout Sections 4 and 5"},{"cited_title":"Soltan, Lectures on convex sets , World Scientific, 2019","cited_arxiv_id":null,"evidence_quote":"supplies the theory of minimal exposed faces and finite face decompositions used in Lemmas 5.8 through 5.10"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the polyhedral subdifferential outer-semicontinuity result used in Proposition 3.1"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"motivating application and special case: simplex-to-sphere Hadamard parametrization recovered by the second-order characterization"},{"cited_title":"Tang and K.-C","cited_arxiv_id":null,"evidence_quote":"motivating application and special case: polyhedral Hadamard parametrizations whose stationary-point characterization is recovered"},{"cited_title":"Li and T","cited_arxiv_id":null,"evidence_quote":"supplies the known KL exponent $1/2$ for convex piecewise linear-quadratic objective forms used in Remark 6.4"}],"review_version":1}