{"id":"d3a47df7-5f0d-4558-adb6-55b404de81e0","arxiv_id":"2504.17193","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"If a continuous map between a Hilbert and a reflexive Banach space satisfies the Hessian-free Jensen inequality, then it is Frechet differentiable and its derivative is Lipschitz continuous.","lead":"A short proof shows that a Hessian-free Jensen-type inequality, used in modern first-order optimization, is equivalent to Lipschitz continuity of the derivative for continuous functions from a Hilbert space to a reflexive Banach space. This closes the converse of a known result and generalizes the equivalence beyond scalar-valued functions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified.","rationale":"I read the full proof carefully. The main risk areas were (a) the application of Baillon–Haddad in Lemma 2, (b) the reflexivity-based lifting in Lemma 3, and (c) the imported forward direction. I checked (a) by recomputing the cocoercivity expansion: from ⟨Ld+A,d⟩ ≥ (1/2L)||Ld+A||² one obtains ||A||² ≤ L²||d||², so the constant L is correct. I checked (b): part (1) uses continuity of F to bound y* ↦ φ'_{y*}(x); part (2) uses reflexivity exactly to place fx(h) in Y; part (3) uses Hahn–Banach and the uniform Lipschitz condition; all steps are valid. I checked (c) by an independent derivation via the descent lemma and the variance identity, which reproduces (1.2) with the same constant L/2. The reader's weakest assumption, reflexivity of Y, is indeed the least secure input in the sense that the proof technique would break without it, but it is a stated hypothesis and does not undermine Theorem 1 as it stands. Therefore no load-bearing concern changes the verdict.","tokens_in":6191,"tokens_out":45488,"duration_ms":407987,"concrete_test":"Re-derive the forward implication (i)=>(ii) for the Banach-valued case directly from the descent lemma and the variance identity; if this does not reproduce the constant L/2 in (1.2), then the equivalence would fail at the level of the constant. This check is independent of the cited Lemma 1 and would also confirm that the generalization from Rd to Hilbert/reflexive-Banach spaces preserves the statement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is sound as stated. The converse proof is internally consistent: Lemma 2 derives convexity of the two quadratic perturbations of each slice from the n=2 case of (1.2), applies the Baillon–Haddad theorem to obtain L-Lipschitz gradients of slices, and Lemma 3 assembles these slices into a Frechet derivative using reflexivity of Y to identify the Y**-valued map with an element of Y. The only step that could be seen as fragile is this lifting step: if Y were not reflexive, y* ↦ φ'_{y*}(x)h would be a genuine element of Y** and the construction would not deliver a derivative in Y. However, reflexivity is an explicit hypothesis in Theorem 1, not a hidden assumption, and no part of the proof requires more. The forward direction is not re-derived but is a standard consequence of the descent lemma ||F(y)-F(x)-F'(x)(y-x)|| ≤ L/2||y-x||² together with the variance identity Σλ_i||x_i-Σλ_jx_j||² = Σ_{i<j}λ_iλ_j||x_i-x_j||², so the citation is reliable. No internal inconsistency or missing proof that affects the theorem was found.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper establishes Theorem 1, an equivalence between a Jensen-type inequality (1.2) and Lipschitz continuity of the Fréchet derivative for continuous maps from a Hilbert space into a reflexive Banach space. The forward direction is cited to Marumo–Takeda [7]; the paper's contribution is the converse. The converse proof has two parts: Lemma 2 uses the n=2 case of (1.2) and the Baillon–Haddad theorem to show that every unit-ball slice y*∘F has an L-Lipschitz continuous gradient; Lemma 3 shows, via reflexivity of the codomain, that these slices assemble into a Fréchet derivative of F itself with the same Lipschitz constant. Corollary 1 yields the converse of Lemma 1: a differentiable scalar function whose gradient satisfies (1.1) is twice differentiable with L-Lipschitz Hessian.","tokens_in":6455,"tokens_out":15147,"duration_ms":130209,"significance":"The result is clean and exact: the constant L is preserved and the characterization is genuinely Hessian-free. The proof is transparent and carefully handles the infinite-dimensional codomain, with reflexivity explicitly used to lift the derivative from Y** to Y. The paper thus clarifies why reflexivity is the natural hypothesis and strengthens the known forward direction into a full equivalence. The use of the Baillon–Haddad theorem is appropriate, the derivation is verifiable, and I found no circularity. This is a worthwhile contribution to the optimization literature on Hessian-free methods.","major_comments":[],"minor_comments":[{"comment":"The forward implication (i)⇒(ii) is not proved in the manuscript; the text after Theorem 1 only states that the proof is essentially the same as [7, Lemma 3.1]. Since (i)⇒(ii) is part of the stated equivalence, please add a short proof—for example, apply the standard descent inequality |φ(y)−φ(x)−φ'(x)(y−x)| ≤ L/2 ||y−x||² to each slice φ_y* = y*∘F and use the identity Σ_i λ_i ||x_i − Σ_j λ_j x_j||² = Σ_{i<j} λ_i λ_j ||x_i − x_j||²—so that the paper is self-contained.","section":"Theorem 1 (Section 1)"},{"comment":"The sentence beginning \"In particular, we have L||·||² − (L/2||·||² + φ_y*) = L/2||·||² − φ_y* convex\" is potentially misleading, since a difference of two convex functions is not convex in general. The intended statement is that L/2||·||² − φ_y*, which was already proved convex, is the function to which Theorem 2 applies with β = 2L; please rephrase to avoid the appearance of an invalid inference.","section":"Lemma 2"},{"comment":"In the displayed chain for ||fx − fy||, the equality between sup_{||h||≤1, y*∈Y*_1} y*(fx(h) − fy(h)) and sup_{y*∈Y*_1} ||φ'_y*(x) − φ'_y*(y)|| uses the fact that sup_{||h||≤1} y*((Tx − Ty)h) = ||y*∘(Tx − Ty)|| and that the supremum over the unit ball of Y* of these norms is the operator norm; please state this briefly to help the reader.","section":"Lemma 3, part (3)"},{"comment":"Please fix typographical errors: \"continuos\" in the abstract, \"Lipschi tz\" in the title/first line, and the garbled accents in \"Fr´echet\" and \"H¨ older\" in the references.","section":"General"}],"recommendation":"minor_revision","confidential_remarks":"The paper is well within the journal's scope and the proof is sound. The only substantive request is a self-contained proof of the forward direction, which is a local addition; otherwise the issues are presentation-level. I support publication after minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper answers a real open question and does it honestly. The Hessian-free Jensen inequality (1.2) really does characterize Lipschitz differentiability for maps from Hilbert to reflexive Banach spaces. The converse is new; the forward direction was already in Marumo-Takeda. The result is meaningful for optimization theory, not earth-shaking, but solid.\n\nWhat is new is the converse proof. The approach is to slice F with functionals y*, turn (1.2) into a two-sided convexity statement for the quadratic perturbations of each slice, apply a piece of Baillon-Haddad to get Lipschitz gradients on the slices, then assemble a Frechet derivative via reflexivity of Y to lift the Y**-valued object back to Y. That assembly step is the heart of the paper and it is well executed. The Hahn-Banach norm identity is used cleanly. The proof is reproducible from the text.\n\nSoft spots, in proportion: the reflexivity assumption is essential, not hidden; the authors state it explicitly. But it means the vector-valued generalization stops exactly where the lifting argument needs it to stop. If Y is not reflexive, the derivative candidate lives in Y** and the argument does not deliver Frechet differentiability. That is a boundary, not a flaw. The forward direction is cited rather than re-derived; since it follows from the standard descent lemma plus a variance identity, the citation is reliable.\n\nOne minor quibble: the title and abstract talk about Lipschitz continuous Hessian, but the theorem concerns L-Lipschitz continuity of the derivative F'. For a vector-valued map between Banach spaces the Hessian is not defined in the usual sense. The scalar Corollary does give twice differentiability with Lipschitz Hessian, so the title is defensible, but the emphasis is slightly off.\n\nThe paper is for people working on smoothness structure in optimization, especially the line of work from Carmon et al. that uses Hessian smoothness through gradient inequalities. A reader who wants a precise characterization rather than folklore will get value.\n\nSerious referee? Yes. The central argument holds up, the citations check out, and the result deserves to become a standard lemma. It needs no heavy revision.","headline":"Solid, honest resolution of the converse of the Hessian-free Jensen inequality; the proof via slicing and Baillon-Haddad is clean and the reflexivity assumption is exactly where it should be.","tokens_in":6937,"tokens_out":1405,"would_cite":true,"duration_ms":13188,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49J50","47H05","90C25","46B10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The Hessian-free Jensen inequality exactly characterizes Lipschitz-continuous derivatives.","keywords":["Hessian-free inequality","Lipschitz continuous Hessian","Jensen-type inequality","Baillon-Haddad theorem","Fréchet differentiability","reflexive Banach space","cocoercivity","first-order optimization"],"falsifier":"The decisive test is the two-point case with weights $1/2,1/2$. For $F(x)=|x|^\\alpha$ on $\\mathbb{R}$ with $0<\\alpha<2$, taking $y=0$ gives $|F(x/2)-F(x)/2|=(2^{-\\alpha}-1/2)|x|^\\alpha$, which must be no larger than $\\frac{L}{8}|x|^2$; letting $x\\to 0$ shows no finite $L$ works, so non-differentiable Hölder maps are excluded. A continuous non-Fréchet-differentiable map from a Hilbert space to a reflexive Banach space that satisfies (1.2) for some $L$ would refute Theorem 1, and the scaling calculation shows why the inequality is strong enough to rule out the simplest nonsmooth candidates.","tokens_in":6061,"feed_emoji":"📐","tokens_out":12443,"duration_ms":111205,"temperature":0.7,"pith_summary":"The paper establishes a converse: a continuous map $F$ from a real Hilbert space $X$ into a real reflexive Banach space $Y$ has an $L$-Lipschitz Fréchet derivative if and only if it satisfies the Hessian-free Jensen inequality (1.2) with the same constant $L$, for every finite convex combination. The inequality involves only values of $F$ at finitely many points, never the Hessian or any second derivative. The proof works by slicing $F$ with dual functionals, applying convexity and the Baillon–Haddad theorem to each slice, and then gluing the slice derivatives back together using reflexivity. If correct, this turns a second-order regularity condition into a first-order, checkable condition and answers the natural question left open by the known forward implication from nonconvex optimization.","feed_headline":"Hessian-free inequality forces Lipschitz-continuous derivative","feed_subtitle":"A Jensen-type bound on function values is enough to certify differentiability and pin down the Lipschitz constant.","key_machinery":"The load-bearing object is inequality (1.2) in its two-point form, combined with the identity $t(1-t)\\|x-y\\|^2=(1-t)\\|x\\|^2+t\\|y\\|^2-\\|x+t(y-x)\\|^2$. This identity converts the two-sided bound on $\\varphi_{y^*}$ into convexity of both quadratic perturbations $\\frac{L}{2}\\|\\cdot\\|^2-\\varphi_{y^*}$ and $\\frac{L}{2}\\|\\cdot\\|^2+\\varphi_{y^*}$. The Baillon–Haddad theorem — which identifies cocoercivity of a gradient with convexity of a quadratic perturbation — then yields Fréchet differentiability and cocoercivity of the perturbed gradients, giving the $L$-Lipschitz property for each slice; reflexivity of $Y$ is the bridge that assembles these slice derivatives into the derivative of $F$.","core_discovery":"The central claim is Theorem 1: for continuous $F$ and $L>0$, condition (i) — $F$ is Fréchet differentiable and $F'$ is $L$-Lipschitz — is equivalent to condition (ii), namely $\\|F(\\sum_i\\lambda_i x_i)-\\sum_i\\lambda_i F(x_i)\\|\\le \\frac{L}{2}\\sum_{i<j}\\lambda_i\\lambda_j\\|x_i-x_j\\|^2$ holding for all convex weights. The proof of (ii)$\\Rightarrow$(i) is the paper's contribution; (i)$\\Rightarrow$(ii) is quoted from the known lemma. The proof shows each scalar slice $\\varphi_{y^*}=y^*\\circ F$ has an $L$-Lipschitz gradient, then uses reflexivity of $Y$ to define a candidate derivative $f_x(h)\\in Y$ by $y^*(f_x(h))=\\varphi'_{y^*}(x)h$, and identifies it as the Fréchet derivative of $F$ with the promised Lipschitz constant. In the scalar case this gives Corollary 1: a differentiable real-valued $f$ whose gradient satisfies (1.1) is twice differentiable with $L$-Lipschitz Hessian.","pith_inferences":["A natural test of sharpness is to drop reflexivity: if a continuous non-differentiable map from a Hilbert space into a non-reflexive space such as $c_0$ satisfied (1.2), the reflexivity assumption would be shown necessary; the paper's gluing step gives the precise place such an example would have to fail.","Because the proof checks only two-point convex combinations, the inequality could in principle be verified empirically on finitely many point pairs, yielding a computable lower bound on the Lipschitz constant of the derivative; algorithms might exploit this as a certificate, though the paper does not discuss this.","The same quadratic-perturbation technique suggests that analogues for higher-order smoothness would require a generalized Baillon–Haddad statement, which the paper's final remarks identify as a nontrivial obstacle."],"forward_implications":["Every continuous map satisfying (1.2) is automatically Fréchet differentiable; no separate regularity assumption is needed.","The constant $L$ in the inequality is exactly the Lipschitz constant of the derivative, so the Hessian-free condition supplies quantitative information, not just qualitative smoothness.","For real-valued functions on Hilbert space, the gradient inequality (1.1) forces twice differentiability with $L$-Lipschitz Hessian, recovering the converse of the lemma used in first-order nonconvex optimization.","The equivalence holds for maps into every reflexive Banach space, so finite-dimensional intuition about the target space is not essential."],"supporting_citations":[{"why":"Supplies the original Hessian-free inequality and the forward implication (i)$\\Rightarrow$(ii), which the paper uses without re-proving.","marker":"[7]"},{"why":"The Baillon–Haddad theorem quoted as Theorem 2 converts convexity of the quadratic perturbation into Fréchet differentiability and cocoercivity of the slice gradient.","marker":"[2]"},{"why":"Provides the Hahn–Banach corollary used to express norms by suprema over the dual unit ball when gluing the slice derivatives.","marker":"[4]"}],"fun_headline_variants":["Inequality alone yields Lipschitz derivative","Hessian-free inequality implies Lipschitz derivative","Jensen-type bound forces Fréchet differentiability","No Hessian required for Lipschitz derivative"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument collapses if the target space $Y$ is not reflexive: the proof constructs the derivative at each point only by identifying a certain element of the second dual $Y^{**}$ with an actual vector in $Y$, and reflexivity is exactly the property that makes this identification valid.","fun_headline_variants_meta":{"raw":{"variants":["Inequality alone yields Lipschitz derivative","Hessian-free inequality implies Lipschitz derivative","Jensen-type bound forces Fréchet differentiability","No Hessian required for Lipschitz derivative"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000771,"raw_usage":{"total_tokens":3391,"prompt_tokens":896,"completion_tokens":2495,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":2434}},"tokens_in":512,"tokens_out":2495,"duration_ms":17742,"temperature":1.0,"reasoning_tokens":2434,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:48:45.120803+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The decisive test is the two-point case with weights $1/2,1/2$. For $F(x)=|x|^\\alpha$ on $\\mathbb{R}$ with $0<\\alpha<2$, taking $y=0$ gives $|F(x/2)-F(x)/2|=(2^{-\\alpha}-1/2)|x|^\\alpha$, which must be no larger than $\\frac{L}{8}|x|^2$; letting $x\\to 0$ shows no finite $L$ works, so non-differentiable Hölder maps are excluded. A continuous non-Fréchet-differentiable map from a Hilbert space to a reflexive Banach space that satisfies (1.2) for some $L$ would refute Theorem 1, and the scaling calculation shows why the inequality is strong enough to rule out the simplest nonsmooth candidates.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Baillon–Haddad theorem quoted as Theorem 2 converts convexity of the quadratic perturbation into Fréchet differentiability and cocoercivity of the slice gradient."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Hahn–Banach corollary used to express norms by suprema over the dual unit ball when gluing the slice derivatives."}],"review_version":1}