{"id":"6455fe4e-b3b8-4168-a7d0-a2272f5c5b7c","arxiv_id":"2508.06002","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Kahan's KGD step-size is shown to converge at least R-linearly with rate 1-1/cond(H) for quadratics, and an adaptive generalization for general optimization is proved and tested.","lead":"A new analysis of Kahan's gradient descent step-size shows it is equivalent to the long Barzilai-Borwein step for quadratics, with a proven linear convergence rate depending on the condition number. The paper also proposes an adaptive framework for general unconstrained optimization and tests it on standard benchmark problems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"General-case convergence depends on unspecified step-size safeguards; the supplied full text is corrupted and cannot support the adaptive-framework proof.","rationale":"Read in good faith, the quadratic-model claim is standard: the long BB step is exact line search, and the stated rate 1 - 1/cond(H) is weaker than the classical (κ-1)/(κ+1)^2 bound, so it is not a likely source of error. The genuinely load-bearing claim is the general adaptive framework. The reader identified the quadratic-to-general transfer as the weakest assumption; I agree. The supplied full text is not merely lacking proofs—it is corrupted and interleaved with an unrelated arXiv identifier, so the framework's definition, safeguards, and assumptions cannot be inspected. I therefore do not assert a mathematical flaw; I assert that the load-bearing premise is unverified. This leaves the reader's UNVERDICTED verdict unchanged. No ad hominem is intended: the issue is the evidence available, not the authors' integrity.","tokens_in":3869,"tokens_out":14432,"duration_ms":155455,"concrete_test":"Retrieve the clean TeX source of arXiv:2508.06002. In the proof of the general global-convergence theorem, locate the explicit bound on the adaptive step-size sequence. Verify that a lemma shows α_k ∈ [α_min, α_max] with α_min>0 and α_max<2/L (or that a sufficient-decrease condition is enforced before each KGD step is accepted). If such a bound is present, the proposed framework is plausible; if the proof instead transfers the quadratic Rayleigh-quotient recurrence without a safeguard, run the method on a Dai–Fletcher-type non-quadratic example and check whether the gradient norm fails to converge to zero.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is the adaptive framework that turns the quadratic-model KGD step-size into a globally convergent method for general unconstrained minimization. The abstract leaves the actual safeguards unspecified, and the supplied full text cannot fill that gap: it is mostly unreadable mojibake and contains the line 'arXiv:2508.06013v2 [physics.acc-ph] 14 Aug 2025,' an identifier from a different paper. Because the quadratic equivalence is to the long BB step, and unsafeguarded long-BB steps are known to fail on smooth non-quadratic problems (Dai–Fletcher examples), the general convergence theorem must rely on a step-size bound or a descent/backtracking safeguard. Without access to that safeguard and its proof, the global-convergence claim is unsupported. This is not a claim of error; it is an unverifiable premise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims to provide a rigorous analysis of Kahan's gradient descent (KGD) step-size strategies. It states that, for strongly convex quadratic objectives, the KGD step-size is equivalent to the long Barzilai-Borwein (BB) step, and that both a newly derived short KGD variant and the long variant (hence BB) converge R-linearly with rate 1 - 1/cond(H). For general unconstrained minimization, it proposes an adaptive framework that uses KGD step-sizes and proves global convergence and local R-linear convergence. Numerical experiments on CUTEst and logistic-regression problems are said to compare the proposed methods against BB variants and other adaptive gradient methods.","tokens_in":4112,"tokens_out":5059,"duration_ms":61143,"significance":"If the theorems can be established, the paper would contribute a useful theoretical bridge between Kahan's KGD step-size and the classical BB step, including a short-step variant with a clean R-linear rate in the quadratic case. The adaptive framework for general objectives could also be practically relevant, provided the convergence proof is valid and the safeguards are explicit. The topic is of interest to the unconstrained-optimization community. However, the submitted full text is unreadable, so none of these contributions can actually be verified from the manuscript as supplied. No code, data, or machine-checked proofs are supplied; the only concrete content is the abstract.","major_comments":[{"comment":"The supplied body text is unreadable mojibake throughout, with a stray line 'arXiv:2508.06013v2 [physics.acc-ph] 14 Aug 2025' from an unrelated paper embedded in the first page. No equation, algorithm, theorem statement, or proof is legible. This is not a local formatting issue: the paper's central claims—the KGD/long-BB equivalence, the R-linear rate, and the adaptive-framework convergence—cannot be checked in any way. A clean, complete manuscript is a prerequisite for review.","section":"Full Text (all sections)"},{"comment":"The central contribution is the adaptive framework for general unconstrained minimization, but the abstract does not specify the safeguards or parameter choices on which the convergence proof rests. Because the KGD step-size is equivalent to the long BB step only in the quadratic model, and unsafeguarded long-BB steps are known to fail on smooth non-quadratic problems (Dai-Fletcher examples), the global-convergence theorem must depend on a step-size bound, a descent/backtracking condition, or another safeguard. No such condition is stated in the abstract, and the full text provides no readable proof. This is a load-bearing gap for the paper's main claim.","section":"Abstract, 'For the general unconstrained minimization...'"},{"comment":"The claimed R-linear rate is stated without proof or even definitions of the step-size recursions, the iteration matrices, and the norm/metric in which the rate holds. The asserted mathematical equivalence between KGD and the long BB step is also not demonstrated in any legible form. To verify this theorem, the authors need to present the exact recursions, the error evolution, and the argument leading to the stated rate. As submitted, this is unverifiable.","section":"Abstract, 'both the long and short KGD ... rate 1 - 1/cond(H)'"},{"comment":"The abstract promises extensive comparisons on CUTEst and logistic-regression problems, but the supplied text contains no tables, performance profiles, or reproducibility details. Since the paper also claims to demonstrate efficiency and robustness, the numerical evidence is part of the support for those claims; its absence in the submitted file prevents assessment of the experimental component.","section":"Abstract, numerical experiments"}],"minor_comments":[{"comment":"The notation cond(H) should be defined as the spectral condition number (or explicitly as the ratio of extreme eigenvalues) when the full text is restored.","section":"Abstract"},{"comment":"The reference to Kahan's 2019 technical report is indicated, but the report's availability should be made explicit (e.g., a stable URL) so readers can consult the original KGD step-size strategy.","section":"Abstract"},{"comment":"When the manuscript is restored, the authors should ensure that section headings, theorem environments, and equation numbering are visible and consistent, since none can currently be identified.","section":"Full Text"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the supplied full text is corrupted/unreadable and even contains an arXiv identifier from a different paper. This is a serious submission-integrity/file-format problem and warrants desk-return or a request for a clean, complete manuscript before further review. The abstract's mathematical claims may well be sound, but they cannot be refereed from the present file; the load-bearing proof of the adaptive framework is entirely absent in readable form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: if the abstract is accurate, this is a useful contribution, but I can't verify it from what was sent. The supplied full text is mojibake, with an unrelated arXiv identifier embedded in the middle, so my read is based on the abstract alone. That's not a mark against the math, but it caps what I can fairly say.\n\nThe genuinely new pieces are the short KGD step-size, the R-linear rate 1 - 1/cond(H) for strongly convex quadratics, and an adaptive framework that claims global and local R-linear convergence for general unconstrained problems. To the authors' credit, they don't present the quadratic equivalence to the long BB step as their own discovery; they attribute it to Kahan's 2019 report and build from there. That is honest and sensible. If the proofs are clean, this gives a rigorous rate for Kahan's heuristics and a nice companion to the known behavior of BB methods.\n\nThe soft spot is the general-case convergence claim. Unsafeguarded long-BB steps are known to fail on smooth non-quadratic problems (Dai-Fletcher examples), so the adaptive framework must include something—a step-size bound, backtracking, or a descent safeguard—that makes global convergence possible. The abstract doesn't say what it is, and the text I was given doesn't either, because it isn't readable. I don't read that as evidence of error; it's the main thing a referee needs to check. I'd also want to see the CUTEst and logistic-regression comparisons presented in a way that doesn't rely on a small hand-picked subset, but the abstract at least suggests standard practice.\n\nBottom line: send it to a serious referee with a request to focus on the adaptive framework's assumptions and the proof for the short variant. The quadratic part is probably fine; the general case decides the value. I wouldn't cite it until I've seen and verified the full version.","headline":"Plausible and potentially useful formalization of Kahan's step-size, but the supplied text is unreadable and the general-case convergence hinges on safeguards I cannot see; worth a real referee, not a desk reject.","tokens_in":4510,"tokens_out":4160,"would_cite":false,"duration_ms":45829,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C53","65K05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Kahan's step-size-updating gradient rule is shown to be equivalent to the long Barzilai-Borwein step in the quadratic model, and an adaptive version is proven to converge globally with local R-linear rate on general smooth unconstrained pro","keywords":["Barzilai-Borwein method","Kahan gradient descent","step-size control","unconstrained optimization","R-linear convergence","quadratic model","CUTEst","logistic regression"],"falsifier":"For the quadratic rate bound, simulate the KGD/BB recurrence on $f(x)=\\frac12\\sum_{i=1}^n d_i x_i^2$ with $d_1=1$, $d_n=\\kappa$, and measure the worst-case asymptotic contraction $\\limsup_{k\\to\\infty}\\|x_k-x^*\\|^{1/k}$ over starting points; exceeding $1-1/\\kappa$ would disprove the rate theorem. For the general claim, run the adaptive KGD on a smooth strongly convex non-quadratic function with a known optimum and check whether the safeguarded step-size sequence stays within the interval the proof requires; divergence or stagnation while the paper's assumptions hold would refute global converge","tokens_in":3832,"feed_emoji":"📉","tokens_out":8823,"duration_ms":98737,"temperature":0.7,"pith_summary":"This paper tries to give Kahan's proposed Gradient Descent (KGD) step-size rule a rigorous theoretical footing. The authors prove that in the strongly convex quadratic model the KGD step size is exactly the long Barzilai-Borwein (BB) step, and they derive a short KGD variant; both long and short KGD (hence BB) converge at least R-linearly with rate $1-1/\\mathrm{cond}(H)$. For general unconstrained minimization they propose an adaptive KGD framework and prove global convergence plus local R-linear convergence. Numerical tests on the CUTEst collection and logistic regression problems indicate that the resulting methods compete with standard BB variants and recent adaptive gradient methods. If the proofs are right, a step-size rule previously treated as a heuristic now carries the same worst-case guarantees as BB.","feed_headline":"Kahan step-size rule matched to long BB; convergence proved","feed_subtitle":"Adaptive KGD gains global and local convergence guarantees; quadratic rate depends on condition number.","key_machinery":"The KGD step-size recurrence, which updates the step size from its previous value and the observed gradient change instead of recomputing it each iteration. In the quadratic model this recurrence is mathematically identical to the long BB step $\\alpha_k^{\\mathrm{BB1}}=\\|s_{k-1}\\|^2/(s_{k-1}^Ty_{k-1})$, with $s_{k-1}=x_k-x_{k-1}$ and $y_{k-1}=g_k-g_{k-1}$; the short version corresponds to $s_{k-1}^Ty_{k-1}/\\|y_{k-1}\\|^2$. This equivalence lets BB convergence machinery be imported into KGD, and the adaptive framework's proof controls the step-size recurrence for non-quadratic objectives.","core_discovery":"The central claim is that Kahan's step-size iteration is not merely a heuristic cousin of BB. For any strongly convex quadratic with Hessian $H$, the KGD step size equals the long BB step, so the step-size recurrence inherits the entire quadratic convergence theory of BB; the paper also constructs a short KGD step, and both variants have worst-case error contraction at least $1-1/\\mathrm{cond}(H)$. Outside quadratics, the paper's adaptive framework—with safeguards on the step-size recurrence—is proved to be globally convergent and locally R-linearly convergent for general unconstrained smooth minimization.","pith_inferences":["Editorial inference: because the quadratic rate $1-1/\\mathrm{cond}(H)$ is the same as plain steepest descent's worst-case rate, KGD's advantage is unlikely to be a better tail rate on ill-conditioned quadratics; the practical gain must come from avoidance of line-search cost or better early behavior, which the paper's experiments address.","Editorial inference: the proof's dependence on the quadratic-model equivalence suggests a direct test of the safeguards—run the raw (unsafeguarded) KGD recurrence on a smooth non-quadratic strongly convex function; if it diverges while the safeguarded version converges, the safeguards are the load-bearing ingredient.","Editorial inference: the long/short BB equivalence opens a two-way street with known BB results, so existing special-case facts about BB could be re-derived for KGD, giving quick testable predictions."],"forward_implications":["In quadratic models, KGD and BB step-size choices are interchangeable: analyses, worst-case rates, and any future improvements for one transfer to the other.","The derived short KGD step provides a second step-size strategy with the same $1-1/\\mathrm{cond}(H)$ worst-case rate, giving practitioners a parameter-free alternative to long-step choices.","The adaptive framework extends Kahan's recurrence to general smooth unconstrained problems with global and local R-linear convergence guarantees, so the method can be used outside the quadratic setting without relying on heuristic rationale.","The numerical comparisons on CUTEst and logistic regression suggest the safeguarded KGD variants remain competitive with established BB and adaptive gradient methods on real problems."],"supporting_citations":[],"fun_headline_variants":["Kahan step-size equals long BB; convergence proved","Adaptive KGD: global convergence plus local R-linear rate","Kahan step-size recurrence: quadratic equivalence to long BB","KGD step-size matches BB; linear convergence guaranteed"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The general-case proof assumes that a step-size recurrence whose exact equivalence to the long BB step is proved only for quadratic models remains well-behaved for arbitrary smooth non-quadratic objectives once the proposed adaptive safeguards are added; if some allowed objective defeats those safeguards, the global convergence claim falls.","fun_headline_variants_meta":{"raw":{"variants":["Kahan step-size equals long BB; convergence proved","Adaptive KGD: global convergence plus local R-linear rate","Kahan step-size recurrence: quadratic equivalence to long BB","KGD step-size matches BB; linear convergence guaranteed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000408,"raw_usage":{"total_tokens":1978,"prompt_tokens":793,"completion_tokens":1185,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":1119}},"tokens_in":537,"tokens_out":1185,"duration_ms":10324,"temperature":1.0,"reasoning_tokens":1119,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:59:02.861671+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For the quadratic rate bound, simulate the KGD/BB recurrence on $f(x)=\\frac12\\sum_{i=1}^n d_i x_i^2$ with $d_1=1$, $d_n=\\kappa$, and measure the worst-case asymptotic contraction $\\limsup_{k\\to\\infty}\\|x_k-x^*\\|^{1/k}$ over starting points; exceeding $1-1/\\kappa$ would disprove the rate theorem. For the general claim, run the adaptive KGD on a smooth strongly convex non-quadratic function with a known optimum and check whether the safeguarded step-size sequence stays within the interval the proof requires; divergence or stagnation while the paper's assumptions hold would refute global converge","supporting_citations":[],"review_version":1}