{"id":"9ba06c20-fead-4c3c-999b-4685697c9aac","arxiv_id":"2603.09581","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Adam achieves local linear convergence on highly degenerate polynomials without schedulers through second-moment decoupling that exponentially amplifies the effective learning rate.","lead":"The paper claims Adam auto-converges with local linear rates on a class of highly degenerate polynomials without learning-rate schedulers, via a second-moment decoupling that amplifies the effective step size. This would explain a concrete regime where Adam’s adaptive moments beat Gradient Descent and Momentum by design rather than by tuning.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Provided full text is an unrelated robotics paper, so Adam claims remain completely unverifiable.","rationale":"The Reader correctly diagnosed the manuscript mismatch and correctly refused to score soundness or reproducibility. No further load-bearing technical concern inside the Adam argument can be raised because that argument is not present. The only honest action is to leave the verdict UNVERDICTED until the correct full text is supplied. The Reader’s weakest-assumption remark about the polynomial class is therefore moot; it presupposes a manuscript that was never given. Agreement is total; no adjustment is warranted.","tokens_in":12407,"tokens_out":394,"duration_ms":8317,"concrete_test":"Retrieve the actual PDF or source of arXiv:2603.09581 (the Adam paper) and re-run the review pipeline on that text; if the correct manuscript still fails to supply an explicit polynomial class, local-stability Lyapunov or contraction argument, and a quantitative comparison of linear vs sub-linear rates, keep UNVERDICTED; otherwise re-evaluate.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim (local linear convergence of Adam on a class of highly degenerate polynomials via v_t / g_t^{2} decoupling, without schedulers) cannot be inspected at all. The CACHEABLE PAPER SOURCE CONTEXT and full manuscript text belong to an entirely different work (terrain-aware CBF-MPC locomotion for quadrupeds, arXiv 2603.09585). No definitions of the polynomial family, no stability conditions, no proof of linear rate, no decoupling argument, no phase diagram, and no experiments for Adam appear. The abstract alone supplies only high-level assertions; every technical step required for the claim is missing. This is not a soft spot inside an argument—it is the total absence of the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The abstract claims that Adam exhibits natural auto-convergence (without external schedulers or β₂ near 1) on a class of highly degenerate polynomials, that local asymptotic stability conditions can be derived for these functions, that Adam attains local linear convergence via a decoupling of the second moment v_t from g_t² that exponentially amplifies the effective step size, and that a hyperparameter phase diagram with three regimes (stable convergence, spikes, SignGD-like oscillation) can be characterized. The supplied full manuscript body, however, is an unrelated robotics paper on proprioceptive 2.5-D terrain estimation, coupled contact/state estimation, and CBF-MPC safety constraints for quadrupedal locomotion (Unitree Go1). No definitions of the polynomial class, stability conditions, linear-rate proofs, decoupling argument, phase diagram, or Adam experiments appear anywhere in the body.","tokens_in":12604,"tokens_out":665,"duration_ms":9313,"significance":"If the abstract claims were substantiated by correct proofs and experiments, the result would be of genuine interest to the optimization and deep-learning communities: an intrinsic regime in which Adam’s adaptive second-moment mechanism yields a linear rate that Gradient Descent and Momentum cannot match, without artificial schedulers. Because the manuscript body contains none of the claimed analysis, the significance of the actual submission cannot be assessed and is currently zero.","major_comments":[{"comment":"Title/abstract versus body mismatch: the entire technical content (Sections I–VII, Algorithms 1, Eqs. (1)–(19), Tables I–II, Figs. 1–10) belongs to a different work on terrain-aware quadruped locomotion (arXiv:2603.09585). None of the abstract’s load-bearing objects—highly degenerate polynomials, local asymptotic stability conditions for Adam, the v_t–g_t² decoupling mechanism, linear-rate proofs, or the three-regime phase diagram—are defined, stated, or proved. The central claim is therefore completely unsupported by the submitted manuscript.","section":null},{"comment":"Absence of any verifiable derivation or experiment for Adam: without the polynomial family, the discrete dynamical system of Adam iterates, the Jacobian or Lyapunov analysis that would establish local linear convergence, or the numerical alignment of theoretical bounds with runs, the abstract’s assertions cannot be checked for correctness, circularity, or scope. This is not a local gap; it is the total absence of the argument required by the title.","section":null}],"minor_comments":[{"comment":"Even the abstract alone leaves the polynomial class, non-degeneracy conditions, and basin of attraction unspecified; once a correct body is supplied these definitions will need to be stated precisely.","section":null}],"recommendation":"reject","confidential_remarks":"The supplied PDF/text is almost certainly a copy-paste or arXiv-ID error (2603.09581 vs 2603.09585). The authors should be asked to resubmit the correct manuscript; until then there is nothing to review. I would not treat this as scientific misconduct, only as a fatal production error that forces rejection of the present submission."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The situation is simple: the abstract for 2603.09581 promises a clean story—Adam gets local linear convergence on a class of highly degenerate polynomials without schedulers or β₂→1, via a v_t–g_t² decoupling that amplifies the effective step size, plus a three-regime hyperparameter phase diagram—while GD and Momentum stay sublinear. That would be a useful, concrete addition to the “when does Adam actually help” literature if the proofs and experiments are there.\n\nThey are not here. The full manuscript text is a completely different paper (terrain-aware CBF-MPC locomotion for quadrupeds, 2603.09585). No polynomial family, no stability conditions, no linear-rate argument, no decoupling derivation, no phase diagram, no Adam experiments. The reader and the stress-test are right: this is not a soft spot inside an argument; it is the absence of the argument.\n\nFrom the abstract alone one can say the claimed contribution is well-scoped and not obviously tautological, and the free parameters (Adam hyperparameters, degeneracy parameters) look like the right ones to track. That is all. Soundness, novelty relative to prior analyses, and experimental alignment cannot be checked. I would not bring this to reading group or cite it until the correct PDF is in hand. A serious editor should not send the current package to referees; they should request the right manuscript. Once that arrives, the abstract is interesting enough to deserve a real look. Right now there is nothing to engage with.","headline":"We only have the Adam abstract; the supplied full text is an unrelated robotics paper, so every technical claim is unverifiable.","tokens_in":13217,"tokens_out":400,"would_cite":false,"duration_ms":7332,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Adam automatically converges at a local linear rate on highly degenerate polynomials, without schedulers, by decoupling its second-moment estimate from the squared gradient and exponentially amplifying the effective step size.","keywords":["Adam","adaptive optimization","degenerate polynomials","local linear convergence","second-moment decoupling","hyperparameter phase diagram","auto-convergence"],"falsifier":"On any concrete member of the claimed polynomial class, measure the local contraction rate of Adam versus Gradient Descent/Momentum from a neighbourhood of a minimiser; if Adam fails to exhibit linear convergence while the others remain sub-linear, or if the observed stability region violates the derived theoretical bounds, the central claim is false.","tokens_in":13293,"feed_emoji":"📈","tokens_out":622,"duration_ms":12811,"temperature":0.7,"pith_summary":"Adam is ubiquitous in deep learning, yet it remains unclear for which objectives it is inherently better than simpler methods. This paper isolates a class of highly degenerate polynomials on which Adam converges by itself—no external learning-rate schedules and no need to push the second-moment decay near 1. On these functions Adam attains a local linear rate, while Gradient Descent and Momentum remain only sub-linear. The speed-up is traced to a decoupling: the second-moment accumulator v_t drifts away from the squared gradient g_t^{2}, which in turn makes the effective learning rate grow exponentially. The same analysis yields a clean hyper-parameter phase diagram with three regimes—stable convergence, spikes, and SignGD-like oscillation—and the theoretical stability bounds match numerical experiments closely. A sympathetic reader cares because the result supplies the first concrete, scheduler-free setting in which Adam’s adaptivity is provably superior.","feed_headline":"Adam hits linear rate on degenerate polynomials, no scheduler","feed_subtitle":"Decoupling of second moment from squared gradient exponentially boosts the effective step size","key_machinery":"The decoupling mechanism between the second-moment accumulator v_t and the squared gradient g_t^{2}: once v_t lags behind g_t^{2} the adaptive denominator shrinks, the effective step size grows exponentially, and local linear convergence follows.","core_discovery":"There exists a class of highly degenerate polynomials for which the plain Adam iteration (fixed step-size, fixed β₂ away from 1) is locally asymptotically stable and converges linearly; the linear rate is produced by a spontaneous decoupling of the second-moment estimate v_t from the squared gradient, which exponentially inflates the effective learning rate and thereby outperforms the sub-linear rates of Gradient Descent and Momentum.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Adam auto-converges linearly on highly degenerate polynomials","Plain Adam yields local linear rate on degenerate polynomials","v_t decoupling drives Adam linear convergence on degenerates","Adam linear-stable on degenerate polynomials without schedulers","Second-moment decoupling gives Adam linear rate over GD Momentum"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The identified family of highly degenerate polynomials is both mathematically well-defined and rich enough that the linear-rate advantage and the claimed decoupling transfer beyond the abstract examples studied.","fun_headline_variants_meta":{"raw":{"variants":["Adam auto-converges linearly on highly degenerate polynomials","Plain Adam yields local linear rate on degenerate polynomials","v_t decoupling drives Adam linear convergence on degenerates","Adam linear-stable on degenerate polynomials without schedulers","Second-moment decoupling gives Adam linear rate over GD Momentum"]},"model":"grok-4.5","effort":"low","cost_usd":0.00366,"raw_usage":{"total_tokens":1118,"prompt_tokens":714,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":36600000,"prompt_tokens_details":{"text_tokens":714,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":325,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":714,"tokens_out":79,"duration_ms":3125,"temperature":1.0,"reasoning_tokens":325,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T00:07:22.609367+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On any concrete member of the claimed polynomial class, measure the local contraction rate of Adam versus Gradient Descent/Momentum from a neighbourhood of a minimiser; if Adam fails to exhibit linear convergence while the others remain sub-linear, or if the observed stability region violates the derived theoretical bounds, the central claim is false.","supporting_citations":[],"review_version":1}