{"id":"b18e86c4-4722-490e-86bf-ddd30593a8e3","arxiv_id":"2604.14669","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Mean-square linear stability of two-point ZO methods is governed by the full Hessian spectrum and admits explicit bounds in terms of trace and top eigenvalue; full-batch ZO-GD/GDM/Adam empirically operate at that boundary.","lead":"Zeroth-order optimizers that use only function values, not gradients, become unstable according to the full Hessian spectrum (especially its trace), not just the top eigenvalue. This predicts and explains why full-batch ZO training of neural nets sits at a mean-square edge of stability and implicitly regularizes the Hessian trace.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim is tightly supported: the mean-square critical step size is characterized exactly by an explicit spectral equation (Theorems 1–3), the tractable bounds depend only on trace and λ_max, and the full-batch experiments show the predicted threshold remaining inside those intervals across CNN/ResNet/ViT and sequence models. The quadratic linearization is the same modeling step used throughout the FO EoS literature; the commutativity assumption is both stated and empirically checked. Because the paper already flags the full-batch restriction and the smoothing-bias interaction as future work, there is no load-bearing gap that would change an ACCEPT verdict. The concrete test above is a useful verification of the bound tightness but is not expected to falsify the claim.","tokens_in":41587,"tokens_out":451,"duration_ms":4726,"concrete_test":"On the CNN CIFAR-10 setup of Figure 2, recompute the exact mean-square critical step size η*_ms by solving the full-spectrum equation of Theorem 1 (via a few hundred Hutchinson/power-iteration probes for the leading eigenvalues) at several late-training checkpoints and verify that the observed 2/η continues to lie inside the reported [Tr(H), Tr(H)+2λ_max] interval; a systematic violation would indicate that the tractable bounds no longer track the true boundary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (quadratic linearization + P H = H P for frozen ZO-Adam) is correctly identified and is standard for EoS-style analyses. The paper already supplies the exact spectral-radius characterizations (Theorems 1–3 via the cone-preserving covariance operator and Krein–Rutman), the matching tractable bounds used in the experiments, and direct empirical support that the relative commutator falls below 0.05 and stays there (Appendix D.3). Higher-order terms and mini-batch noise are acknowledged as open extensions rather than hidden contradictions. No internal inconsistency or circularity appears that would overturn the central claim under the stated full-batch linearized setting.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper derives exact mean-square linear stability conditions for ZO-GD, ZO-GDM, and frozen ZO-Adam under the standard two-point Gaussian estimator applied to the local quadratic model of the loss. Unlike FO methods, whose critical step size depends only on λ_max, the ZO mean-square critical step size is the unique positive root of an explicit spectral equation involving the full Hessian (or preconditioned Hessian) spectrum; tractable bounds depending only on Tr(H) and λ_max are also obtained. Empirically, full-batch ZO-GD, ZO-GDM, and ZO-Adam on CNN/ResNet/ViT (CIFAR-10 subset) and on LSTM/Mamba sorting tasks stabilize so that the predicted thresholds 2/η, 2(1-β)/η, and 2/η remain inside the corresponding [Tr, Tr+2λ_max]-type intervals throughout training, indicating a mean-square edge-of-stability regime that primarily regularizes the Hessian trace.","tokens_in":41788,"tokens_out":813,"duration_ms":6695,"significance":"The work supplies a clean, FO-contrasting stability theory for the ZO estimators used in MeZO-style LLM fine-tuning, together with complete second-moment recursions, a cone-preserving covariance operator analyzed via Krein–Rutman (Theorem 4 / Appendix A), and matching tractable bounds that are directly plotted against independently estimated curvature. The empirical mean-square EoS observation across architectures, the opposite momentum dependence relative to FO-GDM, and the explicit link to trace regularization are new and practically relevant. Strengths include machine-checkable spectral characterizations, falsifiable threshold predictions with no free fit parameters, and supporting commutator measurements (Appendix D.3).","major_comments":[],"minor_comments":[{"comment":"In §4.2 / Remark 2 the commutativity assumption PH=HP is load-bearing for the exact spectral formula of frozen ZO-Adam; while Appendix D.3 shows the relative commutator drops below 0.05, a short forward pointer in the main text would help readers who skip the appendix.","section":"§4.2, Remark 2"},{"comment":"Figure 1 caption and the surrounding text state that the ZO-GD trace stabilizes 'slightly below 2/η'; the later theory (Eq. (5)) places the threshold inside [Tr, Tr+2λ_max]. Aligning the early narrative with the precise interval would avoid a minor inconsistency of language.","section":"Figure 1, §1"},{"comment":"Appendix B sketches extensions to forward differences, non-Gaussian directions, and multi-query averages; a one-sentence pointer in the main text (e.g., after Theorem 1) would make these results more discoverable.","section":"Appendix B"},{"comment":"The squared-loss / one-hot CIFAR-10 subset protocol is standard for EoS studies but should be flagged more prominently when claiming relevance to cross-entropy LLM fine-tuning.","section":"§5.2"}],"recommendation":"accept","confidential_remarks":"The manuscript is already at a high technical standard; the empty major-comments list is intentional. The only residual risk is that the community may over-extrapolate the full-batch quadratic theory to mini-batch ZO LLM fine-tuning; the authors already flag this as future work, so no action is required beyond the minor presentation notes."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This paper gives the first exact mean-square linear-stability thresholds for the two-point ZO methods people actually use (ZO-GD, ZO-GDM, frozen ZO-Adam). The thresholds depend on the whole spectrum, not just λ_max, and the tractable bounds that only need Tr(H) and λ_max are what they track in the experiments. That is the real novelty relative to both the FO EoS literature and classical ZO convergence work.\n\nThe math is solid. They reduce the second-moment recursion to a cone-preserving covariance operator with a rank-one coupling, then apply Krein–Rutman to get the spectral radius exactly (Theorem 4 and the special cases). The proofs in Appendix A are complete; the bounds follow by elementary majorization. Empirically they show, across CNN/ResNet/ViT and LSTM/Mamba, that full-batch ZO training sits inside the predicted interval for long stretches, with the trace term doing most of the work. The opposite momentum dependence (increasing β shrinks the ZO stable region) is a clean, non-obvious contrast with FO-GDM. Commutativity for frozen ZO-Adam is checked directly (relative commutator drops below 0.05 and stays there).\n\nSoft spots are real but proportionate. Everything is linearized on the local quadratic and full-batch; mini-batch noise and higher-order terms are left open. Smoothing μ is treated only empirically. None of that undercuts the claims they actually make under the stated setting. Citations are appropriate; no circularity or free parameters in the equalities.\n\nThis is for people who care about ZO dynamics, memory-efficient fine-tuning theory, or EoS-style implicit bias. It deserves a serious referee. I would bring it to reading group and expect to cite the stability formulae.","headline":"Clean first exact mean-square stability theory for practical ZO methods, with matching full-batch EoS evidence that the governing quantity is the Hessian trace rather than λ_max.","tokens_in":42343,"tokens_out":481,"would_cite":true,"duration_ms":6599,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Zeroth-order optimizers train at a mean-square edge of stability set by the full Hessian spectrum, not just its top eigenvalue.","keywords":["zeroth-order optimization","edge of stability","mean-square linear stability","Hessian spectrum","Hessian trace","two-point estimator","ZO-GD","ZO-Adam"],"falsifier":"Train full-batch ZO-GD on a neural net and track Hessian trace and top eigenvalue: if the fixed threshold 2/η systematically leaves the interval [Tr(Ht), Tr(Ht)+2λ max(Ht)] for a sustained portion of training, the mean-square edge-of-stability claim fails.","tokens_in":42535,"feed_emoji":"⚡","tokens_out":752,"duration_ms":6389,"temperature":0.7,"pith_summary":"Zeroth-order (ZO) optimizers update parameters from function values alone, via random two-point probes. Their short-term stability cannot be read from the familiar first-order rule that depends only on the largest Hessian eigenvalue. This paper derives the exact step-size threshold that keeps the second-moment of the linearized ZO iterates bounded; that threshold is an explicit spectral equation involving every eigenvalue. Tractable bounds that use only the trace and the top eigenvalue follow at once. Full-batch experiments on CNNs, ResNets, ViTs and sequence models show that ZO-GD, ZO-GDM and ZO-Adam settle so that the theoretical threshold remains inside those bounds throughout training. Large step sizes therefore push ZO trajectories toward small-trace Hessians, an implicit bias distinct from the sharpness control seen in first-order training.","feed_headline":"ZO optimizers sit at a stability edge set by Hessian trace","feed_subtitle":"Large step sizes regularize the full spectrum, not just the top eigenvalue, unlike first-order methods.","key_machinery":"A cone-preserving linear covariance operator with a rank-one global coupling term that evolves the second-moment blocks of the ZO iterates; its spectral radius is characterized by the Krein–Rutman theorem and yields the exact mean-square stability condition.","core_discovery":"For the standard Gaussian two-point estimator the mean-square critical step size of ZO-GD, ZO-GDM and frozen ZO-Adam is the unique positive root of an explicit equation that involves the entire spectrum of the (preconditioned) Hessian; the same methods empirically stabilize so that this threshold stays inside the computable interval built from the Hessian trace and top eigenvalue.","pith_inferences":["Mini-batch ZO-SGD should inherit a hybrid stability threshold that mixes estimator noise with sampling noise and may settle at still flatter trace values.","A central-flow description of ZO training could be obtained by replacing the FO sharpness flow with the mean-square spectral equation derived here.","Because ZO fine-tuning already reduces memory, the trace bias may offer a free generalization knob that FO memory-efficient methods do not share."],"forward_implications":["Large ZO step sizes primarily regularize Hessian trace rather than top eigenvalue.","Increasing momentum shrinks the ZO mean-square stable region, the opposite of its effect on first-order momentum methods.","Practical ZO training can monitor only trace and top eigenvalue to stay near the stability boundary without computing the full spectrum.","The same covariance-operator framework extends immediately to forward differences, non-Gaussian directions and multi-query averages."],"fun_headline_variants":["ZO methods stabilize at edge set by full Hessian spectrum","ZO stability bound depends on entire Hessian, not just top eigenvalue","Large ZO steps regularize Hessian trace unlike FO top-eigenvalue focus","ZO-GD rides mean-square stability edge fixed by Hessian spectrum","Full-batch ZO optimizers sit at predicted trace-based stability boundary"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The exact thresholds are derived for the local quadratic Taylor model of the loss, and for ZO-Adam the frozen preconditioner is assumed to commute with the Hessian.","fun_headline_variants_meta":{"raw":{"variants":["ZO methods stabilize at edge set by full Hessian spectrum","ZO stability bound depends on entire Hessian, not just top eigenvalue","Large ZO steps regularize Hessian trace unlike FO top-eigenvalue focus","ZO-GD rides mean-square stability edge fixed by Hessian spectrum","Full-batch ZO optimizers sit at predicted trace-based stability boundary"]},"model":"grok-4.5","effort":"low","cost_usd":0.004442,"raw_usage":{"total_tokens":1313,"prompt_tokens":763,"num_sources_used":0,"completion_tokens":76,"cost_in_usd_ticks":44420000,"prompt_tokens_details":{"text_tokens":763,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":474,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":763,"tokens_out":76,"duration_ms":4203,"temperature":1.0,"reasoning_tokens":474,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T20:05:58.434900+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train full-batch ZO-GD on a neural net and track Hessian trace and top eigenvalue: if the fixed threshold 2/η systematically leaves the interval [Tr(Ht), Tr(Ht)+2λ max(Ht)] for a sustained portion of training, the mean-square edge-of-stability claim fails.","supporting_citations":[],"review_version":2}