{"id":"d9b9d2c7-59a6-4901-8461-b550412930e4","arxiv_id":"2605.21292","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Large constant learning rates in a two-factor linear transformer model can induce cycles, bounded chaos, or divergence rather than convergence to a single in-context linear-regression solution.","lead":"This paper examines the finite-step gradient descent dynamics of a simplified one-prompt linear transformer model. It shows that large learning rates can shift the training attractor to cycles, bounded chaos, or divergence instead of the expected in-context linear regression solution.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Whether normalization exactly reduces the one-prompt dynamics to a two-factor product map with parameter μ","rationale":"The reader's weakest_assumption pinpoints the single step on which every subsequent phase-diagram and stability statement depends. Because the manuscript is now accessible, the concrete algebraic verification above directly tests whether that step is exact. If it is, the attractor-changing claim stands on firm ground for the one-prompt case; if not, the headline conclusion requires qualification before extension to mini-batch training.","tokens_in":1739,"tokens_out":313,"duration_ms":18227,"concrete_test":"Starting from the explicit one-prompt loss and gradient updates in §2, apply the normalization transformation given in §3 and algebraically expand the resulting map for the two factors; confirm that all cross terms vanish and the update is exactly the claimed product map parameterized solely by μ.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that, after the stated normalization, the gradient-descent updates on the one-prompt linear-transformer loss become precisely the two-factor product map whose balanced slice recovers the cubic map and whose full 2D flow admits an explicit invariant Chebyshev ellipse. If this reduction is only approximate, or if it holds only for specific prompt statistics or initializations not stated in the derivation, then the sharp stability thresholds, transverse repulsion of the ellipse, and the conclusion that large constant steps can replace the in-context regression attractor with cycles/chaos/divergence do not follow for the original model.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper analyzes the finite-step gradient descent dynamics of a simplified one-prompt linear transformer model. It claims that after a normalization step, the training dynamics exactly reduce to a two-factor product map parameterized by an effective step size μ. On the balanced slice, this map reproduces the phase diagram of the scalar cubic map, including transitions to periodic, chaotic, and divergent behaviors. For the full 2D system, an explicit invariant Chebyshev ellipse is identified for 0 < μ < 2, which is transversely repelling while balanced attractors can be attracting. The results imply that large constant learning rates can alter the training attractor away from the in-context linear regression solution towards cycles, bounded chaos, or divergence.","tokens_in":1889,"tokens_out":573,"duration_ms":40136,"significance":"If the exact reduction holds, this provides a rigorous mathematical framework for understanding how large learning rates affect transformer training dynamics beyond the gradient-flow limit. It offers explicit stability thresholds and an invariant set analysis that could explain empirical instabilities in high-LR training. The explicit construction of the invariant ellipse and the transverse stability analysis are notable strengths, as is the connection to the known cubic map phenomenology without post-hoc fitting. This could inform the design of training schedules for transformers.","major_comments":[{"comment":"Abstract and the normalization step in the model section: the claim that the one-prompt linear-transformer gradient-descent updates reduce exactly to the two-factor product map with effective step-size μ is load-bearing for all stability thresholds and attractor conclusions, yet the algebraic cancellations that achieve this reduction after normalization are not shown explicitly; it is therefore unclear whether the reduction is exact for arbitrary prompt statistics or requires additional assumptions on data or initialization.","section":"Abstract and model/normalization section"},{"comment":"The 2D system analysis section: the assertion of an explicit invariant Chebyshev ellipse for 0<μ<2 that separates forward-invariant regions and is transversely repelling requires the explicit verification that the ellipse is mapped into itself and the computation of the transverse Lyapunov exponent or linearization; without these steps the conclusion that off-balanced chaotic dynamics do not attract from the balanced slice does not follow.","section":"Full 2D system analysis"}],"minor_comments":[{"comment":"The relation between the original learning rate and the effective parameter μ should be written as a single displayed equation immediately after the normalization is introduced.","section":"Notation and parameters"},{"comment":"Phase portraits or bifurcation diagrams for the 2D map would benefit from explicit annotation of the Chebyshev ellipse and the basins of attraction.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their thorough review and valuable feedback on our work. We address the major comments below and will incorporate the suggested clarifications in the revised manuscript.","responses":[{"response":"We appreciate this observation. Upon review, we recognize that while the reduction is derived in the manuscript, the intermediate algebraic steps were not presented in full detail. In the revised version, we will expand the model section to explicitly show the cancellations leading to the two-factor product map. This derivation holds for arbitrary prompt statistics under the normalization procedure described, without further assumptions on data or initialization beyond those stated in the paper.","revision_made":"yes","referee_comment":"[Abstract and model/normalization section] Abstract and the normalization step in the model section: the claim that the one-prompt linear-transformer gradient-descent updates reduce exactly to the two-factor product map with effective step-size μ is load-bearing for all stability thresholds and attractor conclusions, yet the algebraic cancellations that achieve this reduction after normalization are not shown explicitly; it is therefore unclear whether the reduction is exact for arbitrary prompt statistics or requires additional assumptions on data or initialization."},{"response":"We agree that a more explicit verification is necessary to fully support the claims regarding the invariant ellipse. In the updated manuscript, we will provide the step-by-step verification that the Chebyshev ellipse is mapped into itself under the dynamics for 0 < μ < 2. Additionally, we will include the computation of the transverse linearization and the associated Lyapunov exponent to rigorously demonstrate that the ellipse is transversely repelling, while balanced attractors remain attracting in the transverse direction. This will strengthen the conclusion that off-balanced chaotic dynamics do not attract trajectories starting from the balanced slice.","revision_made":"yes","referee_comment":"[Full 2D system analysis] The 2D system analysis section: the assertion of an explicit invariant Chebyshev ellipse for 0<μ<2 that separates forward-invariant regions and is transversely repelling requires the explicit verification that the ellipse is mapped into itself and the computation of the transverse Lyapunov exponent or linearization; without these steps the conclusion that off-balanced chaotic dynamics do not attract from the balanced slice does not follow."}],"tokens_in":1474,"tokens_out":476,"duration_ms":66291,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core result is that after normalization the one-prompt dynamics collapse exactly to a two-factor map with step-size parameter mu. On the balanced slice this recovers the cubic-map transitions from convergence to periodic, chaotic, or divergent behavior. In the full 2D system they exhibit an explicit invariant Chebyshev ellipse that carries off-balance chaos but repels transversely, while balanced fixed points can attract transversely. This is the new piece: a concrete 2D phase diagram that extends the scalar case and directly ties large constant learning rates to attractor switching rather than just faster convergence.","headline":"The paper reduces normalized one-prompt linear transformer training to an exact 2D product map, constructs an invariant Chebyshev ellipse that is transversely repelling, and shows large steps can replace the regression attractor with cycles or chaos.","tokens_in":2369,"tokens_out":205,"would_cite":false,"duration_ms":24852,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"After normalization, the dynamics reduce to a two-factor product map with an effective step-size parameter μ. On the balanced slice, this map recovers the known scalar cubic transition... the full two-dimensional system... has an explicit invariant Chebyshev ellipse... e+ = C(e) = e³−3e"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/ArithmeticFromLogic.lean","rs_theorem":"LogicNat recovery and embed_strictMono","paper_passage":"the map Φ_μ(a,b)=(a−(ab−μ)b,b−(ab−μ)a)"}],"headline":"Finite-step GD on bilinear attention factors yields Chebyshev ellipse and cubic map; no overlap with RS J-cost or distinction forcing","alignment":"orthogonal","rationale":"The paper reduces one-prompt linear self-attention GD to the explicit product map Φ_μ(a,b) whose balanced slice is the cubic F_μ and whose full dynamics admit the invariant ellipse E_μ carrying the Chebyshev iteration e ↦ e³−3e. These are standard discrete dynamical objects arising from quadratic loss ℓ_μ=½(ab−μ)². RS instead forces the specific reciprocal cost J(x)=½(x+x⁻¹)−1, φ-ladders, 8-tick periodicity and parameter-free constants from a single distinction (reality_from_one_distinction, AbsoluteFloorClosure, Cost.FunctionalEquation). No structural isomorphism or shared forcing mechanism is present; the domains (stat.ML training bifurcations vs. logic-to-physics derivation) are disjoint.","tokens_in":66152,"confidence":"moderate","tokens_out":407,"duration_ms":25195,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Large constant learning rates can shift linear transformer training from in-context regression to cycles, chaos or divergence.","keywords":["linear transformers","training dynamics","large learning rates","gradient descent","chaotic dynamics","in-context learning","phase transitions","attractors"],"falsifier":"Running numerical simulations of the gradient descent updates for the linear transformer with effective step sizes mu just above and below the predicted stability thresholds and checking if the trajectories enter periodic cycles or diverge as forecasted by the cubic map.","tokens_in":2629,"feed_emoji":"🌀","tokens_out":509,"duration_ms":51316,"temperature":0.7,"pith_summary":"The paper investigates the training dynamics of a simplified one-prompt linear transformer using gradient descent with large constant learning rates. By normalizing the problem, it reduces the dynamics to a two-factor product map parameterized by an effective step size mu. Analysis of this map on balanced and full two-dimensional slices reveals transitions from convergence to periodic behavior, bounded chaos, and divergence as mu increases. This is important because it indicates that large learning rates do not simply accelerate training but can alter the final learned behavior or prevent convergence altogether in transformer models.","feed_headline":"Large learning rates alter transformer attractors to cycles and chaos","feed_subtitle":"Finite-step gradient descent with high constant step sizes can lead to periodic or chaotic behavior rather than the expected in-context  ","key_machinery":"The two-factor product map obtained after normalization of the one-prompt linear-transformer training dynamics, which allows reduction to a scalar cubic map on the balanced slice and admits an explicit invariant Chebyshev ellipse in the full 2D system.","core_discovery":"Gradient-flow analyses show that simplified linear transformers can learn the in-context linear-regression algorithm, but they do not explain the finite-step behavior of gradient descent at large learning rates. Motivated by empirical work on high-learning-rate transformer instabilities and by the cubic-map phase diagram for quadratic regression, we study an exactly reducible one-prompt linear-transformer training problem. After normalization, the dynamics reduce to a two-factor product map with an effective step-size parameter mu. On the balanced slice, this map recovers the known scalar cubic transition from monotone convergence to catapult convergence, periodic and chaotic bounded noncon ","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Large steps reshape transformer attractors into cycles and chaos","High step sizes yield periodic and chaotic training in linear transformers","Two-factor dynamics expose chaos at large constant learning rates","Beyond mu stability linear models enter bounded chaos or divergence"],"cache_read_input_tokens":64,"weakest_assumption_plain":"After normalization, the one-prompt linear-transformer training dynamics reduce exactly to a two-factor product map with effective step-size mu.","fun_headline_variants_meta":{"raw":{"variants":["Large steps reshape transformer attractors into cycles and chaos","High step sizes yield periodic and chaotic training in linear transformers","Two-factor dynamics expose chaos at large constant learning rates","Beyond mu stability linear models enter bounded chaos or divergence"]},"model":"grok-4.3","cost_usd":0.01333,"raw_usage":{"total_tokens":5795,"prompt_tokens":711,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":133299500,"prompt_tokens_details":{"text_tokens":711,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":5022,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":711,"tokens_out":62,"duration_ms":44646,"temperature":1.0,"reasoning_tokens":5022,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T03:46:19.481174+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running numerical simulations of the gradient descent updates for the linear transformer with effective step sizes mu just above and below the predicted stability thresholds and checking if the trajectories enter periodic cycles or diverge as forecasted by the cubic map.","supporting_citations":[],"review_version":1}