{"id":"5a10fbc9-5674-4391-95ef-71dbfe5e8caf","arxiv_id":"2412.17070","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Rescaled fast and slow iterates of two-time-scale stochastic approximation converge weakly to decoupled Ornstein-Uhlenbeck processes, with coupling entering only through the coefficient matrices.","lead":"A new theorem proves functional central limit theorems for two-time-scale stochastic approximation, showing that the fast and slow iterates, rescaled by their own step sizes, converge weakly to separate Ornstein-Uhlenbeck processes. This gives a sharper trajectory-level asymptotic description than earlier single-point or single-scaling results.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 4.1 imports non-asymptotic rate bounds whose formal proof needs strong monotonicity; the theorem's own assumptions do not guarantee them.","rationale":"The reader's weakest assumption identifies exactly the load-bearing issue: the main theorem depends on Assumption 4.1, which is imported from prior work and whose formal justification requires stronger conditions than those stated in the theorem. The paper is transparent about this in Appendix B.1, where Proposition B.1 adds strong monotonicity and an extra step-size restriction, but those conditions are not part of Theorem 4.1's hypothesis list. Because Assumption 4.1 underpins all residual bounds, tightness, and the martingale approximation, a gap here directly affects the central claim. However, this is a conditional gap rather than an error: Theorem 4.1 is valid under the stated assumptions including 4.1, and the paper does not claim to prove 4.1 from 3.1-3.6. No internal inconsistency or mathematical error was found in the proof framework itself. Hence the reader's CONDITIONAL verdict is appropriate and should remain unchanged.","tokens_in":43499,"tokens_out":24667,"duration_ms":216747,"concrete_test":"Analytical check: revisit the proof of [23, Theorem 3.1] and locate every use of the strong monotonicity constants mu_F and mu_G. For each use, determine whether Assumption 3.5 (local Hurwitz) plus the global Lipschitz bounds in Assumption 3.1 can supply the same bound after a finite burn-in. If any essential use cannot be replaced, then Theorem 4.1's hypothesis set is incomplete: it must either include the strong monotonicity conditions of Proposition B.1 or prove Assumption 4.1 from the existing assumptions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every residual control in Lemmas 5.1 and 5.2, and therefore the tightness in Lemma 5.3 and the martingale approximation in Lemma 5.4, is built on Assumption 4.1: E||xhat_n||^2 = O(alpha_n), E||yhat_n||^2 = O(beta_n), and E||xhat_n||^4 + E||yhat_n||^4 = O(alpha_n^2). The paper never derives these bounds from Assumptions 3.1-3.6. The only formal statement, Proposition B.1, adds global strong monotonicity of F and of G(H(y),y), plus the step-size restriction b/a <= 1 + delta_F/2, none of which appear in Theorem 4.1. Assumption 3.5 is a local Hurwitz condition; it controls behavior near the root but not the transient from an arbitrary initialization, so it cannot by itself justify the uniform-in-n second-moment bounds that make the rescaled initial values tight. If Assumption 4.1 fails for some system satisfying 3.1-3.6, the residual terms R^x_n, R^y_n, R^z_n do not vanish at the required orders and the limiting SDEs in (14)-(16) need not follow. This is not an internal contradiction because Theorem 4.1 explicitly assumes 4.1, but it means the central claim is conditional on a black-box rate that is not a consequence of the listed hypotheses.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper establishes decoupled functional central limit theorems for two-time-scale stochastic approximation. The authors construct continuous-time processes by rescaling the fast error x_n - H(y_n) by sqrt(alpha_n) and the slow error y_n - y* by sqrt(beta_n), and prove, under Assumptions 3.1-3.6 together with an imported non-asymptotic rate assumption (Assumption 4.1), that these processes converge weakly to stationary Ornstein-Uhlenbeck processes with drift and diffusion coefficients given in (14)-(17). The proof proceeds through one-step recursions, an auxiliary sequence that removes the dominant fast-scale influence on the slow scale, tightness of the trajectory sequences, and the martingale problem approach.","tokens_in":1452,"tokens_out":2108,"duration_ms":97572,"significance":"If the result is valid, it is a meaningful advance over earlier single-point decoupled convergence results and over the prior FCLT of Faizal and Borkar [18], because it captures a stochastic limit on each time scale and identifies the coupling only through the coefficients of the limiting SDEs. The paper contains detailed appendix proofs of the recursions, tightness, and approximate martingale problem, and the limiting SDEs are explicit, falsifiable predictions with no fitted parameters. The main reservation is that the central theorem is conditional on Assumption 4.1, whose formal verification in Appendix B.1 requires strong monotonicity conditions and step-size restrictions that do not appear in the main hypotheses of Theorem 4.1.","major_comments":[{"comment":"The central claim of Theorem 4.1 is conditional on Assumption 4.1, which imports the rate bounds E||xhat_n||^2 = O(alpha_n), E||yhat_n||^2 = O(beta_n), and E||xhat_n||^4 + E||yhat_n||^4 = O(alpha_n^2) from [23, Theorem 3.1]. These bounds are never derived from Assumptions 3.1-3.6; the only formal statement, Proposition B.1, requires global strong monotonicity of F and of G(H(y),y), and it obtains the second-moment rates only under the additional step-size restriction b/a <= 1 + delta_F/2. Neither the strong monotonicity nor the step-size restriction appears in Theorem 4.1. Since every residual control in Lemmas 5.1 and 5.2, and therefore the tightness Lemma 5.3 and the approximate martingale problem Lemma 5.4, is built on Assumption 4.1, a system satisfying Assumptions 3.1-3.6 but not the stronger rate conditions could have residual terms R^x_n, R^y_n, R^z_n that do not vanish at the required orders, so the limiting SDEs in (14)-(16) need not follow. Please either prove Assumption 4.1 under the main assumptions or add the strong monotonicity and step-size conditions to the hypotheses of Theorem 4.1 and discuss the impact on Examples 4.1-4.3.","section":"Section 4.1 and Appendix B.1"},{"comment":"The residual properties stated in Lemmas 5.1 and 5.2 are used for p in the interval (2, 4/(1+(delta_H vee delta_F vee delta_G)/2)], but the moment computations in the proof of Lemma 5.1 in Appendix C.1 are only sketched. In particular, the bounds for E||R^y_n||^p and E||R^x_n||^p require a careful interpolation between the second-moment and fourth-moment rates of Assumption 4.1 under the step-size condition (v) of Assumption 3.4; the text says this follows from Young and Jensen inequalities but does not display the full argument. Please provide the complete interpolation step or state the needed moment hypotheses explicitly, so that the residual bounds are fully verifiable.","section":"Section 5.1"}],"minor_comments":[{"comment":"The phrase 'depend sorely on their respective step sizes' should read 'depend solely on their respective step sizes'.","section":"Section 1, Related Work"},{"comment":"In the displayed decomposition for the slow scale, the generator term is written as A_y f(bar{Z}_n(s)); since the test function is g, this should be A_y g(bar{Z}_n(s)).","section":"Lemma 5.4(ii)"},{"comment":"In Step 1 and Step 2 of the proof, the text refers to 'the SDE in (16)' when the generator A_x corresponds to equation (14); the same substitution is needed wherever the invariant distribution and semigroup of the fast process are discussed. Also, near the end of Step 2, 'Letting epsilon -> infinity' should be 'Letting epsilon -> 0'.","section":"Section C.4, Proof of Theorem 4.1(i)"},{"comment":"The diffusion matrix of the slow SDE appears in inconsistent notation: tilde Sigma psi, ~Sigma psi, and Sigma psi are all used. Please unify the notation.","section":"Equations (16)-(17)"},{"comment":"Proposition 4.1 is labelled 'informal' and Assumption 4.1 refers back to it; to help the reader, point directly to the formal Proposition B.1 in the main body when introducing Assumption 4.1.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The proof strategy is substantial and appears sound conditional on the imported rate bounds. The main issue is that Assumption 4.1 is a black-box condition whose formal verification requires stronger hypotheses than those stated in Theorem 4.1; this is a fixable gap if the authors either prove the rates under the stated assumptions or restate the theorem with the stronger conditions. I do not see an internal contradiction, and the novelty claim appears reasonable, though I have not independently audited the entire related literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does what it says: it gives the first decoupled functional CLT for general two-time-scale SA, with the fast and slow iterates each rescaled by the square root of their own step size. The limit is the expected OU process from standard SA, and the coupling between scales enters only through the coefficients. The auxiliary sequence ˇzn is a genuine new analytical trick: it subtracts the leading cross-term from the slow iterate and converts an intractable recursion into a standard SA form. The tightness proof for ¯Yn is the hard part, and the authors handle it honestly, expanding the troublesome cross term over the fast scale rather than hiding it.\n\nWhere the paper is soft is exactly where the stress-test note puts it. The main theorem is conditional on Assumption 4.1, which imports the non-asymptotic rate bounds E‖x̂n‖² = O(αn), E‖ŷn‖² = O(βn), and the fourth-moment bound. These rates are not derived from Assumptions 3.1–3.6 in this paper. The formal statement in Appendix B.1 (Proposition B.1) requires global strong monotonicity of F and of G(H(y),y), plus the step-size restriction b/a ≤ 1 + δF/2, none of which appear in Theorem 4.1. The authors are transparent that they are assuming the rates as given, so this is not a hidden circularity or an internal contradiction. But it does mean the headline theorem is a reduction to a black box: the decoupled FCLT is proven for systems satisfying the stronger monotonicity conditions, not for all systems satisfying Assumptions 3.1–3.6. This is fixable by either moving the strong monotonicity into the main theorem or explicitly expanding Assumption 4.1 to match Proposition B.1.\n\nMinor issues: the GTD2/TDC example is mostly checked by inspection, though the covariance computation is in the appendix; the geometric ergodicity step uses a cited result from Mattingly–Stuart–Higham, fine. There are no fitted parameters and the proofs are detailed enough to be checked. The comparison with the earlier FCLT of Faizal–Borkar is accurate: their common scaling misses slow-scale randomness, and this paper fixes that.\n\nFor a reader working on asymptotic theory of two-time-scale SA or RL algorithms, this is a solid, publishable contribution. It deserves serious peer review, with the recommendation that the authors align the assumptions across Theorem 4.1 and Proposition B.1.","headline":"A credible, genuinely new decoupled FCLT for two-time-scale SA, with a smart auxiliary-sequence device, but the main theorem leans on non-asymptotic rate bounds imported from the authors' earlier paper under stronger assumptions than those stated.","tokens_in":44316,"tokens_out":2532,"would_cite":true,"duration_ms":25529,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60F17","60F05","62L20","60J60"],"pacs":[],"model":"deepseek-v4-flash","headline":"Decoupled functional CLTs give two Ornstein-Uhlenbeck limits for two-time-scale stochastic approximation.","keywords":["two-time-scale stochastic approximation","functional central limit theorem","Ornstein-Uhlenbeck process","decoupled convergence","martingale problem","weak convergence","tightness","stochastic approximation"],"falsifier":"Run two-time-scale SA on a linear problem with known $B_1,B_2,B_3$ and known noise covariances, using step sizes $\\alpha_n=(n+1)^{-a}$ and $\\beta_n=(n+1)^{-b}$ with $b/a<1$, so Assumption 3.4(v) fails. If the rescaled slow trajectory still converges weakly to the stationary OU solution of (16), the condition is not necessary; if convergence fails or the limit changes, it is load-bearing. A second check is to estimate $E\\|\\hat x_n\\|^4$ directly: if it is not $O(\\alpha_n^2)$, Assumption 4.1 fails and the residual terms $R_n^y$ and $R_n^z$ in Lemmas 5.1–5.2 no longer vanish at the required orders.","tokens_in":2139,"feed_emoji":"📈","tokens_out":5399,"duration_ms":101602,"temperature":0.7,"pith_summary":"Two-time-scale stochastic approximation updates a fast and a slow iterate with different step sizes, each influencing the other. This paper proves that, after rescaling each error term by the square root of its own step size, the fast trajectory converges weakly to a stationary Ornstein-Uhlenbeck process with drift $-B_1$ and diffusion $\\Sigma_\\xi$, and the slow trajectory converges weakly to a stationary Ornstein-Uhlenbeck process with drift $-(B_3-\\tilde{\\beta}I/2)$ and diffusion $\\tilde{\\Sigma}_\\psi$. The limiting dynamics on each time scale has the same form as standard stochastic approximation run with that step size alone; the only trace of the coupling between the two scales is in the coefficients of the slow limit. A reader should care because trajectory-level limits, not just single-point central limit theorems, describe how two-time-scale algorithms fluctuate over time and support inference for methods such as actor-critic and gradient temporal-difference learning.","feed_headline":"Two-time-scale stochastic approximation splits into two OU limits","feed_subtitle":"Both rescaled trajectories converge to standard SA limits; coupling appears only in coefficients.","key_machinery":"The central object is an auxiliary sequence $\\check z_n=\\check y_n-\\sqrt{\\kappa_{n-1}}B_2B_1^{-1}\\check x_n$, where $\\check x_n=(x_n-H(y_n))/\\sqrt{\\alpha_{n-1}}$, $\\check y_n=(y_n-y_\\star)/\\sqrt{\\beta_{n-1}}$, and $\\kappa_n=\\beta_n/\\alpha_n$. Subtracting this correction removes the dominant $O(\\sqrt{\\beta_n\\alpha_n})$ influence of the fast iterate on the slow recursion, so $\\check z_n$ evolves like a standard stochastic approximation iterate. The proof builds one-step recursions for $\\check x_n$, $\\check y_n$, and $\\check z_n$ with explicit residual bounds, establishes tightness of the continuous-time trajectories formed by linear interpolation, and then uses the martingale problem approach to identify the unique limiting Ornstein-Uhlenbeck processes.","core_discovery":"Theorem 4.1 shows that, under Assumptions 3.1–3.6 and 4.1, the linearly interpolated rescaled fast trajectory $\\bar X_n(\\cdot)$ converges weakly to the stationary solution of $dX(t)=-B_1X(t)\\,dt+\\Sigma_\\xi^{1/2}\\,dW_{d_x}(t)$, and the rescaled slow trajectory $\\bar Y_n(\\cdot)$ converges weakly to the stationary solution of $dY(t)=-(B_3-\\tilde{\\beta}I/2)Y(t)\\,dt+\\tilde{\\Sigma}_\\psi^{1/2}\\,dW_{d_y}(t)$, where $\\tilde{\\Sigma}_\\psi=\\Sigma_\\psi-B_2B_1^{-1}\\Sigma_{\\xi,\\psi}-\\Sigma_{\\xi,\\psi}^{\\top}B_1^{-\\top}B_2^{\\top}+B_2B_1^{-1}\\Sigma_\\xi B_1^{-\\top}B_2^{\\top}$. Equivalently, each rescaled iterate converges in distribution to the invariant Gaussian distribution $N(0,\\Sigma_x)$ or $N(0,\\Sigma_y)$ solving the Lyapunov equations (15) and (18). This is the first decoupled functional central limit theorem for two-time-scale stochastic approximation in which each time scale is rescaled by its own step size, rather than by a common factor.","pith_inferences":["A next step the paper does not take is to use the auxiliary sequence $\\check z_n$ as a debiased slow iterate for constructing confidence intervals or tests in two-time-scale algorithms; the FCLT for $\\check z_n$ makes its asymptotic distribution explicit.","If Assumption 4.1's non-asymptotic rates are established under Markovian or state-dependent noise, the same tightness-and-martingale skeleton should yield decoupled FCLTs there; the residual bounds in Lemmas 5.1 and 5.2 are the main place the noise model enters.","The coefficient-only coupling suggests that in actor-critic and TD-style algorithms, the fast auxiliary variable and the slow parameter have asymptotically independent Gaussian trajectory fluctuations on their own timescales, which could be checked empirically by comparing simulated rescaled paths to the predicted OU limits."],"forward_implications":["Corollary 4.1 recovers the classical single-point central limit theorems: $\\alpha_n^{-1/2}(x_n-H(y_n)) \\Rightarrow N(0,\\Sigma_x)$ and $\\beta_n^{-1/2}(y_n-y_\\star) \\Rightarrow N(0,\\Sigma_y)$.","For algorithms such as SGD with Polyak-Ruppert averaging, normalized stochastic heavy ball, GTD2, and TDC, the trajectory-level limit is the stationary Ornstein-Uhlenbeck process (14) or (16), so asymptotic fluctuations over finite time horizons are characterized, not just the marginal distributions.","The fast-scale limit is independent of the slow scale, while the slow-scale limit depends on the fast scale only through the coefficient matrices $B_1,B_2$ and the noise cross-covariance, plus the step-size constant $\\tilde{\\beta}$ when $\\beta_n$ decays like $1/n$.","When $\\beta_n \\asymp 1/n$, the initial slow step size enters the limiting slow drift through $\\tilde{\\beta}=\\beta_0^{-1}$, so the asymptotic slow dynamics depends on the step-size schedule in a concrete, testable way."],"supporting_citations":[{"why":"Supplies Assumption 4.1, the non-asymptotic decoupled rates $E\\|\\hat x_n\\|^2=O(\\alpha_n)$, $E\\|\\hat y_n\\|^2=O(\\beta_n)$, and $E\\|\\hat x_n\\|^4+E\\|\\hat y_n\\|^4=O(\\alpha_n^2)$ on which all residual controls rest.","marker":"[23]"},{"why":"Gives the linear-case central limit theorem and convergence-rate framework that the decoupled FCLT extends to trajectory level.","marker":"[33]"},{"why":"Establishes the nonlinear single-point joint CLT for $\\alpha_n^{-1/2}(x_n-x_\\star)$ and $\\beta_n^{-1/2}(y_n-y_\\star)$, which Corollary 4.1 recovers.","marker":"[49]"},{"why":"Provides the only prior FCLT for two-time-scale SA, whose common-step-size rescaling this paper refines into decoupled rescalings.","marker":"[18]"},{"why":"Supplies the martingale problem characterization and tightness machinery used to prove weak convergence.","marker":"[17]"},{"why":"Gives finite-time convergence rates for linear two-time-scale SA with Markovian noise, contextualizing the decoupled rates being lifted to FCLTs.","marker":"[27]"},{"why":"Provides the classical Polyak-Ruppert averaging CLT recovered in Example 4.1, anchoring the slow-scale limit.","marker":"[55]"}],"fun_headline_variants":["First decoupled functional CLT for two-time-scale SA","Two-time-scale SA decouples into standard SA limits","Each rescaled trajectory converges to an OU process","Decoupled CLT: two-time-scale SA gets per-scale OU limits","First CLT showing two-time-scale SA splits into OU limits"],"cache_read_input_tokens":46464,"weakest_assumption_plain":"The paper assumes as given the non-asymptotic rates $E\\|\\hat x_n\\|^2=O(\\alpha_n)$, $E\\|\\hat y_n\\|^2=O(\\beta_n)$, and $E\\|\\hat x_n\\|^4+E\\|\\hat y_n\\|^4=O(\\alpha_n^2)$ from an earlier paper; it does not prove these rates from Assumptions 3.1–3.6, and every residual bound collapses if they fail.","fun_headline_variants_meta":{"raw":{"variants":["First decoupled functional CLT for two-time-scale SA","Two-time-scale SA decouples into standard SA limits","Each rescaled trajectory converges to an OU process","Decoupled CLT: two-time-scale SA gets per-scale OU limits","First CLT showing two-time-scale SA splits into OU limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000581,"raw_usage":{"total_tokens":2763,"prompt_tokens":1001,"completion_tokens":1762,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":1678}},"tokens_in":617,"tokens_out":1762,"duration_ms":13128,"temperature":1.0,"reasoning_tokens":1678,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:48:33.868944+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run two-time-scale SA on a linear problem with known $B_1,B_2,B_3$ and known noise covariances, using step sizes $\\alpha_n=(n+1)^{-a}$ and $\\beta_n=(n+1)^{-b}$ with $b/a<1$, so Assumption 3.4(v) fails. If the rescaled slow trajectory still converges weakly to the stationary OU solution of (16), the condition is not necessary; if convergence fails or the limit changes, it is load-bearing. A second check is to estimate $E\\|\\hat x_n\\|^4$ directly: if it is not $O(\\alpha_n^2)$, Assumption 4.1 fails and the residual terms $R_n^y$ and $R_n^z$ in Lemmas 5.1–5.2 no longer vanish at the required orders.","supporting_citations":[{"cited_title":"Convergence rate of linear two-time-scale stochastic approximation","cited_arxiv_id":null,"evidence_quote":"Gives the linear-case central limit theorem and convergence-rate framework that the decoupled FCLT extends to trajectory level."},{"cited_title":"Convergen ce rate and averaging of nonlinear two-time-scale stochastic approximation algorithms","cited_arxiv_id":null,"evidence_quote":"Establishes the nonlinear single-point joint CLT for $\\alpha_n^{-1/2}(x_n-x_\\star)$ and $\\beta_n^{-1/2}(y_n-y_\\star)$, which Corollary 4.1 recovers."},{"cited_title":"Finite time analysis of linear two-timescale stochastic approximatio n with Markovian noise","cited_arxiv_id":null,"evidence_quote":"Gives finite-time convergence rates for linear two-time-scale SA with Markovian noise, contextualizing the decoupled rates being lifted to FCLTs."},{"cited_title":"Acceleration of st ochastic approximation by averaging","cited_arxiv_id":null,"evidence_quote":"Provides the classical Polyak-Ruppert averaging CLT recovered in Example 4.1, anchoring the slow-scale limit."}],"review_version":1}