{"id":"a54706af-599b-48ad-bee8-b8351f10d60a","arxiv_id":"2501.16703","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"For continuously observed high-dimensional diffusions with linear drift, the adaptive Lasso is shown to be sign-consistent and asymptotically normal under explicit rate conditions and a sub-Gaussian concentration assumption.","lead":"An adaptive Lasso estimator for the drift of a continuously observed high-dimensional diffusion is proved to recover the correct sparse set of nonzero coefficients and to be asymptotically normal, under explicit rate conditions and a concentration assumption. The paper also gives a marginal pre-estimator for the p>>d regime and confirms the theory in simulations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption (C) drives every probability bound in Theorems 2.5–2.6, yet the only verification offered (Prop. 4.2) is invalid for the simulation basis, so no concrete high-dimensional model is shown to satisfy the core hypothesis.","rationale":"Good-faith reading: the paper's core contribution is conditional theory — if (B1)–(B3) and (C) hold, then the adaptive Lasso is sign-consistent and asymptotically normal. The proof structure is standard, the rates are explicit, and the KKT-based decomposition in Theorem 2.5 is coherent. The load-bearing issue is that (C) is a high-level concentration assumption on the random Fisher information, and the paper provides only one concrete route to verify it (Prop. 4.2). That route's hypotheses are already too strong to cover the paper's own simulation model, so the reader cannot tell whether any high-dimensional diffusion with p growing actually satisfies (C). This is an addressable gap: state a verified example (e.g., an OU process with a fixed dictionary) or give a theorem that derives (C) from growth/mixing conditions with explicit K. A second, separate error strengthens the conditional verdict: Section 3.2's marginal estimator (3.1) omits the φ_0 drift term from model (1.3); if φ_0 ≠ 0, E[eθ] = C∞θ0 + E[Φ^*φ_0], and the proof of Theorem 3.2 silently assumes φ_0=0. This undermines the p≫d pre-estimator claim but does not invalidate the main theorems, since they are assumption-driven. The reader's identification of (C) as the weakest assumption is correct; I add the simulation-verification failure and the φ_0 omission. These are fixable but real, so the verdict remains CONDITIONAL rather than ACCEPT or REJECT.","tokens_in":23658,"tokens_out":21957,"duration_ms":196190,"concrete_test":"Compute the Lipschitz-constant matrix Q for the simulation basis: Q_ij = sup_x ∥∇⟨φ_i,φ_j⟩(x)∥ with φ_i(x)=cos((i+1)x). Show max_{i,j} Q_ij = 2p+O(1), so the uniform boundedness condition of Prop. 4.2 is violated. Then compute (or upper-bound) ∥Q∥_op and substitute the resulting K into Assumption (B3) for the simulation's p,s,T; if the condition fails, the experiment cannot be cited as a check of the theorem. Optionally, simulate the diffusion and estimate the tail of H_v for random unit vectors v to see whether P(H_v ≥ x) is bounded by exp(−T x²/(4K)) with any K satisfying (B3); heavier tails would refute (C) for this model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorems 2.5 and 2.6 reduce to Lemmas 6.1, 6.5–6.9 and Props. 6.2–6.4, all of which invoke Assumption (C) through the sub-Gaussian bound on H_v (Eq. 2.2). The only explicit sufficient condition for (C) is Proposition 4.2, which requires the inner products ⟨φ_i,φ_j⟩ to be Lipschitz with constants uniformly bounded in p,s,d. The paper's numerical basis φ_i(x)=cos((i+1)x) (Eq. 5.1) fails this: ∂_x ⟨φ_i,φ_j⟩ has sup norm at least (i+1)+(j+1), so max_{1≤i,j≤p} Q_ij = Θ(p), unbounded as p grows. Hence Section 5's claim 'Assumption (C) is verified by Proposition 4.2' is false. Even repairing the argument via Theorem 4.1 with K = ∥Q∥_op gives K = Θ(p^2) for this basis, and the convergence condition in Assumption (B3), which contains K s log p / (τmin² T), would not hold for the simulated values p=30, T=10, s≈5–10 (K s log p/T ≈ 10^5–10^6). Thus the numerical study does not evidence the theorem's regime. No other verifiable sufficient condition is supplied for general growing-p dictionaries, so the central claim is conditional on a concentration property that is not shown to hold in a single concrete instance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies continuous-time observation of a d-dimensional ergodic diffusion whose drift is linear in a p-dimensional parameter θ0 through a dictionary Φ. Under a sparsity assumption on θ0, it analyzes the adaptive Lasso estimator of the drift parameter. The main results, Theorems 2.5 and 2.6, assert that under Assumptions (A1)-(A3), (B1)-(B3) and (C), the adaptive Lasso is sign-consistent and that the estimated active subvector is asymptotically normal with oracle covariance. The proof decomposes the failure of sign consistency into five bad events whose probabilities are bounded using the sub-Gaussian concentration Assumption (C) on the quadratic functional H_v, together with bounds on the empirical Fisher information and the martingale term. The paper also proposes a marginal pre-estimator for p >> d under a partial orthogonality condition, and reports simulations for the basis φ_i(x)=cos((i+1)x) with d=5, p=30, T=10.","tokens_in":24099,"tokens_out":14642,"duration_ms":136474,"significance":"If the main claims hold, the paper makes a useful contribution: it extends adaptive-Lasso oracle properties from static high-dimensional regression to ergodic diffusions with growing parameter dimension, gives explicit tuning-parameter relationships, and proposes a pre-estimator for genuinely high-dimensional cases. The proof is organized as a standard sequence of probability bounds derived from Assumption (C); there is no circularity, and the handling of the marginal estimator is constructive. However, the paper's only explicit sufficient condition for Assumption (C) in the growing-p regime, Proposition 4.2, is not satisfied by the paper's own simulation basis, and no other concrete model is shown to satisfy (C). Because every probability bound in Theorems 2.5 and 2.6 passes through Assumption (C), this is a load-bearing gap that must be addressed before the numerical claims and the scope of the theorems can be accepted.","major_comments":[{"comment":"The claim in Section 5 that \"Assumption (C) is verified by Proposition 4.2\" is not correct for the simulation basis. For φ_i(x)=cos((i+1)x), the function g_{ij}(x)=⟨φ_i,φ_j⟩ has Lipschitz constant d(i+j+2): already in dimension d=1, the derivative of cos((i+1)x)cos((j+1)x) has sup norm i+j+2. Hence the Lipschitz constants are not uniformly bounded in p, and Proposition 4.2, which explicitly requires uniformly bounded constants, cannot be invoked. If one instead tries to use Theorem 4.1 with K=∥Q∥_op, then for Q_{ij}=i+j+2 one has ∥Q∥_op=Θ(p^2); for the simulated values p=30, T=10 and s≈5-10, the quantity K s log p/(τ_min^2 T) appearing in Assumption (B3) is of order 10^3 and does not vanish. Thus the numerical study does not provide evidence for the asymptotic regime of Theorems 2.5-2.6.","section":"Section 5, Eq. (5.1), and Proposition 4.2"},{"comment":"Every probability bound used in the proofs of Theorems 2.5 and 2.6, including Lemmas 6.1, 6.5-6.9 and Propositions 6.2-6.4, relies on Assumption (C). The only sufficient condition supplied for (C) is Proposition 4.2, and Major Comment 1 shows that this condition fails for the paper's own example. As the manuscript stands, no concrete high-dimensional dictionary is proved to satisfy (C), so the main theorems are conditional on a concentration property that is neither verified nor instantiated. The revision should either prove (C) for a nontrivial class of dictionaries with explicit control of K, or state the results purely under Assumption (C) and remove the verification sentence in Section 5 unless a valid example is supplied.","section":"Section 2, Assumption (C), and Lemmas 6.1, 6.4-6.9"},{"comment":"For nonlinear dictionaries the statement that l_min(C_∞)>0 implies p≤d is not correct. Since C_∞=E[Φ(X_0)^TΦ(X_0)], its rank can exceed d when the dictionary functions are nonlinear; for instance, with d=1, φ_1(x)=x, φ_2(x)=x^2 and X_0 uniform on ±1, one obtains C_∞=I_2. This does not affect the main theorems, but the motivation for the marginal estimator in Section 3.2 should be reworded.","section":"Section 3.1"}],"minor_comments":[{"comment":"The symbol d is reused for both the dimension of the diffusion and the exponent in Assumption (C'), which is confusing; an exponent like q or γ would be clearer.","section":"Remark 4.3"},{"comment":"There is a small typo in the statement of Assumption (C): \"there exists a constant K such that that for all μ∈R\" contains a duplicated \"that\".","section":"Assumption (C)"},{"comment":"In the simulation model b_{θ0}(x)=3s x + Σ θ_i cos((i+1)x), the scalar s is used for a model constant even though s denotes the sparsity level elsewhere in the paper; this is potentially misleading and should be renamed.","section":"Eq. (5.1)"},{"comment":"References [7] and [8] are identical, and the bibliography would benefit from a consistency check for other duplicates or missing page ranges.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The main obstacle is the verification of Assumption (C). The theorems themselves may well be correct under (C), and the proof strategy is standard, but the paper's own numerical section claims a verification that is invalid, and no other example is provided. I would ask the authors to supply a verifiable sufficient condition for (C) with a concrete model, or to remove the verification claim and add an example where (C) is checked. This is fixable within the scope of the paper, so I do not recommend rejection at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious paper that extends the adaptive Lasso oracle program to continuous-time ergodic diffusions with growing p, and the main theorems are probably right. But the load-bearing Assumption (C) is never verified for a single concrete high-dimensional dictionary, and the one verification the paper claims—in the numerical section—is wrong for the basis they actually simulate. That combination means the paper needs real revision before it is a citable result.\n\nWhat is genuinely new: they give the first sign-consistency and asymptotic-normality conditions for the adaptive Lasso in high-dimensional diffusions, and they handle p >> d via a marginal pre-estimator under partial orthogonality. The proof strategy is standard—KKT conditions plus concentration bounds—and they are careful to track how L, M, K, and τ_min scale. The explicit discussion of why the Lasso pre-estimator requires p ≤ d, and the partial-orthogonality alternative, is a real contribution. No circularity: the reliance on their earlier Lasso bound is only as an example pre-estimator, not in the central proof.\n\nThe soft spots are real. Assumption (C) is a black-box sub-Gaussian bound on H_v, and the only sufficient condition they offer (Prop 4.2) requires uniformly bounded Lipschitz constants for the inner products ⟨φ_i, φ_j⟩. Their simulation basis φ_i(x)=cos((i+1)x) fails that: the Lipschitz constant of the product grows like i+j. So the Section 5 sentence \"Assumption (C) is verified by Proposition 4.2\" is false, and the numerical study does not actually exercise the theorem's regime. That does not falsify the theorems, but it leaves the central assumption without a single concrete instance where it is shown to hold for growing p. The paper should either prove (C) for a nontrivial dictionary or state clearly that it is an unverified condition.\n\nWorth fixing too: the proof of Theorem 2.5 displays a wrong necessary condition in (2.5)—the first line should require the estimator has the right sign, not sign(θ0)(θ̂ − θ0) ≤ |θ0|. And the paper should be explicit that A3, positive definiteness of I_S, forces s ≤ d; the \"p >> d\" claim is about zero coordinates only. Both are addressable.\n\nWho is this for? People working on sparse inference for diffusions, and anyone designing pre-estimators for high-dimensional stochastic models. It deserves a serious referee—the contribution is real and the gaps are fixable—but not acceptance as-is. I would send it out and ask for a revision that either verifies (C) in a concrete growing-p example or honestly relegates it to a conjecture, plus the small proof fixes.","headline":"Solid conditional extension of the adaptive Lasso to high-dimensional diffusions, but the core concentration assumption is unverified in any concrete growing-p example and the paper's own numerical verification is wrong.","tokens_in":24516,"tokens_out":5699,"would_cite":false,"duration_ms":51594,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M05","60G15","62H12"],"pacs":[],"model":"deepseek-v4-flash","headline":"The adaptive Lasso recovers the exact support of the drift in high-dimensional ergodic diffusions and is asymptotically normal on the selected coefficients.","keywords":["Adaptive Lasso","support recovery","concentration inequality","diffusion models","high-dimensional statistics","ergodic diffusion","asymptotic normality","drift estimation"],"falsifier":"Take a specific ergodic diffusion with a high-frequency basis such as $\\phi_i(x)=\\cos((i+1)x)$, estimate the bounding constant $K$ in the assumed sub-Gaussian concentration inequality from long simulated trajectories, and check whether $K$ stays bounded as $p$ and the basis index grow; if $K$ grows, the dimension-free concentration that the theorems need is not satisfied, and one can test whether the adaptive Lasso's support-recovery probability still tends to 1 in simulations.","tokens_in":23443,"feed_emoji":"🎯","tokens_out":12713,"duration_ms":107445,"temperature":0.7,"pith_summary":"The paper establishes that the adaptive Lasso, applied to the drift of a high-dimensional ergodic diffusion, can do two things at once: recover the exact set of nonzero drift coefficients with probability tending to one, and estimate those nonzero coefficients at the oracle rate with a Gaussian limit. Such a simultaneous support-recovery and oracle-normality guarantee had been missing for diffusion models with a growing number of parameters. The practical payoff is that in sparse high-dimensional models of processes arising in finance, biology, and epidemiology, one estimator and one tuning parameter can deliver both variable selection and valid inference, something the standard Lasso cannot do.","feed_headline":"Adaptive Lasso nails exact drift support in high-dimensional diffusions","feed_subtitle":"With sparsity and one tuning parameter, it selects the true nonzero coefficients and estimates them at oracle precision","key_machinery":"The machinery is the Karush-Kuhn-Tucker characterization of the adaptive Lasso solution, which turns sign consistency into control of five bad events $B_1$ through $B_5$: the inverse active Fisher information $C^{SS}_T$ acting on the martingale noise, the weighted sign vector, the cross-correlation between active and inactive blocks, and the martingale noise itself. Two conditions carry the argument: the adaptive irrepresentable condition (B2), which bounds $C^{S^C S}_\\infty (C^{SS}_\\infty)^{-1}u$ in terms of the pre-estimator weights and prevents inactive coefficients from entering the model, and Assumption (C), a sub-Gaussian concentration inequality for the quadratic functional $H_v$, which ensures the empirical Fisher information matrix stays close to its expectation with constants independent of dimension.","core_discovery":"The central claim is that for a linear drift model $b_\\theta = \\phi_0 + \\sum_{j=1}^p \\theta^j_0 \\phi_j$ with $s$ nonzero coefficients, the adaptive Lasso estimator $\\hat\\theta$ defined by minimizing $L_T(\\theta) + \\lambda \\sum_j w_j |\\theta_j|$ satisfies $\\mathrm{P}(\\hat\\theta =_s \\theta_0) \\to 1$ as $T \\to \\infty$ under assumptions (A1)-(A3), (B1)-(B3), and (C) (Theorem 2.5). Moreover, for any direction $\\alpha \\in \\mathbb{R}^s$, $\\sqrt{T}\\, s_T^{-1} \\alpha^\\top (\\hat\\theta_S - \\theta_{0,S})$ converges in distribution to a standard normal (Theorem 2.6), with the normalization $s_T^2 = \\alpha^\\top (C^{SS}_\\infty)^{-1}\\alpha$. The paper frames this as the adaptive Lasso combining exact variable selection with oracle-style efficiency for the active drift coefficients, which the standard Lasso cannot achieve with a single tuning parameter.","pith_inferences":["The proof is modular in the concentration assumption: any verification of (C), for instance through exponential bounds derived from the process's mixing or ergodicity, plugs into the same theorems, so the scope of the results is set by how widely (C) can be checked.","For the numerical basis $\\phi_i(x)=\\cos((i+1)x)$, the pairwise inner products have Lipschitz constants growing with $i$, so Proposition 4.2's sufficient condition for (C) is not literally satisfied in the simulations; verifying (C) directly for that basis is a concrete open check.","The sub-exponential relaxation (C') should preserve support recovery but shrink the allowed sparsity $s$; spelling out the exact rate trade-off would quantify how much generality costs in statistical power.","Theorem 2.6 makes it feasible to build confidence intervals for individual drift coefficients in high-dimensional diffusions, a practical procedure the paper illustrates but does not fully develop."],"forward_implications":["Exact support recovery: the estimated nonzero set equals the true nonzero set with probability approaching one, so downstream analysis can condition on the selected support.","Oracle inference on active coefficients: $\\sqrt{T}\\, s_T^{-1} \\alpha^\\top(\\hat\\theta_S - \\theta_{0,S})$ is asymptotically standard normal, enabling confidence intervals and tests for the drift coefficients.","Feasible regimes: $p$ can grow exponentially in the observation horizon, like $\\exp((dT)^\\delta)$, while the sparsity $s$ grows only polynomially, making the theory relevant for $p\\gg d$ problems.","Pre-estimator options: with a positive-definite expected Fisher information, the Lasso itself works as the pre-estimator; under a partial-orthogonality condition, a cheap marginal estimator covers the $p\\gg d$ case.","One tuning parameter suffices: the same $\\lambda$ can satisfy both the support-recovery condition (B3) and the normality condition (2.8), overcoming the standard Lasso's inability to do both."],"supporting_citations":[{"why":"Supplies the adaptive Lasso estimator and the oracle-property framework that the paper adapts to diffusion models.","marker":"[45]"},{"why":"Provides the Lasso drift estimator used as the pre-estimator and its convergence rate $r_T$ for the support-recovery proof.","marker":"[13]"},{"why":"Gives the multivariate martingale central limit theorem that yields the asymptotic normality in Theorem 2.6.","marker":"[31]"},{"why":"Provides the concentration inequality for additive functionals that Theorem 4.1 uses to verify Assumption (C).","marker":"[40]"},{"why":"Supplies concentration bounds for multivariate elliptic diffusions, an alternative route to controlling the random Fisher information matrix.","marker":"[43]"},{"why":"Establishes high-dimensional Lasso estimation under discrete sampling for diffusion drift, motivating the numerical setup and the pre-estimator discussion.","marker":"[4]"}],"fun_headline_variants":["Adaptive Lasso recovers true drift support with oracle efficiency","Exact support recovery for diffusions via adaptive Lasso","Adaptive Lasso: exact selection + normal limits in high-dim diffusions","One tuning parameter, oracle accuracy: adaptive Lasso for diffusions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the random fluctuations of the empirical information matrix concentrate around their averages with sub-Gaussian tails and a constant that stays bounded as the dimension grows; if that concentration fails, the probability bounds behind exact support recovery and asymptotic normality collapse.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive Lasso recovers true drift support with oracle efficiency","Exact support recovery for diffusions via adaptive Lasso","Adaptive Lasso: exact selection + normal limits in high-dim diffusions","One tuning parameter, oracle accuracy: adaptive Lasso for diffusions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000245,"raw_usage":{"total_tokens":1516,"prompt_tokens":907,"completion_tokens":609,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":535}},"tokens_in":523,"tokens_out":609,"duration_ms":5506,"temperature":1.0,"reasoning_tokens":535,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T11:18:11.953180+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a specific ergodic diffusion with a high-frequency basis such as $\\phi_i(x)=\\cos((i+1)x)$, estimate the bounding constant $K$ in the assumed sub-Gaussian concentration inequality from long simulated trajectories, and check whether $K$ stays bounded as $p$ and the basis index grow; if $K$ grows, the dimension-free concentration that the theorems need is not satisfied, and one can test whether the adaptive Lasso's support-recovery probability still tends to 1 in simulations.","supporting_citations":[{"cited_title":"The adaptive lasso and its oracle properties","cited_arxiv_id":null,"evidence_quote":"Supplies the adaptive Lasso estimator and the oracle-property framework that the paper adapts to diffusion models."},{"cited_title":"On Lasso estimator for the drift function in diffusion models","cited_arxiv_id":null,"evidence_quote":"Provides the Lasso drift estimator used as the pre-estimator and its convergence rate $r_T$ for the support-recovery proof."},{"cited_title":"A note on limit theorems for multivariate martingales","cited_arxiv_id":null,"evidence_quote":"Gives the multivariate martingale central limit theorem that yields the asymptotic normality in Theorem 2.6."},{"cited_title":"Transportation inequalities for stochastic differential equations driven by a fractional Brownian motion","cited_arxiv_id":null,"evidence_quote":"Provides the concentration inequality for additive functionals that Theorem 4.1 uses to verify Assumption (C)."},{"cited_title":"Concentration anal- ysis of multivariate elliptic diffusions","cited_arxiv_id":null,"evidence_quote":"Supplies concentration bounds for multivariate elliptic diffusions, an alternative route to controlling the random Fisher information matrix."},{"cited_title":"Sampling effects on lasso esti- mation of drift functions in high-dimensional diffusion processes","cited_arxiv_id":null,"evidence_quote":"Establishes high-dimensional Lasso estimation under discrete sampling for diffusion drift, motivating the numerical setup and the pre-estimator discussion."}],"review_version":1}