REVIEW 3 major objections 4 minor 44 references
Consistent support recovery for high-dimensional diffusions
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The adaptive Lasso recovers the exact support of the drift in high-dimensional ergodic diffusions and is asymptotically normal on the selected coefficients.
desk verdict Solid conditional extension of the adaptive Lasso to high-dimensional diffusions, but the core concentration assumption is unverified in any concrete growing-p example and the paper's own numerical verification is wrong. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the Karush-Kuhn-Tucker characterization of the adaptive Lasso solution, which turns sign consistency into control of five bad events $B_1$ through $B_5$: the inverse active Fisher information $C^{SS}_T$ acting on the martingale noise, the weighted sign vector, the cross-correlation between active and inactive blocks, and the martingale noise itself. Two conditions carry the argument: the adaptive irrepresentable condition (B2), which bounds $C^{S^C S}_\infty (C^{SS}_\infty)^{-1}u$ in terms of the pre-estimator weights and prevents inactive coefficients from entering the model, and Assumption (C), a sub-Gaussian concentration inequality for the quadratic functional $H_v$, which ensures the empirical Fisher information matrix stays close to its expectation with constants independent of dimension.
What would settle it
Take a specific ergodic diffusion with a high-frequency basis such as $\phi_i(x)=\cos((i+1)x)$, estimate the bounding constant $K$ in the assumed sub-Gaussian concentration inequality from long simulated trajectories, and check whether $K$ stays bounded as $p$ and the basis index grow; if $K$ grows, the dimension-free concentration that the theorems need is not satisfied, and one can test whether the adaptive Lasso's support-recovery probability still tends to 1 in simulations.
Extended reading notes
Core claim
The central claim is that for a linear drift model $b_\theta = \phi_0 + \sum_{j=1}^p \theta^j_0 \phi_j$ with $s$ nonzero coefficients, the adaptive Lasso estimator $\hat\theta$ defined by minimizing $L_T(\theta) + \lambda \sum_j w_j |\theta_j|$ satisfies $\mathrm{P}(\hat\theta =_s \theta_0) \to 1$ as $T \to \infty$ under assumptions (A1)-(A3), (B1)-(B3), and (C) (Theorem 2.5). Moreover, for any direction $\alpha \in \mathbb{R}^s$, $\sqrt{T}\, s_T^{-1} \alpha^\top (\hat\theta_S - \theta_{0,S})$ converges in distribution to a standard normal (Theorem 2.6), with the normalization $s_T^2 = \alpha^\top (C^{SS}_\infty)^{-1}\alpha$. The paper frames this as the adaptive Lasso combining exact variable selection with oracle-style efficiency for the active drift coefficients, which the standard Lasso cannot achieve with a single tuning parameter.
Load-bearing premise
The load-bearing premise is that the random fluctuations of the empirical information matrix concentrate around their averages with sub-Gaussian tails and a constant that stays bounded as the dimension grows; if that concentration fails, the probability bounds behind exact support recovery and asymptotic normality collapse.
Editorial extensions
If this is right
- Exact support recovery: the estimated nonzero set equals the true nonzero set with probability approaching one, so downstream analysis can condition on the selected support.
- Oracle inference on active coefficients: $\sqrt{T}\, s_T^{-1} \alpha^\top(\hat\theta_S - \theta_{0,S})$ is asymptotically standard normal, enabling confidence intervals and tests for the drift coefficients.
- Feasible regimes: $p$ can grow exponentially in the observation horizon, like $\exp((dT)^\delta)$, while the sparsity $s$ grows only polynomially, making the theory relevant for $p\gg d$ problems.
- Pre-estimator options: with a positive-definite expected Fisher information, the Lasso itself works as the pre-estimator; under a partial-orthogonality condition, a cheap marginal estimator covers the $p\gg d$ case.
- One tuning parameter suffices: the same $\lambda$ can satisfy both the support-recovery condition (B3) and the normality condition (2.8), overcoming the standard Lasso's inability to do both.
Reading between the lines
- The proof is modular in the concentration assumption: any verification of (C), for instance through exponential bounds derived from the process's mixing or ergodicity, plugs into the same theorems, so the scope of the results is set by how widely (C) can be checked.
- For the numerical basis $\phi_i(x)=\cos((i+1)x)$, the pairwise inner products have Lipschitz constants growing with $i$, so Proposition 4.2's sufficient condition for (C) is not literally satisfied in the simulations; verifying (C) directly for that basis is a concrete open check.
- The sub-exponential relaxation (C') should preserve support recovery but shrink the allowed sparsity $s$; spelling out the exact rate trade-off would quantify how much generality costs in statistical power.
- Theorem 2.6 makes it feasible to build confidence intervals for individual drift coefficients in high-dimensional diffusions, a practical procedure the paper illustrates but does not fully develop.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies continuous-time observation of a d-dimensional ergodic diffusion whose drift is linear in a p-dimensional parameter θ0 through a dictionary Φ. Under a sparsity assumption on θ0, it analyzes the adaptive Lasso estimator of the drift parameter. The main results, Theorems 2.5 and 2.6, assert that under Assumptions (A1)-(A3), (B1)-(B3) and (C), the adaptive Lasso is sign-consistent and that the estimated active subvector is asymptotically normal with oracle covariance. The proof decomposes the failure of sign consistency into five bad events whose probabilities are bounded using the sub-Gaussian concentration Assumption (C) on the quadratic functional H_v, together with bounds on the empirical Fisher information and the martingale term. The paper also proposes a marginal pre-estimator for p >> d under a partial orthogonality condition, and reports simulations for the basis φ_i(x)=cos((i+1)x) with d=5, p=30, T=10.
Significance. If the main claims hold, the paper makes a useful contribution: it extends adaptive-Lasso oracle properties from static high-dimensional regression to ergodic diffusions with growing parameter dimension, gives explicit tuning-parameter relationships, and proposes a pre-estimator for genuinely high-dimensional cases. The proof is organized as a standard sequence of probability bounds derived from Assumption (C); there is no circularity, and the handling of the marginal estimator is constructive. However, the paper's only explicit sufficient condition for Assumption (C) in the growing-p regime, Proposition 4.2, is not satisfied by the paper's own simulation basis, and no other concrete model is shown to satisfy (C). Because every probability bound in Theorems 2.5 and 2.6 passes through Assumption (C), this is a load-bearing gap that must be addressed before the numerical claims and the scope of the theorems can be accepted.
major comments (3)
- [Section 5, Eq. (5.1), and Proposition 4.2] The claim in Section 5 that "Assumption (C) is verified by Proposition 4.2" is not correct for the simulation basis. For φ_i(x)=cos((i+1)x), the function g_{ij}(x)=⟨φ_i,φ_j⟩ has Lipschitz constant d(i+j+2): already in dimension d=1, the derivative of cos((i+1)x)cos((j+1)x) has sup norm i+j+2. Hence the Lipschitz constants are not uniformly bounded in p, and Proposition 4.2, which explicitly requires uniformly bounded constants, cannot be invoked. If one instead tries to use Theorem 4.1 with K=∥Q∥_op, then for Q_{ij}=i+j+2 one has ∥Q∥_op=Θ(p^2); for the simulated values p=30, T=10 and s≈5-10, the quantity K s log p/(τ_min^2 T) appearing in Assumption (B3) is of order 10^3 and does not vanish. Thus the numerical study does not provide evidence for the asymptotic regime of Theorems 2.5-2.6.
- [Section 2, Assumption (C), and Lemmas 6.1, 6.4-6.9] Every probability bound used in the proofs of Theorems 2.5 and 2.6, including Lemmas 6.1, 6.5-6.9 and Propositions 6.2-6.4, relies on Assumption (C). The only sufficient condition supplied for (C) is Proposition 4.2, and Major Comment 1 shows that this condition fails for the paper's own example. As the manuscript stands, no concrete high-dimensional dictionary is proved to satisfy (C), so the main theorems are conditional on a concentration property that is neither verified nor instantiated. The revision should either prove (C) for a nontrivial class of dictionaries with explicit control of K, or state the results purely under Assumption (C) and remove the verification sentence in Section 5 unless a valid example is supplied.
- [Section 3.1] For nonlinear dictionaries the statement that l_min(C_∞)>0 implies p≤d is not correct. Since C_∞=E[Φ(X_0)^TΦ(X_0)], its rank can exceed d when the dictionary functions are nonlinear; for instance, with d=1, φ_1(x)=x, φ_2(x)=x^2 and X_0 uniform on ±1, one obtains C_∞=I_2. This does not affect the main theorems, but the motivation for the marginal estimator in Section 3.2 should be reworded.
minor comments (4)
- [Remark 4.3] The symbol d is reused for both the dimension of the diffusion and the exponent in Assumption (C'), which is confusing; an exponent like q or γ would be clearer.
- [Assumption (C)] There is a small typo in the statement of Assumption (C): "there exists a constant K such that that for all μ∈R" contains a duplicated "that".
- [Eq. (5.1)] In the simulation model b_{θ0}(x)=3s x + Σ θ_i cos((i+1)x), the scalar s is used for a model constant even though s denotes the sparsity level elsewhere in the paper; this is potentially misleading and should be renamed.
- [References] References [7] and [8] are identical, and the bibliography would benefit from a consistency check for other duplicates or missing page ranges.
Circularity Check
No significant circularity: the support-recovery and asymptotic-normality theorems are proved from the stated assumptions (A1)-(A3), (B1)-(B3), and (C), and no fitted quantity is renamed as a prediction.
full rationale
The main derivation chain is conditional: Theorem 2.5 decomposes non-sign-consistency into events B1-B5 and bounds each with Lemmas 6.5-6.9, all built from Assumption (C) via Lemma 6.1 and Propositions 6.2-6.4. Theorem 2.6 uses Theorem 2.5 and the same concentration bounds. No parameter is fitted to data and then reported as a prediction; the tuning parameter lambda and weights w_j are inputs to the optimization, not calibrated to the target support. The only self-citations appear in auxiliary examples and numerics: [13] supplies a Lasso convergence rate used to illustrate that (B1) can be met, [4] motivates the simulation design, and [12] is cited for tuning and OU-type results. These are not the objects proved by Theorems 2.5-2.6 and are not presupposed by the main proof. Assumption (C) is a hypothesis; the paper's verification via Proposition 4.2 may be questionable for the cosine basis used in simulations, but a failed or incomplete verification is a correctness concern, not a circular reduction. No circular step can be exhibited, because no central claim is assumed as input or reduced to a fitted parameter by construction.
Assumptions & free parameters
free parameters (2)
- adaptive Lasso tuning parameter λ =
not a single value; chosen to satisfy (B3) and (2.8), and by cross-validation in simulations
- Lasso pre-estimator tuning parameter λ_l =
selected by cross-validation in numerical studies
assumptions (8)
- domain assumption Assumption (A1): drift bθ is globally Lipschitz and satisfies ⟨bθ(x), x⟩ ≥ M ∥x∥² for a constant M, giving a unique strong solution and ergodicity.
- domain assumption Assumption (A2): the initial value X0 follows the invariant distribution, so X is strictly stationary.
- domain assumption Assumption (A3): I_S = E[Φ_S(X0)ᵀΦ_S(X0)] is positive definite with minimum eigenvalue τmin > 0.
- domain assumption Assumption (B1): the initial estimator θhat is r_T-consistent for a proxy η0 with the stated bounds on M1,T and M2,T.
- domain assumption Assumption (B2): the adaptive irrepresentable condition with constant κ < 1.
- ad hoc to paper Assumption (B3): the explicit growth conditions on s, p, λ, K, M, L, τmin, M1,T, M2,T, θmin that make the probabilities in Lemmas 6.5-6.9 vanish.
- ad hoc to paper Assumption (C): sub-Gaussian concentration of the quadratic functional H_v(X[0,T]) with constant K, for all unit vectors v.
- ad hoc to paper Proposition 4.2 sufficient condition: each inner product ⟨φ_i, φ_j⟩ is Lipschitz with constants uniformly bounded in p, s, d.
Cite this review
Pith. "Pith review of Consistent support recovery for high-dimensional diffusions." pith.science (2026). https://pith.science/paper/JC3PSD5D
@misc{pith2026250116703,
author = {Pith},
title = {Pith review of: Consistent support recovery for high-dimensional diffusions},
year = {2026},
howpublished = {\url{https://pith.science/paper/JC3PSD5D}},
note = {Machine review of arXiv:2501.16703}
}
read the original abstract
Statistical inference for stochastic processes has advanced significantly due to applications in diverse fields, but challenges remain in high-dimensional settings where parameters are allowed to grow with the sample size. This paper analyzes a d-dimensional ergodic diffusion process under sparsity constraints, focusing on the adaptive Lasso estimator, which improves variable selection and bias over the standard Lasso. We derive conditions under which the adaptive Lasso achieves support recovery property and asymptotic normality for the drift parameter, with a focus on linear models. Explicit parameter relationships guide tuning for optimal performance, and a marginal estimator is proposed for p>>d scenarios under partial orthogonality assumption. Numerical studies confirm the adaptive Lasso's superiority over standard Lasso and MLE in accuracy and support recovery, providing robust solutions for high-dimensional stochastic processes.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Concentration of scalar ergodic dif- fusions and some statistical implications
Aeckerle-Willems, C., and Strauch, C. Concentration of scalar ergodic dif- fusions and some statistical implications. Annales de l’Institut Henri Poincar´ e, Prob- abilit´ es et Statistiques 57, 4 (2021), 1857 – 1887
work page 2021
-
[2]
Polynomial rates via deconvolution for nonparametric estimation in McKean-Vlasov SDEs
Amorino, C., Belomestny, D., Pilipauskait ˙e, V., Podolskij, M., and Zhou, S.-Y. Polynomial rates via deconvolution for nonparametric estimation in McKean-Vlasov SDEs. Probability Theory and Related Fields (2024)
work page 2024
-
[3]
Parameter estimation of discretely observed interacting particle systems
Amorino, C., Heidari, A., Pilipauskait˙e, V., and Podolskij, M. Parameter estimation of discretely observed interacting particle systems. Stochastic Processes and their Applications 163 (2023), 350–386
work page 2023
-
[4]
Sampling effects on lasso esti- mation of drift functions in high-dimensional diffusion processes
Amorino, C., Pina, F., and Podolskij, M. Sampling effects on lasso esti- mation of drift functions in high-dimensional diffusion processes. arXiv preprint arXiv:2408.08638 (2024)
-
[5]
Asteriou, D., and Hall, S. G. Applied econometrics. Bloomsbury Publishing, 2021
work page 2021
-
[6]
Bailey, N. T. Some stochastic models for small epidemics in large populations. Applied Statistics (1964), 9–19
work page 1964
-
[8]
On nonparametric es- timation of the interaction function in particle system models
Belomestny, D., Podolskij, M., and Zhou, S.-Y. On nonparametric es- timation of the interaction function in particle system models. arXiv preprint arXiv:2402.14419 (2024)
arXiv 2024
-
[9]
Bishwal, J. P. N., et al. Estimation in interacting diffusions: Continuous and discrete sampling. Applied Mathematics 2 , 9 (2011), 1154–1158
work page 2011
Show all 44 references
-
[10]
Statistics for high-dimensional data: methods, theory and applications
B¨uhlmann, P., and V an De Geer, S. Statistics for high-dimensional data: methods, theory and applications . Springer Science & Business Media, 2011
2011
-
[11]
Maximum likelihood estimation of potential energy in interacting particle systems from single-trajectory data
Chen, X. Maximum likelihood estimation of potential energy in interacting particle systems from single-trajectory data. Electronic Communications in Probability 26 (2021), 1–13
2021
-
[12]
On Dantzig and Lasso estimators of the drift in a high dimensional Ornstein-Uhlenbeck model
Ciolek, G., Marushkevych, D., and Podolskij, M. On Dantzig and Lasso estimators of the drift in a high dimensional Ornstein-Uhlenbeck model. Electronic Journal of Statistics 14 , 2 (2020)
2020
-
[13]
On Lasso estimator for the drift function in diffusion models
Ciolek, G., Marushkevych, D., and Podolskij, M. On Lasso estimator for the drift function in diffusion models. Bernoulli, to appear (2022)
2022
-
[14]
De Gregorio, A., and Iacus, S. M. Adaptive lasso-type estimation for multi- variate diffusion processes. Econometric Theory 28 , 4 (2012), 838–860
2012
-
[15]
The lan property for mckean–vlasov models in a mean-field regime
Della Maestra, L., and Hoffmann, M. The lan property for mckean–vlasov models in a mean-field regime. Stochastic Processes and their Applications 155 (2023), 109–146
2023
-
[16]
On the lasso for graphical continuous lyapunov models
Dettling, P., Drton, M., and Kolar, M. On the lasso for graphical continuous lyapunov models. Proceedings of Machine Learning Research 236 (2024), 514–550
2024
-
[17]
On Lasso and Slope drift estimators for L´ evy- driven Ornstein–Uhlenbeck processes
Dexheimer, N., and Strauch, C. On Lasso and Slope drift estimators for L´ evy- driven Ornstein–Uhlenbeck processes. Bernoulli 30 , 1 (2024), 88–116
2024
-
[18]
Transportation cost-information in- equalities and applications to random dynamical systems and diffusions.Ann
Djellout, H., Guillin, A., and Wu, L. Transportation cost-information in- equalities and applications to random dynamical systems and diffusions.Ann. Probab. 32, 3B (2004), 2702–2732
2004
-
[19]
Antoniadis
F an, J.Comments on ≪Wavelets in statistics: A review ≫ by A. Antoniadis. Statis- tical Methods & Applications 6 , 2 (1997), 131–138
1997
-
[20]
Variable selection via nonconcave penalized likelihood and its oracle properties
F an, J., and Li, R. Variable selection via nonconcave penalized likelihood and its oracle properties. Journal of the American statistical Association 96 , 456 (2001), 1348–1360
2001
-
[21]
Nonconcave penalized likelihood with a diverging number of parameters
F an, J., and Peng, H. Nonconcave penalized likelihood with a diverging number of parameters. The Annals of Statistics 32 , 3 (2004), 928 – 961
2004
-
[22]
Sparse inference of the drift of a high- dimensional Ornstein–Uhlenbeck process
Ga¨ıffas, S., and Matulewicz, G. Sparse inference of the drift of a high- dimensional Ornstein–Uhlenbeck process. Journal of Multivariate Analysis 169 (2019), 1–20. 30
2019
-
[23]
Parametric inference for small variance and long time horizon mckean-vlasov diffusion models
Genon-Catalot, V., and Lar ´edo, C. Parametric inference for small variance and long time horizon mckean-vlasov diffusion models. Electronic Journal of Statis- tics 15 , 2 (2021), 5811–5854
2021
-
[24]
Parametric inference for ergodic mckean- vlasov stochastic differential equations
Genon-Catalot, V., and Lar´edo, C. Parametric inference for ergodic mckean- vlasov stochastic differential equations. Bernoulli 30 , 3 (2024), 1971–1997
2024
-
[25]
Holden, A. V. Models of the stochastic activity of neurones , vol. 12. Springer Science & Business Media, 2013
2013
-
[26]
A., and Johnson, C
Horn, R. A., and Johnson, C. R.Hermitian and symmetric matrices. Cambridge University Press, 1985, p. 167–256
1985
-
[27]
The Annals of Statistics 36 (05 2008)
Huang, J., Horowitz, J., and Ma, S.Asymptotic properties of bridge estimators in sparse high-dimensional regression model. The Annals of Statistics 36 (05 2008)
2008
-
[28]
Adaptive lasso for sparse high-dimensional regression models
Huang, J., Ma, S., and Zhang, C.-H. Adaptive lasso for sparse high-dimensional regression models. Statistica Sinica (2008), 1603–1618
2008
-
[29]
C., and Basu, S
Hull, J. C., and Basu, S. Options, futures, and other derivatives . Pearson Education India, 2016
2016
-
[30]
yuima: The YUIMA Project: A Framework for Simulation and Inference of SDE Models , 2020
Iacus, A., Masuda, H., and Yoshida, L. yuima: The YUIMA Project: A Framework for Simulation and Inference of SDE Models , 2020. R package version 1.6.7
2020
-
[31]
A note on limit theorems for multivariate martingales
K¨uchler, U., and Sørensen, M. A note on limit theorems for multivariate martingales. Bernoulli 5 , 3 (1999), 483–493
1999
-
[32]
Parameter estimation of path-dependent mckean-vlasov stochastic differential equations
Liu, M., and Qiao, H. Parameter estimation of path-dependent mckean-vlasov stochastic differential equations. Acta Mathematica Scientia 42 , 3 (2022), 876–886
2022
-
[33]
Causal modeling with stationary diffusions
Lorch, L., Krause, A., and Sch ¨olkopf, B. Causal modeling with stationary diffusions. Proceedings of the 27th International Conference on Artificial Intelligence and Statistics (AISTATS) 238 (2024), 1927–1935
2024
-
[34]
W., and Hansen, N
Mogensen, S. W., and Hansen, N. R. Markov equivalence of marginalized local independence graphs. Ann. Statist. 48 , 1 (2020), 539–559
2020
-
[35]
W., and Hansen, N
Mogensen, S. W., and Hansen, N. R.Graphical modeling of stochastic processes driven by correlated noise. Bernoulli 28 , 4 (2022), 3023–3050
2022
-
[36]
Nourdin, I., and Viens, F. G. Density formula and concentration inequalities with malliavin calculus. EJS 14 , 14 (2009), 2287–2309
2009
-
[37]
Papanicolaou, G. C. Diffusion in random media. In Surveys in applied mathe- matics. Springer, 1995, pp. 205–253
1995
-
[38]
Ricciardi, L. M. Diffusion processes and related topics in biology, vol. 14. Springer Science & Business Media, 2013
2013
-
[39]
Rogers, L. C. G., and Williams, D. Diffusions, Markov Processes, and Mar- tingales, 2 ed. Cambridge Mathematical Library. Cambridge University Press, 2000. 31
2000
-
[40]
Transportation inequalities for stochastic differential equations driven by a fractional Brownian motion
Saussereau, B. Transportation inequalities for stochastic differential equations driven by a fractional Brownian motion. Bernoulli 18 , 1 (2012), 1 – 23
2012
-
[41]
Sharrock, L., Kantas, N., Parpas, P., and Pavliotis, G. A. Online pa- rameter estimation for the mckean–vlasov stochastic differential equation. Stochastic Processes and their Applications 162 (2023), 481–546
2023
-
[42]
Sur, P., and Cand `es, E. J. A modern maximum-likelihood theory for high- dimensional logistic regression. Proceedings of the National Academy of Sciences 116, 29 (2019), 14516–14525
2019
-
[43]
Concentration anal- ysis of multivariate elliptic diffusions
Trottner, L., Aeckerle-Willems, C., and Strauch, C. Concentration anal- ysis of multivariate elliptic diffusions. J. Mach. Learn. Res. 24 (2023), Paper No. [106], 38
2023
-
[44]
Electron
V arvenne, M.Concentration inequalities for stochastic differential equations with additive fractional noise. Electron. J. Probab. 24 (2019), Paper No. 124, 22
2019
-
[45]
The adaptive lasso and its oracle properties
Zou, H. The adaptive lasso and its oracle properties. Journal of the American statistical association 101 , 476 (2006), 1418–1429. 32
2006
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.