{"id":"427891f8-7b79-4fde-b798-f3b6a75ee616","arxiv_id":"2412.17779","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A quasi-likelihood and adaptive Lasso framework is proposed for estimating ergodic network SDE models with nonlinear drift, stochastic volatility, and directed graphs.","lead":"This paper introduces a class of network stochastic differential equations where each node's dynamics combine a momentum term, neighbor interactions, and stochastic volatility, and develops quasi-likelihood estimators plus an adaptive Lasso procedure to recover the underlying directed graph. The intended payoff is a parametric, interpretable framework for high-dimensional networked time series, with applications to high-frequency financial data and causal analysis.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Growing-dimension graph recovery is asserted via a fixed-dimensional QLA extension and a pointwise union bound; no proof supports the p>n regime used in the real-data application.","rationale":"The Reader's verdict REJECT is supported. The central advertised contribution is not the fixed-d quasi-likelihood theory but the ability to handle very high-dimensional time series via sparsity, growth of d and |E|, and exact graph recovery. Theorem 2 is qualified 'for every d > 0' with a fixed-d asymptotic event; it does not provide a genuinely non-asymptotic bound uniform in d. More decisively, the graph-recovery section introduces d(d−1) auxiliary edge weights and then cites fixed-dimensional QLA results as Theorem 3 without proof. Adaptive-Lasso selection consistency (Theorem 5) is proved by a union bound over null edges whose terms are only individually o(1); for d growing, this is insufficient. The real-data example is exactly the p > n regime where the required initial QL estimator cannot be expected to be consistent, a point the paper does not disclose. These are internal gaps in the argument, not disagreements with external consensus. A careful re-derivation with explicit rates and a p > n simulation would settle the matter. Because these gaps affect the paper's headline claims, the REJECT verdict stands; no change needed.","tokens_in":19749,"tokens_out":4606,"duration_ms":48285,"concrete_test":"Re-derive the proof of Theorem 5 with d = d_n → ∞ and add explicit tail bounds: check whether the KKT-based bound for a null edge, P(ŵ_{n,ij} ≠ 0), is O(d^{−2}) under (L2)–(L3); if not, the conclusion P(Ĝ_n = G) → 1 fails for growing d. As a numerical companion, run the proposed two-step pipeline on a simulated p > n instance (e.g., d = 100, n = 100, sparse directed graph) and record whether the unpenalized estimator (19) exists and the adaptive weights are finite; a divergence or non-uniqueness confirms the initial-estimator condition is not satisfied in the advertised regime.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's advertised scope is that parameter and topology estimation work for growing networks (Section 1, contribution ii; abstract). Theorems 3–5 are the load-bearing part of that claim, but they inherit a fixed-dimensional quasi-likelihood theory. Theorem 3 is stated as a recall of [25, Thm 13] and [12], yet those results are for a parameter space of fixed dimension; extending them to π_β + d(d−1) parameters with d = d_n → ∞ is nontrivial and no growth conditions on d versus n are given. The proof of Theorem 5 then uses a union bound over the O(d²) null edges: each summand is shown to tend to 0 pointwise, but for d growing this only yields exact recovery if the term is o(1/d²); no such rate is proved. This is not a technicality: the S&P100 application has d = 99 (π_d + d² ≈ 9900) and n = 1596, so the unpenalized QL initial estimator (19) required for adaptive weights (21)–(22) lives in a p > n regime where it need not be consistent and may not be unique. Thus neither Theorem 5's graph-recovery guarantee nor the high-dimensional applicability is supported as written. Fixed-d finite-sample results (e.g., Theorem 2 with d fixed) are not the issue; the gap is specifically the growing-network / p > n regime that the paper advertises.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a network SDE framework in which each node follows an SDE with a momentum term, a network interaction term, and a state-dependent stochastic volatility term. Estimation is based on high-frequency discrete observations and quasi-likelihood. For a known graph, Theorem 2 gives a non-asymptotic ℓ2 error bound under scaling conditions on the number of edges and observation time. For an unknown directed graph, the paper augments the parameter vector with edge weights and proposes an adaptive Lasso estimator, claiming consistency (Theorem 4) and exact graph recovery (Theorem 5). The paper includes simulations for Erdős–Rényi, polymer, and stochastic block model graphs, and applies the method to S&P100 high-frequency data.","tokens_in":20010,"tokens_out":6768,"duration_ms":64515,"significance":"If the claims were established, the framework would be a notable contribution: it extends network autoregressive and graph Ornstein–Uhlenbeck models to nonlinear drift and stochastic volatility, handles directed edges, and provides finite-sample-type bounds that make the role of graph sparsity and sampling frequency explicit. The fixed-dimensional known-graph analysis and the explicit linear-model estimator in Section 3.1 are useful, and the simulation studies are reasonable. However, the advertised central contribution—parameter and topology estimation for growing networks with unknown graph—is not backed by the proofs. Theorem 3 merely recalls fixed-dimensional QLA results and applies them to a parameter space of dimension πβ + d(d−1) without growth conditions; Theorem 5 relies on a union bound over O(d²) null edges without the necessary uniform rates; and the real-data application operates in a p > n regime where the required initial quasi-likelihood estimator is not justified. The paper is also transparent in Remark 3 that Theorem 2 could be replaced by an estimator-scaling assumption that is essentially the conclusion. These gaps are load-bearing for the paper's main claims.","major_comments":[{"comment":"Theorem 3 is stated as a recall of [25, Theorem 13] and [12, Theorem 1, Lemma 4], but those are fixed-dimensional quasi-likelihood asymptotics. In model (18) the augmented parameter vector for (θ, w) has dimension πα + πβ + d(d−1), and the paper explicitly allows d to grow with n (Section 1, contribution (ii)). No growth conditions linking d to n are supplied, and no extension of the polynomial-type large deviation inequalities of [25] to this growing parameter space is proved. Since Theorems 4 and 5 both rely on the Γn-consistency of the initial estimator (19), the graph-recovery guarantees for growing networks rest on an unproven high-dimensional QLA.","section":"Section 4, Theorem 3"},{"comment":"The proof that no null edge is selected uses a union bound over the O(d²) pairs (i, j) not in E. Each summand is shown to tend to 0 pointwise via the KKT conditions, but for d growing the sum tends to 0 only if the individual probabilities are o(1/d²) or if additional structure controls the number of null edges. No such rate is established; condition (L3), which only requires ˇγw_{n,d}√(nΔn) → ∞, does not control the union bound. Consequently, the conclusion P(Ĝn = G) → 1 is not proved for growing d.","section":"Section 8, Proof of Theorem 5, Step 1"},{"comment":"In the S&P100 application, d = 99 and n = 1596, so the augmented parameter vector for (θ, w) has dimension πd + d² ≈ 9900, exceeding the sample size. The initial quasi-likelihood estimator (19), used to construct the adaptive weights (21)–(22), is an unpenalized QL estimator over this p > n parameter space; it need not be unique, consistent, or even well-defined. The paper gives no first-step regularization or alternative construction of the initial estimator, so the adaptive Lasso procedure is not justified in the reported high-dimensional regime.","section":"Section 6, real-data application; Section 4, estimator (19)"},{"comment":"Theorem 2 claims finite-sample bounds for growing networks, but its proof invokes the event {|Γn^{-1}(θ̂n−θ0)| ≤ r} with probability at least 1 − CL/r^L 'due to the results in [25]'. Those results do not provide a uniform bound over the dimension d; the constants CL and L, and the radius r required to make the probability high, may depend on πd. Moreover, Remark 3 acknowledges that Theorem 2 could instead be proved under the estimator-scaling assumption sup_n E|Γn^{-1}(θ̂−θ0)|² ≲ πd, which is essentially the conclusion one would need to prove for the QLE in growing dimension. As written, the known-graph finite-sample guarantee is therefore also not established in the growing-network regime advertised in the abstract.","section":"Section 3, Remark 3; Section 8, Proof of Theorem 2"}],"minor_comments":[{"comment":"The sentence 'The degree distribution is shown in .' is incomplete; a figure or table reference is missing.","section":"Section 6"},{"comment":"The condition labelled (8) is referenced before it is defined; equation (8) appears later in Example 2.2. Please renumber or reorder.","section":"Section 2, Example 2.1"},{"comment":"The notation τmax(−µId) is undefined; the largest singular value of −µId is |µ|, not −µ, and the displayed inequality should be clarified.","section":"Section 2, Example 2.1"},{"comment":"The column heading 'Bound Mean Error' is ambiguous; specify whether 'Mean Error' is the empirical mean of the ℓ2 error or its square.","section":"Table 1"},{"comment":"The function c·tanh(x/c) is a scaled hyperbolic tangent, not a sigmoid; use a consistent term such as 'smooth clipping' throughout.","section":"Section 5, model (25)"},{"comment":"The proof of Theorem 2 spells 'Cauchy-Schwartz'; this should be 'Cauchy-Schwarz'.","section":"Section 8"},{"comment":"The statement 'No regularization is required on the diffusion part in this example' is unclear because the table has no LASSO column for the α parameters; clarify whether α was unpenalized by construction.","section":"Table 2 caption"}],"recommendation":"reject","confidential_remarks":"The main advertised contribution—recovering the directed graph of a network SDE when the number of candidate edges grows with n—is not supported by the proofs. The gap is structural: the authors recall fixed-dimensional QLA results and then apply them in a parameter space of dimension d(d−1) without growth conditions. The S&P100 application cannot be justified by the theory as it stands. I would not rule out a future revision that either establishes high-dimensional QLA bounds or restricts the claims to settings with p < n, but the current manuscript does not deliver what it advertises. Given the journal's standards, I recommend rejection while encouraging the authors to consider a more modest claim set or a rigorous growing-dimension theory."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The bottom line: this is a useful model class with a real gap between its advertised growing-network guarantees and what the proofs actually support. The N-SDE framework—nonlinear drift interactions, state-dependent stochastic volatility, directed edges, high-frequency quasi-likelihood—is genuinely new relative to NAR and GrOU models. The fixed-dimensional known-graph part (Theorem 2, linear drift estimates, simulation table) is plausible, and the empirical work is honest. I'd send this to a careful referee, but I wouldn't rely on the graph-recovery theorems as stated.\n\nWhat's new and good: the model nests NAR and GrOU into a continuous-time setup with nonlinear interaction terms and volatility; Example 2.2's radial-basis ergodicity condition is a nice touch; the explicit linear-drift estimator (17) is a clean contribution. The empirical section tests Erdős–Rényi, directed polymer, and stochastic block models, and the error-bound table aligns with Theorem 2's rate |E|/(nΔ_n). That part looks like real work.\n\nSoft spots: the high-dimensional claims—the advertised 'growing networks' and exact directed graph recovery—rest on Theorem 3, which simply recalls fixed-dimensional QLA asymptotics and applies them to a parameter space of dimension πβ + d(d−1). No growth conditions on d vs n are given. Theorem 5's union bound over O(d^2) null edges shows each exclusion probability goes to zero pointwise; with d growing, you'd need a rate of o(1/d^2), which is not proved. And the S&P100 application runs in a p > n regime (9900 parameters, 1596 observations). The adaptive Lasso step requires a consistent unpenalized initial estimator for the augmented model; in p > n that estimator need not be consistent or unique. The paper does not flag this. The fixed-dimensional parts are not the problem; the load-bearing gap is precisely the growing-network / p > n regime the abstract advertises.\n\nBottom line: the model and the fixed-d case are worth a serious look. The graph-recovery theory is not supported as written. A referee could push for a restriction to fixed d, or for proper concentration inequalities and growth conditions. I'd accept it for review, but with a note that the advertised scope is currently unmet.","headline":"Useful new model class, but the advertised growing-network graph-recovery claims are not backed by the proofs; the fixed-dimensional parts are solid.","tokens_in":20569,"tokens_out":1962,"would_cite":false,"duration_ms":18738,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M05","62F12","60H10","62J07"],"pacs":[],"model":"deepseek-v4-flash","headline":"For network SDEs, drift parameters and the directed graph can both be estimated from high-frequency observations of a single trajectory.","keywords":["network stochastic differential equations","directed graph recovery","adaptive lasso","quasi-likelihood estimation","high-frequency observations","ergodic diffusion","graph scaling","stochastic volatility"],"falsifier":"Run the paper's two-step adaptive Lasso on data from model (25) with $d = 10$, $n = 50$, and $\\Delta_n = 1/50$ (so $n\\Delta_n = 1$ and the 90 candidate edge weights are fewer than the sample size), and repeat with $n = 200, 800$ while holding $d$ and $\\Delta_n$ fixed; if the empirical frequency $P(\\hat G_n = G)$ does not approach one over many replications, the consistency claim of Theorem 5 fails.","tokens_in":19441,"feed_emoji":"🕸️","tokens_out":6977,"duration_ms":63636,"temperature":0.7,"pith_summary":"The paper introduces a network stochastic differential equation (N-SDE) model in which each node evolves according to its own dynamics, feedback from neighbouring nodes, and a stochastic-volatility diffusion term, and it claims that this model can be estimated from high-frequency observations of a single trajectory. When the directed graph is known, the quasi-likelihood estimator is shown to satisfy a non-asymptotic error bound: the squared error is controlled by the number of graph parameters divided by total observation time, up to constants describing model regularity and graph scaling. When the graph is unknown, the paper proposes a two-step adaptive Lasso procedure applied to an augmented model with one weight per candidate directed edge, and proves that the estimated graph equals the true graph with probability tending to one. A closed-form neighbourhood-wise estimator is derived for linear drifts, and simulations plus a real S&P100 data example illustrate the claims. If correct, this gives an interpretable, directed-graph framework for causal and network analysis of high-dimensional continuous-time series.","feed_headline":"Directed network SDEs can be estimated from high-frequency data","feed_subtitle":"Known graph yields explicit error bounds; unknown directed graph is recovered exactly by adaptive Lasso with probability one.","key_machinery":"The workhorse is the quasi-likelihood contrast $\\ell_n(\\alpha,\\beta)$ built from increments of the observed path, together with the scaled parameter rates $\\Gamma_n$, which give $\\alpha$ a $1/\\sqrt{n}$ rate and $\\beta$ a $1/\\sqrt{n\\Delta_n}$ rate. Assumption C(r) imposes a regular contrast: a gradient bound by a square-integrable random variable $\\xi_n$ and a Hessian lower bound $\\mu$, which together turn the Taylor expansion of the contrast into the probability bound of Theorem 2. For graph recovery, the machinery is the least-squares-approximation objective $F_n(\\theta,w) = \\frac12\\langle H_n, (\\theta-\\tilde\\theta, w-\\tilde w)^{\\otimes 2}\\rangle + \\lambda_n\\|(\\theta,w)\\|_{1,\\gamma(n,d)}$ with adaptive weights $\\gamma_{n,ij} \\propto |\\tilde w_{ij}|^{-\\delta}$. The KKT equations for zeroed edges give the no-false-inclusion part of Theorem 5, while sign consistency of the nonzero weights gives the no-missing-edge part.","core_discovery":"The central claim is that the pair consisting of the drift/diffusion parameters and the directed topology of an ergodic N-SDE can be identified from discrete high-frequency data, with explicit rates for both. For a known graph, Theorem 2 states that with high probability $|\\hat\\theta_n - \\theta_0|^2 \\le \\frac{4\\xi_n^2}{\\mu^2}\\,K\\varepsilon$, where $K = \\pi_d/|G_d|$ controls how many parameters each edge carries and $\\varepsilon = |G_d|/(n\\Delta_n)$ is the graph size relative to total observation time; the proof Taylor-expands the quasi-likelihood contrast on an event where the scaled estimation error is bounded, using the regularity constants $\\xi_n$ and $\\mu$. For an unknown graph, the paper augments the drift with edge weights $w_{ij}$, forms an initial quasi-likelihood estimate $(\\tilde\\theta, \\tilde w)$, and then minimises a least-squares-approximation objective with adaptive $\\ell^1$ weights; Theorem 5 shows $P(\\hat G_n = G) \\to 1$ under rate conditions on the adaptive weights and the information matrix. The linear-drift case reduces each node's estimation to a small regression on its neighbourhood, so computation tracks the graph rather than the full $d^2$ problem.","pith_inferences":["The graph-recovery theorems require a consistent initial quasi-likelihood estimator in the augmented $d(d-1)$-parameter model, which is only plausible when the number of candidate edges is smaller than the sample size; the paper's S&P100 application ($p \\approx 9900$, $n = 1596$) sits outside that regime and should be read as heuristic rather than as covered by Theorem 5.","The exact-recovery result suggests a practical diagnostic: replicating the two-step procedure on subsampled data and checking whether the estimated support stabilises would indicate whether the asymptotic regime is being approached.","Because the adjacency matrix enters the drift, a zero estimated weight means no Granger-type predictive effect in continuous time, but a causal reading of the recovered directed graph still requires the usual assumption that no unobserved common drivers exist.","The error bound depends on the finite-sample identifiability constant $\\mu$, so the practical takeaway is that sparsity alone is not enough; the strength of the drift's mean reversion also governs estimation quality."],"forward_implications":["For a known directed graph, the squared parameter error decays like $|E|/(n\\Delta_n)$, so holding the observation window fixed and merely sampling more frequently does not improve precision unless the number of observations grows.","The linear-drift estimator (17) estimates each node's parameters from a regression on that node's neighbours, making estimation feasible for large $d$ when neighbourhoods are small.","When the graph is unknown, the adaptive Lasso recovers the exact directed edge set with probability tending to one, so one-way influence can be read off the support of the estimated edge weights.","The non-asymptotic bound implies a quantitative trade-off: adding edges or adding parameters per edge must be compensated by proportionally longer observation time to keep the error below a fixed level.","The framework extends network autoregressive and graph Ornstein-Uhlenbeck models to nonlinear drifts and state-dependent volatility while keeping parameter estimation interpretable."],"supporting_citations":[{"why":"Supplies the polynomial-type large deviation inequalities and quasi-likelihood analysis used to control the event on which the scaled estimation error is bounded.","marker":"[25]"},{"why":"Provides the quasi-likelihood estimation framework for diffusion processes from discrete observations on which the contrast $\\ell_n$ is based.","marker":"[24]"},{"why":"Provides the adaptive Lasso oracle-property framework and the weight choice $\\gamma \\propto |\\text{initial estimate}|^{-\\delta}$.","marker":"[29]"},{"why":"Introduces adaptive Lasso-type estimation for multivariate diffusion processes, the direct template for the graph-recovery procedure.","marker":"[8]"},{"why":"Provides the least-squares-approximation consistency proof that Theorem 4 follows.","marker":"[10]"},{"why":"Supplies the two-step adaptive estimation scheme and the linear-drift estimating equations used in subsection 3.1.","marker":"[21]"},{"why":"Gives the ergodicity and mixing conditions (restated as Theorem 1) used to verify assumptions (A4) and (A5).","marker":"[15]"},{"why":"Provides the identifiability condition (A6) and the quasi-likelihood asymptotic normality invoked as Theorem 3.","marker":"[12]"}],"fun_headline_variants":["Network SDE parameters estimated from high-frequency data","Unknown directed networks recovered from SDE data","Ergodic N-SDEs: estimation with known or unknown graph","Adaptive Lasso identifies SDE network structure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The graph-recovery results depend on a consistent initial quasi-likelihood estimate of the augmented model containing one weight per possible directed edge, which is only plausible when the number of candidate edges is smaller than the sample size, and the paper does not prove the needed asymptotics when the parameter dimension grows with $n$.","fun_headline_variants_meta":{"raw":{"variants":["Network SDE parameters estimated from high-frequency data","Unknown directed networks recovered from SDE data","Ergodic N-SDEs: estimation with known or unknown graph","Adaptive Lasso identifies SDE network structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00026,"raw_usage":{"total_tokens":1666,"prompt_tokens":1097,"completion_tokens":569,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":713,"completion_tokens_details":{"reasoning_tokens":506}},"tokens_in":713,"tokens_out":569,"duration_ms":4891,"temperature":1.0,"reasoning_tokens":506,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:09:38.910063+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's two-step adaptive Lasso on data from model (25) with $d = 10$, $n = 50$, and $\\Delta_n = 1/50$ (so $n\\Delta_n = 1$ and the 90 candidate edge weights are fewer than the sample size), and repeat with $n = 200, 800$ while holding $d$ and $\\Delta_n$ fixed; if the empirical frequency $P(\\hat G_n = G)$ does not approach one over many replications, the consistency claim of Theorem 5 fails.","supporting_citations":[{"cited_title":"Polynomial type large deviation inequalities and quasi- likelihood analysis for stochastic differential equations","cited_arxiv_id":null,"evidence_quote":"Supplies the polynomial-type large deviation inequalities and quasi-likelihood analysis used to control the event on which the scaled estimation error is bounded."},{"cited_title":"Estimation for diffusion processes from discrete ob- servation","cited_arxiv_id":null,"evidence_quote":"Provides the quasi-likelihood estimation framework for diffusion processes from discrete observations on which the contrast $\\ell_n$ is based."},{"cited_title":"The adaptive lasso and its oracle properties","cited_arxiv_id":null,"evidence_quote":"Provides the adaptive Lasso oracle-property framework and the weight choice $\\gamma \\propto |\\text{initial estimate}|^{-\\delta}$."},{"cited_title":"Adaptive LASSO-type estimation for multivariate diffusion processes","cited_arxiv_id":null,"evidence_quote":"Introduces adaptive Lasso-type estimation for multivariate diffusion processes, the direct template for the graph-recovery procedure."},{"cited_title":"Regularized bridge-type estimation with multiple penalties","cited_arxiv_id":null,"evidence_quote":"Provides the least-squares-approximation consistency proof that Theorem 4 follows."},{"cited_title":"Adaptive estimation of an er- godic diffusion process based on sampled data","cited_arxiv_id":null,"evidence_quote":"Supplies the two-step adaptive estimation scheme and the linear-drift estimating equations used in subsection 3.1."},{"cited_title":"On the Poisson equation and diffusion approximation. I","cited_arxiv_id":null,"evidence_quote":"Gives the ergodicity and mixing conditions (restated as Theorem 1) used to verify assumptions (A4) and (A5)."},{"cited_title":"Estimation of an ergodic diffusion from discrete obser- vations","cited_arxiv_id":null,"evidence_quote":"Provides the identifiability condition (A6) and the quasi-likelihood asymptotic normality invoked as Theorem 3."}],"review_version":1}