REVIEW 4 major objections 5 minor 17 references
The paper derives explicit h³-order bias and variance formulas for Strang splittings of one-dimensional SDEs, and shows that the best splitting depends on model geometry: fixed-point linearization for potential models, slow-manifold piecewi
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 08:44 UTC pith:UHRNEJVR
load-bearing objection Solid h^3 splitting-dependent error expansion for Strang estimators; practical 'optimal splitting' advice is oracle-dependent and the bridge to estimator performance is heuristic, but the theory is checkable and worth citing. the 4 major comments →
Choosing optimal Strang splitting estimators of nonlinear stochastic differential equation models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For a one-dimensional SDE dX_t = F(X_t)dt + σdW_t, with F split as A(x−b)+N(x), the Strang scheme's one-step conditional mean has bias h³ B_θ(x)+O(h⁴) and its conditional variance differs from the true variance by h³ ΔV_θ(x)+O(h⁴), where B_θ and ΔV_θ are explicit in A, b, N and σ². Both error coefficients carry the splitting; the paper computes these quantities exactly (Theorem 3.2) and shows that at any fixed point x*, B_θ(x*, x*)=0, so local linearization around equilibria is unbiased at the equilibrium. Simulations then indicate that fixed-point linearization is a strong default for potential models, while slow-fast excitable systems need split by slow-manifold regions.
What carries the argument
Theorem 3.2's expansions (the functions B_θ and ΔV_θ), obtained by pushing the infinitesimal-generator moment expansion through the Strang composition f_{h/2}∘Φ_h∘f_{h/2}. The formulas convert splitting choice into a numerically evaluable error surface, and the paper uses them to define optimal-splitting candidates (minimize |B| or |ΔV| locally or in expectation under the invariant density). The second main device is the adaptive splitting rule that chooses a linearization point from local dynamics — fixed point of the current basin for potentials, slow-manifold points (±1,α) and ±1/√3 for FitzHugh-Nagumo.
Load-bearing premise
The whole optimal-splitting program leans on the unproven correspondence between one-step density accuracy and estimator accuracy (the paper's own simulations show local minimizers of the derived errors can be worse than simpler splittings), and the criteria assume the true parameters are known to compute the error surfaces and fixed points.
What would settle it
Simulate the double-well model at h = 0.01, 0.02, 0.04 for a fixed split, and compare empirical one-step bias and variance to h³ B_θ and h³ ΔV_θ from Theorem 3.2; if the normalized differences (error/h³) are not constant in h to within O(h), the expansion is wrong. Separately, compare parameter MSE for the local-error-minimizing splittings (MoLB, MoLV) against fixed-point linearization; the paper already reports the former lose, so a sharper test is whether a KL-optimal or piecewise-optimal splitting beats fixed-point on parameter recovery.
If this is right
- For any one-dimensional additive-noise SDE, the h³ error coefficients can be computed in closed form and used to compare splittings without running the estimator.
- Fixed-point linearization is a reliable default for gradient/potential systems, giving near-zero one-step bias at equilibria and accurate parameter estimates, plus an explicit plug-in estimator for the diffusion coefficient.
- For slow-fast excitable systems, the fixed point is not representative: piecewise splittings that follow stable slow manifolds yield lower bias and variance in one-step densities and better parameter estimates, especially for oscillatory regimes and larger h.
- Because the error is O(h³) for every splitting, choice of splitting changes finite-sample constants at third order, not the asymptotic rate.
- Adaptive splittings based on minimizing local error measures are computationally heavy and, in the simulations, less accurate than fixed-point or piecewise schemes, suggesting that global/geometry-informed rules are preferable.
Where Pith is reading between the lines
- Editorial inference: the pointwise local-error minimizers (MoLB and MoLV) underperforming in the double-well suggests that pseudo-likelihood accuracy is driven by higher-order density structure (e.g., the Jacobian term in the transformed Gaussian), not only the first two moments; a testable extension is to minimize a distributional divergence such as KL between the Strang and true transition densi
- Editorial inference: the optimal-splitting criteria require the true parameter θ, making them oracle procedures; one could iterate by plugging in a preliminary estimate to locate fixed points or regions, and the paper's piecewise slow-manifold rule (which needs no θ) hints that such data-driven versions may still work.
- Editorial inference: because the h³ terms are explicit, the same calculation should extend to multidimensional systems via tensor-index expansions; if it does, the bias/variance surfaces could serve as a cheap design tool for Strang estimators in higher-dimensional state spaces.
- Editorial inference: the fixed-point zero-bias property suggests a general principle—splitting about invariant structures (equilibria, slow manifolds) minimizes local error; this may transfer to other numerical schemes beyond Strang.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies how the choice of Strang splitting (the decomposition of the drift into an analytically tractable OU part and a nonlinear deterministic part) affects finite-sample inference in parametric SDEs. The main theoretical result, Theorem 3.2, derives h^3-order expansions for the conditional bias and the conditional variance error of the Strang one-step prediction in one spatial dimension, exhibiting explicit dependence on the linearization parameters A, b and the residual nonlinearity N. Based on these error functions, the paper proposes several optimality criteria (MoAB, MoAV, MoLB, MoLV), and uses simulation studies in the double-well potential model and the FitzHugh–Nagumo model to recommend fixed-point linearization for potential models and a piecewise slow-manifold splitting for slow–fast excitable systems. The paper also provides code for reproducing the simulations.
Significance. If the practical claims are borne out, the paper contributes a nontrivial, checkable expansion that quantifies the previously observed finite-sample dependence of Strang splitting estimators on the chosen splitting. The special cases (pure OU, pure deterministic) behave correctly, the proof of Theorem 3.2 is detailed, and the code availability is a strength. However, the optimality theory is heuristic, and the recommended splittings are constructed using true parameter values, so the practical significance as a standalone inference method is currently limited. The theoretical moment expansion alone is a solid contribution, but the paper's central 'optimal splitting' claim requires substantially more justification.
major comments (4)
- [§4.2, Eqs. (4.5), (4.8), (4.9), Fig. 4; §4.3, Eq. (4.13)] The proposed optimal splittings are defined through quantities that depend on the true parameter θ: fixed points of F, the invariant density π in (4.7), and the minimizers of B_θ and ΔV_θ. The simulation studies use the true parameters from Table 1 to construct these splittings. No data-driven rule is given, so the estimators described in §4.2 and §4.3 are not well-defined for real data; the reported accuracy is an oracle result. The paper should either provide a preliminary estimation or profiling strategy, or explicitly frame the claims as conditional on known θ.
- [§4, first paragraph; Fig. 4(a)] The optimality criterion rests on the assertion that a splitting that approximates transition densities accurately will also yield accurate parameter estimates. This is never formalized. More importantly, the paper's own results in Fig. 4 show that the local error-minimizing schemes MoLB and MoLV, which minimize |B_θ(x,c)| and |ΔV_θ(x,c)| pointwise, are noticeably less accurate than MoAV and the fixed-point scheme. Thus pointwise minimization of the derived error measures is not a reliable proxy for estimator accuracy, and the phrase 'optimal splitting' is not justified without an explicit connection between these local errors and the estimator's objective.
- [§3, Theorem 3.2; §6 proof; §2.1, Assumption A1] Theorem 3.2 assumes F ∈ C^5, and the proof uses fourth derivatives of N (e.g., the σ⁴N''''(x) term in (3.1)). The standing assumption A1 in §2.1 only states F ∈ C^3 in x. The paper notes that A1 is stronger than needed for existence, but does not reconcile it with the C^5 requirement of Theorem 3.2. This is not fatal for the polynomial examples, which are C^∞, but the assumption sets should be harmonized so that the theorem is stated under a transparently sufficient smoothness condition.
- [§4.2, Eqs. (4.5), (4.6); §4.3, Eq. (4.13)] Theorem 3.2 is derived for a fixed splitting (2.5), but the recommended schemes are adaptive: the center c is chosen per observation and per region (e.g., (4.5), (4.13)). For such adaptive splittings, the one-step transition is not governed by a single Φ_h^[S], and the h^3 error formulas do not directly apply to the overall adaptive kernel. The paper applies the local error functions at each fixed c but does not analyze the effect of switching centers or the approximation error of the adaptive transition density. This is a particular gap for the FitzHugh–Nagumo recommendation, which relies entirely on the empirical one-step comparisons of Fig. 6–7 rather than on a theoretical error bound.
minor comments (5)
- [Throughout] There are minor typographical issues: 'efficient' (p.1), inconsistent notation 'MoA V'/'MoL V' (p.10), and 'but is should be noted' (p.10). A proofreading pass is needed.
- [§4.2, Eq. (4.10)] The formula for the analytical maximizer of σ² in (4.10) is stated without derivation. Since it is used as a practical advantage of fixed-point linearization, a brief derivation or reference would improve the paper.
- [§4.2, Figure 3] Figures 3a and 3b are informative but the color scheme and the many dots make them hard to read in print. The caption should state which parameters are used and how the 'minimizers' are computed.
- [§4.3] The statement 'When γ > 1, A is invertible for all c' is not immediately obvious from (4.11) because the matrix depends on c1 and ε; a one-line explanation would help.
- [§2.2] The adaptive pseudo-likelihood in (2.7) changes the objective with θ when the splitting depends on θ. The paper does not discuss whether the asymptotic results of Pilipovic et al. [2024] extend to this adaptive setting; a sentence acknowledging this would calibrate the reader's expectations.
Circularity Check
Theorem 3.2 is a non-circular expansion; the practical optimal-splitting claims are partially oracle-based because the criteria and simulations use the true parameter that is the target of estimation.
specific steps
-
self definitional
[Section 4.2, eqs. (4.8)–(4.9), (4.5), (4.7), Table 1 and Fig. 4; also Section 4.3, eq. (4.13)]
"Minimizer of Average Bias (MoAB) and Minimizer of Average Variance (MoAV) defined, respectively, as ci 2 arg min c2R\{±1/sqrt(3)} Eπ(jBθ(X, c)j j X 2 Ai) ...; True parameter values used in the simulation studies."
The optimality criteria and the recommended splittings are functions of the true θ: Bθ and ΔVθ are the h^3 error coefficients at the true drift, π in (4.7) is the true invariant density, and the centers x−, x†, x+ in (4.5) are roots of the drift using the true y (similarly c2=α in (4.13) for FitzHugh-Nagumo). The simulation study then generates data from exactly those true parameters and evaluates the splittings built from them. Thus the practical claim that fixed-point/slow-manifold splittings yield accurate estimates is established only for an oracle that already knows the quantity being estimated. No data-driven rule is given for choosing c from data, and re-computing c during optimization would make (2.7) a different objective not covered by the fixed-splitting theory. This makes the o
full rationale
The central mathematical contribution, Theorem 3.2, is self-contained: it is a Taylor expansion of the Strang flow (Lemma 6.1 proved in the paper, Lemma 3.1 from an external source) compared with the true moment expansion, and no fitted quantity is renamed as a prediction. The paper's citations to Pilipovic et al. [2024] set up the estimator and lower-order expansion; although Ditlevsen is a co-author, these are mathematical results and the paper proves its own extension, so this is not load-bearing circularity. The real issue is in Section 4: the error measures Bθ, ΔVθ and the invariant density π are evaluated at the true θ, and the recommended splittings use true fixed points and true α. The simulations then use those same true values to construct the splittings and to simulate the data. Therefore the practical 'optimal splitting' recommendations are demonstrated in an oracle setting, and the paper does not specify a data-driven fixed-point rule or analyze the pseudo-likelihood when the splitting changes with θ. This is a partial self-referentiality of the optimality program, not of the mathematical derivation, so the score is moderate.
Axiom & Free-Parameter Ledger
free parameters (2)
- Linearization center c (and derived A, b, N) =
c ∈ R for double-well; c = (c1, c2) with c1 ∈ {−1, −1/√3, 1/√3, 1}, c2 = α for FHN
- Region partition A1, A2, A3 (double-well) at ±1/√3 =
±1/√3 (boundaries where F'(x)=0)
axioms (9)
- standard math Conditional moment expansion E[φ(X_{t+h})|X_t=x] = Σ (h^j/j!) L^j φ(x) + O(h^{n+1}) (Lemma 3.1, from Sørensen 2012, Lemma 1.10)
- domain assumption Assumptions A1–A4 (F ∈ C³, one-sided Lipschitz, polynomial growth, ΣΣᵀ > 0) guarantee a unique strong solution; the model is well-posed
- domain assumption Expansion of the deterministic flow f_h (Lemma 6.1, extending Pilipovic et al. [2024, Prop. 2.2]) to order h³
- domain assumption The pseudo-likelihood (2.6)–(2.7) is the right objective: the last term (log |det Df^{-1}|) correctly accounts for the nonlinear transformation
- ad hoc to paper A splitting that approximates transition densities well also yields accurate parameter estimates
- domain assumption EM-simulated distributions with very small time steps (h_sim = 0.00004–0.0001) and 2000–5000 repetitions represent the true transition densities
- ad hoc to paper True parameters are available for computing the splittings in the simulation studies (fixed points, minimizers of error measures)
- standard math For the double-well, the process is ergodic with invariant density π (Kutoyants 2004, Thm 1.15) so time averages of |B| and |ΔV| converge to E_π
- standard math Invariant density π(x) ∝ exp(−2U(x)/σ²) for gradient SDEs
read the original abstract
The Strang splitting estimator is a powerful estimator for parametric inference in multivariate stochastic differential equation models with nonlinear drift and additive noise. While the choice of splitting does not affect the asymptotic distribution of the estimator, it makes a huge impact in finite-sample settings and it has not yet been shown how the splitting can be chosen optimally. We derive error measures for the transition densities of the Strang splitting scheme, in particular calculating the bias up to the order of $h^3$, where $h$ is the length of the time step. We study the connection between these error measures and the performance of the Strang splitting estimator in the double-well potential model and the stochastic FitzHugh-Nagumo model, respectively. Our simulation studies suggest that linearization around fixed points yields accurate parameter estimates for potential models, while other splittings perform better for slow-fast excitable models.
Figures
Reference graph
Works this paper leans on
-
[1]
Parameter Estimation in Nonlinear Multivariate Stochastic Differential Equations Based on Splitting Schemes
Predrag Pilipovic and Adeline Samson and Susanne Ditlevsen. Parameter Estimation in Nonlinear Multivariate Stochastic Differential Equations Based on Splitting Schemes. The Annals of Statistics. 2024
2024
-
[2]
2023 , publisher=
Stochastic Differential Equations for Science and Engineering , author=. 2023 , publisher=
2023
-
[3]
Kutoyants
Yury A. Kutoyants. Statistical Inference for Ergodic Diffusion Processes. 2004
2004
-
[4]
A splitting method for SDE s with locally L ipschitz drift: I llustration on the F itz H ugh- N agumo model
Evelyn Buckwar and Adeline Samson and Massimiliano Tamborrino and Irene Tubikanec. A splitting method for SDE s with locally L ipschitz drift: I llustration on the F itz H ugh- N agumo model. Applied Numerical Mathematics. 2022
2022
-
[5]
Applied Stochastic Differential Equations
Simo Särkkä and Arno Solin. Applied Stochastic Differential Equations. 2019
2019
-
[6]
1961 , author =
Biophysical Journal , volume =. 1961 , author =
1961
-
[7]
and Arimoto, S
Nagumo, J. and Arimoto, S. and Yoshizawa, S. , journal=. 1962 , volume=
1962
-
[8]
Tellus , volume=
Stochastic resonance in climatic change , author=. Tellus , volume=
-
[9]
SIAM Journal on Applied Mathematics , volume=
A theory of stochastic resonance in climatic change , author=. SIAM Journal on Applied Mathematics , volume=. 1983 , publisher =
1983
-
[10]
Six decades of the FitzHugh–Nagumo model: A guide through its spatio-temporal dynamics and influence across disciplines
Daniel Cebrián-Lacasa and Pedro Parra-Rivas and Daniel Ruiz-Reynés and Lendert Gelens. Six decades of the FitzHugh–Nagumo model: A guide through its spatio-temporal dynamics and influence across disciplines. Physics Reports. 2024
2024
-
[11]
2012 , month =
Sørensen, Michael , editor =. 2012 , month =
2012
-
[12]
Nature Communications , Year =
Ditlevsen, Peter and Ditlevsen, Susanne , Title =. Nature Communications , Year =. doi:10.1038/s41467-023-39810-w , Article-Number =
-
[13]
Strang splitting for parametric inference in second-order stochastic differential equations , journal =. 2025 , issn =. doi:https://doi.org/10.1016/j.spa.2025.104650 , url =
arXiv 2025
-
[14]
2026 , doi =
Ditlevsen, Susanne and Rahbek, Anders and Damsholt, Gabriel and Ditlevsen, Peter , title =. 2026 , doi =
2026
-
[15]
Jensen and Susanne Ditlevsen and Mathieu Kessler and Omiros Papaspiliopoulos , title =
Anders C. Jensen and Susanne Ditlevsen and Mathieu Kessler and Omiros Papaspiliopoulos , title =. Physical Review E. 2012
2012
-
[16]
Strogatz
Steven H. Strogatz. Nonlinear Dynamics and Chaos: With Applications to Physics, Chemistry, and Engineering. 2018
2018
-
[17]
Noise-Induced Phenomena in Slow-Fast Dynamical Systems
Nils Berglund and Barbara Gentz. Noise-Induced Phenomena in Slow-Fast Dynamical Systems. 2006
2006
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.