REVIEW 2 major objections 6 minor 41 references
Learning with Expected Signatures: Theory and Applications
T0 review · 2 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper proves the expected signature of a latent continuous-time process is consistently estimated by averaging signatures of discretely observed paths, with asymptotic normality and a martingale variance-reduction modification.
desk verdict Unified expected-signature asymptotics and a useful martingale control variate, but the proof of Theorem 2.8 has a genuine gap and the abstract overstates the experiments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the expected signature $\phi_I(T)=\mathbb{E}[S_I(X)[0,T]]$, where $S_I(X)[0,T]$ is an entry of the signature, the sequence of iterated integrals of the path. The argument is carried by the error decomposition separating an in-fill discretization term from a long-span statistical term; the first is controlled by the $L^m$ in-fill convergence of Theorem 2.8, proved with a Cauchy-sequence argument and moment assumptions on increments, and the second by Birkhoff's ergodic theorem or a dependent central limit theorem. The martingale correction machinery is the Itô integral in place of the outermost Stratonovich integral, which yields a mean-zero control variate whose optimal coefficient is the slope of a simple linear regression of the signature term on the control.
What would settle it
Take a stationary ergodic process, fix the observation partition mesh at $h>0$, and let $N$ grow without further refinement: the estimator will converge to $\mathbb{E}[S_I(X_\pi)[0,T]]$, which differs from $\mathbb{E}[S_I(X)[0,T]]$ by the unvanished bias term in decomposition (8), and the discrepancy can be measured against a fine-mesh Monte Carlo benchmark. For the normality claim, use chopped windows with non-summable mixing coefficients, such as fractional Brownian motion under the chop scheme, and check that confidence intervals built from the Corollary 2.12 covariance estimator lose coverage.
Extended reading notes
Core claim
On its own terms, the paper establishes that for a canonical geometric stochastic process $X$, the estimator $\hat\phi^{\Pi(N)}_I(T)=N^{-1}\sum_{n=1}^N S_I(X_{n,\pi_{N,n}})[0,T]$ is $L^2$-consistent for the continuous-time expected signature $\phi_I(T)=\mathbb{E}[S_I(X)[0,T]]$ under stationarity and ergodicity, and becomes $\sqrt{N}$-asymptotically normal with long-run covariance $\Sigma_I$ when the chopped sequence $\{X_n\}$ is strongly mixing and the partition refinement is fast enough. The regularity conditions in Assumption 2.6 and the summable refinement schedule make the discretization bias in the error decomposition vanish in $L^m$, while the long-span term is handled by an ergodic theorem or a mixing central limit theorem. The paper further claims that when $X$ is a square-integrable martingale, replacing the outermost Stratonovich integral in a signature word with an Itô integral creates a mean-zero control variate $S^c_I$, and subtracting the optimally scaled control lowers the estimator variance by the factor $1-\rho^2_{I,\pi}$ without changing its bias.
Load-bearing premise
The central claim collapses if the observed discrete paths are not refinement-based discretizations of a latent continuous-time process satisfying the stated moment and fast-refinement conditions, because then the in-fill bias term in the error decomposition does not vanish and the sample average converges to the expected signature of the discretized paths instead of the continuous-time expected signature.
Editorial extensions
If this is right
- Signature-based machine learning features computed from discrete time series can be interpreted as estimates of the expected signature of a latent continuous-time process, not merely as ad hoc empirical averages.
- Consistency holds with irregular and sample-dependent observation partitions and with dependent, chopped observations, so the estimator applies to long single recordings by splitting them into windows.
- The feasible long-run covariance estimator of Corollary 2.12 allows confidence intervals for expected signature terms, enabling calibrated uncertainty quantification in downstream predictions.
- For martingale data, the control-variate estimator has the same bias as the naive estimator and variance reduced by $1-\rho^2_{I,\pi}$; the reported experiments show lower mean squared error and improved predictive accuracy.
- For Gaussian processes with decaying increment covariances, consistency holds without the strong mixing assumption, extending the results to processes whose increments are not mixing.
Reading between the lines
- The paper leaves implicit that the fast-refinement schedule is not merely technical: algorithms that keep the observation grid fixed while increasing the number of samples will converge to the expected signature of the discretized path, not of the continuous-time process, so reporting partition schedules should matter in experimental practice.
- The martingale correction can be read as a drift-removal operation on signature features; for non-martingale data the optimal coefficient trades bias against variance, and a cross-validated shrinkage version of the correction is a natural testable extension.
- The Gaussian consistency theorem suggests a practical diagnostic: compute expected signature estimates on non-overlapping windows of increasing length and check stability, giving an empirical probe of whether the latent-process and covariance-decay assumptions hold.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper establishes asymptotic theory for the empirical expected signature estimator in a double in-fill / long-span regime: Theorem 2.8 gives L^m convergence of the discretized signature to the continuous-time signature under regularity conditions on the latent process, Theorem 2.10 gives L^2 consistency and asymptotic normality under stationarity/ergodicity and strong mixing, and Theorem 2.14 gives a Gaussian consistency result under covariance decay. The second half proposes a control-variate modification of the estimator that exploits a martingale property to reduce variance, and it reports numerical experiments on GPES, signature pricing/hedging, SES, and controlled linear regression. The paper is clearly written, carefully links the assumptions to concrete processes (Brownian motion, fractional Brownian motion, CAR, Heston), and provides code for the experiments.
Significance. If the central proof gap is repaired, the paper would be a valuable contribution: it unifies and extends earlier in-fill and long-span results for expected signatures, gives checkable conditions for common continuous-time models, and introduces a simple, practical martingale correction with an oracle variance-reduction guarantee. The explicit treatment of dependent samples and irregular partitions is a genuine step beyond prior work. The manuscript also ships reproducible code and is unusually candid about limitations (e.g., the bias of the correction for non-martingales and the absence of significance in some experiments). These strengths are substantial, but the current proof of the main in-fill theorem is incomplete, and the load-bearing gap must be resolved before the theoretical claims can be accepted.
major comments (2)
- [Appendix B.1.1, Eq. (24)] The proof of Theorem 2.8 under cases (iii) and (iv) applies Lemma B.1 to the family G_[s,t] = F_s ∨ σ(X_{v,w}, [v,w]∈π_n, [t,τ]). This family is not increasing in the interval order, and the asserted measurability of Z^I_[v,w] with respect to G_[s,t] fails when i3>0: the factor S_{i3}(X_{π_n})[w,τ1] contains increments on [w,t] that are in neither F_s nor the tail σ-algebra generated from t onward. The failure already occurs at level k'+1=3 with (i1,i2,i3)=(0,2,1). Consequently the BDG-based bound (24) is not established by Lemma B.1, and the claimed L^m in-fill convergence rate of Theorem 2.8 is not proved as written. Since Theorem 2.10 and Theorem 2.14 both rely on Theorem 2.8, this gap is load-bearing. The result may still be true and the argument repairable, for example via a genuinely two-sided stochastic sewing lemma, but the current proof does not supply that argument.
- [Section 2.2 and Appendix C] The oracle variance-reduction formula Var(phi_hat^{c*}) = (1-ρ^2) Var(phi_hat) is stated for the infeasible optimal coefficient c*_π, but in the experiments and in the ML pipelines the coefficient is estimated from the same data (e.g., \(c\)hat\(c*_{π,1}\) in Section C.2). No theorem is given for the feasible estimator; in particular, the additional variability of the estimated coefficient is not analyzed. Thus the abstract's claim of 'significantly lower mean squared error' is only partially supported by the theory, and the empirical confirmation in Table 1 is mixed (the FBM row is not significant, t-stat 1.49, p=0.15). The paper should either add a theoretical analysis of the feasible estimator or explicitly restrict the variance-reduction claim to the oracle setting and tone down the abstract.
minor comments (6)
- [Section 3.2.1, Table 1] The caption and text state that the martingale correction significantly improves performance, but on the FBM dataset the t-statistic is 1.49 and the p-value is 0.15, so the improvement is not statistically significant there. The claim should be qualified to the datasets where the test is significant.
- [Appendix C.3] There is a typo in the proof of Lemma C.3: 'traingle inequality' should be 'triangle inequality'.
- [Section 3.1, CAR example] The phrase 'CARA bidimensional Continuous-time Autoregressive (CAR) process' appears to contain a typo; it should likely be 'A bidimensional Continuous-time Autoregressive (CAR) process'.
- [Corollary 2.12] The consistency of the kernel estimator is made conditional on an assumed rate ρ(N)∼N^{-υ} that is not derived from Theorem 2.10. This should be stated as an additional condition in the corollary statement, not introduced as an assumption in the proof sketch.
- [Appendix F.2.1] Algorithm 1 applies the martingale correction inside the GPES model, where the target is a conditional expected signature given X_{π1}=x. As the paper notes, the correction can introduce bias in that setting; this point is important and should be repeated in the main text near the experiments, not only in the appendix.
- [Notation in Definition 2.5 and Theorem 2.8] The notation S_I(X_π)[s,t] in Definition 2.5 is used for the signature of the linear interpolation X_π restricted to [s,t], but the dependence on the partition π and the interval [s,t] is sometimes ambiguous (e.g., in the proof of Theorem 2.8 the 'abuse of notation' is acknowledged). A short notational clarification in the main text would improve readability.
Circularity Check
No significant circularity: the central convergence theorems are proved from stated regularity and mixing assumptions, and the control-variate estimator is a standard plug-in variance-reduction device rather than a fitted quantity passed off as a prediction.
full rationale
The paper's central derivation chain is not circular. Theorem 2.8 takes as input a pointwise 'signature-defining' convergence (Definition 2.5) and upgrades it to L^m convergence under Assumption 2.6; the conclusion is strictly stronger than the input and is not assumed. The decomposition in Equation (8) is algebraic, and Theorem 2.10 controls the in-fill term through Theorem 2.8 plus the partition refinement condition (11), while the long-span term is handled by Birkhoff's ergodic theorem and the Ibragimov CLT. No parameter is fitted to a subset of data and then renamed a prediction. The martingale correction's coefficient c* is derived as Cov(S_I, S_c)/Var(S_c), an exact variance-minimizing expression for a control variate, and its sample version is a standard plug-in estimator rather than a quantity calibrated to the target. Self-citations such as Lucchese, Pakkanen and Veraart (2023) are used only to verify that a CAR example satisfies the assumptions, not as load-bearing support for the main theorems; no 'uniqueness' result from the authors' prior work is imported to force a choice. The skeptical note about the measurability hypothesis of Lemma B.1 in Appendix B.1 is a potential proof gap, but a proof gap is not circularity: it does not make the theorem's conclusion equivalent to its inputs by construction. Overall, the derivation is self-contained relative to external probabilistic results, and the empirical sections benchmark against existing methods rather than claiming a prediction that reduces to a fit.
Assumptions & free parameters
assumptions (7)
- standard math Rough path signature extension and continuity (Lyons et al. 2007, Theorems 3.7 and 3.10).
- domain assumption Semimartingales admit Stratonovich geometric rough path lifts and linear interpolations converge in p-variation.
- domain assumption Fractional Brownian motion with Hurst parameter H > 1/4 admits a canonical geometric lift via dyadic partitions.
- standard math Ibragimov's dependent central limit theorem and Birkhoff's ergodic theorem apply to stationary strongly mixing sequences with the stated mixing condition.
- standard math Isserlis theorem for moments of Gaussian random vectors.
- standard math Burkholder-Davis-Gundy inequality and the martingale estimate behind Lemma B.1.
- domain assumption Stationarity, ergodicity, and strong mixing of CAR and Heston processes under stated parameter conditions.
Cite this review
Pith. "Pith review of Learning with Expected Signatures: Theory and Applications." pith.science (2026). https://pith.science/paper/WCEQ5AXW
@misc{pith2026250520465,
author = {Pith},
title = {Pith review of: Learning with Expected Signatures: Theory and Applications},
year = {2026},
howpublished = {\url{https://pith.science/paper/WCEQ5AXW}},
note = {Machine review of arXiv:2505.20465}
}
read the original abstract
The expected signature maps a collection of data streams to a lower dimensional representation, with a remarkable property: the resulting feature tensor can fully characterize the data generating distribution. This "model-free" embedding has been successfully leveraged to build multiple domain-agnostic machine learning (ML) algorithms for time series and sequential data. The convergence results proved in this paper bridge the gap between the expected signature's empirical discrete-time estimator and its theoretical continuous-time value, allowing for a more complete probabilistic interpretation of expected signature-based ML methods. Moreover, when the data generating process is a martingale, we suggest a simple modification of the expected signature estimator with significantly lower mean squared error and empirically demonstrate how it can be effectively applied to improve predictive performance.
Figures
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Buehler, H., Murray, P., Pakkanen, M. S., and Wood, B. Deep Hedging: Learning to Remove the Drift under Trading Frictions with Minimal Equivalent Near-Martingale Measures , 2022. URL https://arxiv.org/abs/2111.07844
arXiv 2022
-
[3]
Burkholder, D. L., Davis, B., and Gundy, R. F. Integral inequalities for convex functions of operators on martingales. In Le Cam, L. M., Neyman, J., and Scott, E. L. (eds.), Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability, volume 2, pp.\ 223--240, Berkeley, California, 1972. University of California Press. URL https:...
arXiv 1972
-
[4]
Iterated Integrals and Exponential Homomorphisms
Chen, K.-T. Iterated Integrals and Exponential Homomorphisms . Proceedings of the London Mathematical Society, s3--4 0 (1): 0 502--512, 1954. doi:10.1112/plms/s3-4.1.502
-
[5]
Chevyrev, I. and Lyons, T. Characteristic functions of measures on geometric rough paths . The Annals of Probability, 44 0 (6): 0 4049--4082, 2016. ISSN 00911798
work page 2016
-
[6]
Chevyrev, I. and Oberhauser, H. Signature moments to characterize laws of stochastic processes , 2018. URL https://arxiv.org/abs/1810.10971
arXiv 2018
-
[7]
Coutin, L. and Qian, Z. Stochastic analysis, rough path analysis and fractional Brownian motions . Probability Theory and Related Fields, 122: 0 108--140, 01 2002. doi:10.1007/s004400100158
-
[8]
Signature SDEs from an affine and polynomial perspective , 2023
Cuchiero, C., Svaluto-Ferro, S., and Teichmann, J. Signature SDEs from an affine and polynomial perspective , 2023. URL https://arxiv.org/abs/2302.01362
arXiv 2023
Show all 41 references
-
[9]
Dragomir, S. S. Some Gronwall Type Inequalities and Applications . Nova Science, New York, 2003
2003
-
[10]
Problems in stochastic analysis
Fawcett, T. Problems in stochastic analysis. Connections between rough paths and non-commutative harmonic analysis . PhD thesis, University of Oxford, 2003
2003
-
[11]
Partial Differential Equations of Parabolic Type
Friedman, A. Partial Differential Equations of Parabolic Type . Prentice-Hall, 1964
1964
-
[12]
Friz, P. K. and Victoir, N. B. Multidimensional Stochastic Processes as Rough Paths: Theory and Applications . Cambridge Studies in Advanced Mathematics. Cambridge University Press, 2010
2010
-
[13]
K., Hager, P
Friz, P. K., Hager, P. P., and Tapia, N. Unified signature cumulants and generalized Magnus expansions . Forum of Mathematics, Sigma, 10: 0 e42, 2022. doi:10.1017/fms.2022.20
2022 doi
-
[14]
K., Hager, P
Friz, P. K., Hager, P. P., and Tapia, N. On expected signatures and signature cumulants in semimartingale models , 2024. URL https://arxiv.org/abs/2408.05085
2024 arXiv
-
[15]
Signature Trading: A Path-Dependent Extension of the Mean-Variance Framework with Exogenous Signals , 2023
Futter, O., Horvath, B., and Wiese, M. Signature Trading: A Path-Dependent Extension of the Mean-Variance Framework with Exogenous Signals , 2023. URL https://arxiv.org/abs/2308.15135
2023 arXiv
-
[16]
Sparse arrays of signatures for online character recognition , 2013
Graham, B. Sparse arrays of signatures for online character recognition , 2013. URL https://arxiv.org/abs/1308.0371
2013 arXiv
-
[17]
and Lyons, T
Hambly, B. and Lyons, T. Uniqueness for the signature of a path of bounded variation and the reduced path group . Annals of Mathematics, 171 0 (1): 0 109--167, 2005
2005
-
[18]
Ibragimov, I. A. Some Limit Theorems for Stationary Processes . Theory of Probability & Its Applications, 7 0 (4): 0 349--382, 1962. doi:10.1137/1107036
1962 doi
-
[19]
On a Formula for the Product-Moment Coefficient of any Order of a Normal Frequency Distribution in any Number of Variables
Isserlis, L. On a Formula for the Product-Moment Coefficient of any Order of a Normal Frequency Distribution in any Number of Variables . Biometrika, 12 0 (1--2): 0 134--139, 11 1918. ISSN 0006-3444. doi:10.1093/biomet/12.1-2.134
1918 doi
-
[20]
and Shiryaev, A
Jacod, J. and Shiryaev, A. N. Limit Theorems for Stochastic Processes , volume 288 of Grundlehren der mathematischen Wissenschaften. Springer Berlin, Heidelberg, 1987. ISBN 9783662025161. doi:10.1007/978-3-662-02514-7
1987 doi
-
[21]
Foundations of Modern Probability
Kallenberg, O. Foundations of Modern Probability . Probability theory and stochastic modelling. Springer, 2021. ISBN 9783030618728
2021
-
[22]
Kiraly, F. J. and Oberhauser, H. Kernels for Sequentially Ordered Data . Journal of Machine Learning Research, 20 0 (31): 0 1--45, 2019
2019
-
[23]
Ergodic Behavior of Markov Processes
Kulik, A. Ergodic Behavior of Markov Processes . De Gruyter, Berlin, Boston, 2018. ISBN 9783110458930. doi:10.1515/9783110458930
2018 doi
-
[24]
A stochastic sewing lemma and applications
L \^e , K. A stochastic sewing lemma and applications . Electronic Journal of Probability, 25: 0 1 -- 55, 2020. doi:10.1214/20-EJP442
2020 doi
-
[25]
V., and Lyons, T
Lemercier, M., Salvi, C., Damoulas, T., Bonilla, E. V., and Lyons, T. Distribution regression for sequential data . In 24th International Conference on Artificial Intelligence and Statistics (AISTATS 2021), Proceedings of Machine Learning Research, pp.\ 3754--3762. Journal of ...
2021
-
[26]
Learning from the past, predicting the statistics for the future, learning an evolving system , 2016
Levin, D., Lyons, T., and Ni, H. Learning from the past, predicting the statistics for the future, learning an evolving system , 2016. URL https://arxiv.org/abs/1309.0260
2016 arXiv
-
[27]
S., and Veraart, A
Lucchese, L., Pakkanen, M. S., and Veraart, A. E. D. Estimation and Inference for Multivariate Continuous-time Autoregressive Processes , 2023. URL https://arxiv.org/abs/2307.13020
2023 arXiv
-
[28]
and McLeod, A
Lyons, T. and McLeod, A. D. Signature Methods in Machine Learning , 2024. URL https://arxiv.org/abs/2206.14674
2024 arXiv
-
[29]
Differential Equations Driven by Rough Paths: \'E cole D' \'e t \'e de Probabilit \'e s de Saint-Flour XXXIV-2004
Lyons, T., Caruana, M., and L \'e vy, T. Differential Equations Driven by Rough Paths: \'E cole D' \'e t \'e de Probabilit \'e s de Saint-Flour XXXIV-2004 . Number no. 1908 in Differential Equations Driven by Rough Paths: \'E cole D' \'e t \'e de Probabilit \'e s de Saint-Flou...
2004
-
[30]
Non-parametric pricing and hedging of exotic derivatives
Lyons, T., Nejad, S., and P\' e rez Arribas, I. Non-parametric pricing and hedging of exotic derivatives . Applied Mathematical Finance, 27 0 (6): 0 457--494, 2021
2021
-
[31]
Mandelbrot, B. B. and Van Ness, J. W. Fractional Brownian Motions, Fractional Noises and Applications . SIAM Review, 10 0 (4): 0 422--437, 1968. ISSN 00361445. URL http://www.jstor.org/stable/2027184
1968
-
[32]
and Stelzer, R
Marquardt, T. and Stelzer, R. Multivariate CARMA processes . Stochastic Processes and their Applications, 117 0 (1): 0 96--120, Jan 2007. ISSN 03044149. doi:10.1016/j.spa.2006.05.014
2007 doi
-
[33]
Newey, W. K. and West, K. D. A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix . Econometrica, 55 0 (3): 0 703--708, 1987. ISSN 00129682, 14680262
1987
-
[34]
The expected signature of a stochastic process
Ni, H. The expected signature of a stochastic process . PhD thesis, University of Oxford, 2012
2012
-
[35]
On the signature and cubature of the fractional Brownian motion for H>1/2
Passeggeri, R. On the signature and cubature of the fractional Brownian motion for H>1/2 . Stochastic Processes and their Applications, 130 0 (3): 0 1226--1257, 2020. ISSN 0304-4149. doi:10.1016/j.spa.2019.04.013
2020 doi
-
[36]
M., Geddes, J
P\' e rez Arribas, I., Goodwin, G. M., Geddes, J. R., Lyons, T., and Saunders, K. E. A. A signature-based machine learning model for distinguishing bipolar disorder and borderline personality disorder . Translational Psychiatry, 8, 2018
2018
-
[37]
The Signature Kernel Is the Solution of a Goursat PDE
Salvi, C., Cass, T., Foster, J., Lyons, T., and Yang, W. The Signature Kernel Is the Solution of a Goursat PDE . SIAM Journal on Mathematics of Data Science, 3 0 (3): 0 873--899, 2021. doi:10.1137/20M1366794
2021 doi
-
[38]
and Oberhauser, H
Schell, A. and Oberhauser, H. Nonlinear independent component analysis for discrete-time and continuous-time signals . Annals of Statistics, 51 0 (2): 0 487--518, 2023
2023
-
[39]
and Romito, M
Triggiano, F. and Romito, M. Gaussian Processes Based Data Augmentation and Expected Signature for Time Series Classification . IEEE Access, 12: 0 80884--80895, 2024. doi:10.1109/ACCESS.2024.3408712
2024
-
[40]
Willett, D. W. Nonlinear vector integral equations as contraction mappings . Archive for Rational Mechanics and Analysis, 15: 0 79--86, 1964. doi:10.1007/bf00257405
1964 doi
-
[41]
Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text Recognition
Xie, Z., Sun, Z., Jin, L., Ni, H., and Lyons, T. Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text Recognition . IEEE Transactions on Pattern Analysis and Machine Intelligence, 40 0 (8): 0 1903--1917, 2018. doi:10....
1903
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.