Pith. sign in

REVIEW 3 major objections 4 minor 50 references

Truncated signatures learn smooth path functionals at rate K^{-2γ}.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 16:46 UTC pith:CVKAQ7KA

load-bearing objection The minimax rate theorem is a real contribution and the paper should go to review, but the OLS consistency proof has a specific gap: the covariance concentration bound ignores exponential K-dependence in the signature moments. the 3 major comments →

arxiv 2607.17865 v1 pith:CVKAQ7KA submitted 2026-07-20 math.ST stat.MEstat.MLstat.TH

How Fast Do Signatures Learn? Statistical Theory and Applications for Path Regression

classification math.ST stat.MEstat.MLstat.TH MSC 62G0862J0762F1260H10
keywords path regressionpath signaturetruncated signatureItô diffusionminimax approximation rateuniversal approximationsignature LASSOsignature logistic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper's central aim is to replace the qualitative universal approximation theorem for path signatures with a quantitative rate. For smooth functionals built from coefficients h of an Itô diffusion X, the best level-K signature approximation has squared L2 error of order K^{-2γ}, where γ is the regularity of h, and this rate is minimax optimal over the class. That single rate then drives statistical guarantees: the paper proves consistency for Signature-OLS, sparse recovery for Signature-LASSO, and latent-score consistency for Signature-Logistic, subject to balancing the feature dimension d_K, sample size n, and the decay exponent Q. The payoff is practical: truncating a path at level K becomes a principled bias-variance choice, and the paper demonstrates it on foreign-exchange volatility forecasting, battery end-of-life prediction, and EEG seizure detection.

Core claim

The paper proves that for (X,h) in the smooth diffusion-functional class U_{γ,R}, the projection residual E|ξ^X_K(Y)|^2 is bounded by C H_γ(h)^2 K^{-2γ}, and the supremum over the class is of exact order K^{-2γ}. Smoother coefficient functions yield faster convergence of the truncated signature approximation, and the exponent 2γ cannot be improved uniformly. It then shows that, under a uniform spectral-gap condition on the signature covariance matrix and a polynomial decay bound on the truncation residual, Signature-OLS is consistent and asymptotically normal, Signature-LASSO recovers the active signature coordinates with high probability, and Signature-Logistic estimates the latent logit co

What carries the argument

The central object is the time-augmented path signature: the ordered hierarchy of iterated integrals of cX_t = (t, X_t). Theorem 1 is built on a localization argument that confines the diffusion path to a box of size K^β; Jackson polynomial approximation of the coefficient functions on that box; the shuffle identity, which converts polynomial coefficient functionals into level-K signature functionals; and sub-Gaussian tail control for the event that the path leaves the box. For Brownian motion, a signature-chaos lemma shows that the level-K signature space is exactly the polynomial part of the first K Wiener chaos kernels, which is why the minimax lower bound is sharp. The same rate enters t

Load-bearing premise

Every consistency result assumes the covariance matrix of the level-K signature features keeps its smallest eigenvalue bounded away from zero for all K, and that the truncation residual decays at a polynomial rate; the authors themselves note that the full signature dictionary grows exponentially, so these conditions are only plausible for analytic targets or a pre-selected feature subset whose spectral properties are not demonstrated.

What would settle it

Compute the smallest eigenvalue of the sample covariance matrix of the level-K signature features used in the battery or EEG application for K = 1,2,3,4; if λ_min is not bounded below by a positive constant as K grows, Assumption 1(i) fails and the Signature-OLS and Signature-Logistic consistency theorems do not govern those fitted models.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Truncation level K can now be treated as a smoothing parameter: for a target of regularity γ, squared bias decays like K^{-2γ}, so the optimal K in finite samples balances this bias against the d_K/n variance term.
  • Signature-OLS is consistent and asymptotically normal when d_K^2/n → 0 and K^Q/d_K^2 → ∞; equivalently, the truncation bias must be negligible relative to estimation error.
  • Signature-LASSO can achieve support recovery even when the full signature dictionary is much larger than n, provided the active set stays small and the irrepresentability condition holds; the exponent Q drops out of the selection rate.
  • Signature-Logistic gives consistent latent-score estimates under d_K^2/n → 0 and d_K K^{-Q} → 0, so binary classification with path covariates inherits the same bias-variance logic.
  • For analytic coefficient functions, the truncation error becomes exponential, which makes a logarithmic choice of K (roughly log n / (2 log ρ)) theoretically optimal.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the approximation rate extends to barrier, stopping-time, and occupation-time functionals—listed by the authors as open—the same consistency framework would transfer to simulation-based pricing and optimal stopping, where the truncation level would play the role of a basis dimension with known error decay.
  • A testable consequence of the exponential-rate result is that for very smooth path functionals a logarithmic truncation schedule should outperform deeper signatures; a cross-validated comparison on synthetic analytic targets would settle how tight the constant ρ is.
  • The paper's own remark after the OLS theorem shows that the full signature dictionary cannot satisfy the polynomial-rate conditions, so the applied value of the theory hinges on the spectral gap of a selected feature subset; checking λ_min for the reported 155- and 858-feature dictionaries would directly test whether the theorems govern those implementations.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper develops a quantitative theory of truncated signatures for path regression. Its central approximation result, Theorem 1, states that for a class of smooth functionals of Itô diffusions — functionals built from time integrals and Stratonovich integrals with coefficient functions of smoothness γ — the squared L2 truncation error of the level-K signature is O(K^{-2γ}), and that this rate is minimax sharp over the class U_{γ,R}. The paper then propagates a generic residual rate E|ξ|^2=O(K^{-Q}) through three estimators: Signature-OLS, Signature-LASSO, and Signature-Logistic, giving consistency and, for OLS, asymptotic normality. Simulations illustrate the K^{-2γ} ordering and the approximation–estimation tradeoff. Three applications (FX volatility, battery end-of-life, EEG seizure detection) compare signature features with handcrafted benchmarks.

Significance. If correct, the paper supplies a genuinely useful quantitative complement to the qualitative signature universal approximation theorem, and the explicit residual-propagation framework is a step forward for the statistical signature literature. The paper is also well served by its technical apparatus: Lemma A.4 gives a concrete, admissible sub-Weibull constant, Theorem 1 is proved with a localization-plus-Jackson argument, and the appendices contain full proofs and substantial simulation/application detail. The empirical sections honestly report overfitting (notably Signature-OLS in the battery application), which strengthens credibility. The central approximation theorem appears sound. However, I find the skeptic's concern about Lemma A.3(i) valid and load-bearing: the covariance-concentration step that underpins the OLS and logistic theorems ignores the K-dependence of the signature-coordinate tail constants, so the stated sufficient conditions do not prove the claimed statistical rates. This is repairable, but it is more than a local edit.

major comments (3)
  1. [Appendix B.1, Lemma A.3(i); Theorem 2] The proof asserts E‖s_i‖^4 = O(d_K^2) from "bounded fourth moment" of each coordinate. This is inconsistent with the paper's own Lemma A.4, where a level-k coordinate satisfies ‖S^I−ES^I‖_{ψ_{2/k}} ≤ C M_ψ^k with M_ψ ≥ 32e > 1. Consequently E[(S^I)^4] grows like M_ψ^{4k}, and summing over the d^k words of level k gives E‖s_K‖^4 = O((d M_ψ^4)^K), not O(d_K^2) with K-independent constants. The Frobenius bound (A.68) therefore needs an additional factor of order M_ψ^{4K}, and consistency in Lemma A.3(i) requires a condition such as n/(M_ψ^{4K} d_K^2)→∞ (up to polynomial factors), strictly stronger than d_K^2/n→0. Theorem 2(i)–(ii) relies on this step; as written, the displayed rate conditions do not establish the OLS consistency or asymptotic normality claims.
  2. [Sec. 4.1 Assumption 1 and Remark after Theorem 2; Sec. 6.1.2, 6.2.2, 6.3.2] Assumption 1(i) postulates a uniformly well-conditioned Gram matrix for all K, but the full time-augmented signature is exactly singular because of pure-time coordinates, as the paper itself notes. The statistical theorems therefore apply only to a preselected feature subset whose spectral properties are assumed rather than proved, and the relation between the full-signature rate in Theorem 1 and the residual of the selected subset is not established. In the applications, no eigenvalue diagnostics are reported, and the feature counts/sample sizes do not satisfy the displayed theorem conditions: 858 features with n≈2000 in Sec. 6.1.2 gives d_K^2/n≈368; 155 features with n=166 in Sec. 6.2.2 gives d_K^2/n≈145; the level-2 EEG features are even larger. Thus the consistency theorems are not demonstrated for the estimators actually implemented.
  3. [Appendix B.4, Theorem 4] The proof of Theorem 4 localizes the empirical logistic objective by asserting sup_{‖L−L0K‖≤r} ‖∇²L_n(L) − H_K(L)‖ = o_P(1) "by the same concentration argument as in the least-squares case." Since the Hessian involves rank-one matrices s_i s_i^T, the same missing M_ψ^{4K} factor from Lemma A.3(i) enters here. Without an additional condition such as n/(M_ψ^{4K} d_K^2)→∞, the localization and the displayed rate in Eq. (26) are unproved. The asymmetry with Theorem 3, which explicitly tracks M_ψ^{2K} in its rate condition, underscores that this is not just a stylistic omission.
minor comments (4)
  1. [Appendix B.3, Step 2] The proof of Theorem 3 writes "∥Y∥_{L_p} ≤ C p^{K/2} for every K≥1". Since Y is the target functional and does not depend on K, this should be p^{1/2} (Y is sub-Gaussian under the bounded-coefficient assumptions). As written it is confusing and technically wrong, although the subsequent bound can absorb the correct factor.
  2. [Sec. 2.1, Eq. (1)] The iterated integral in Eq. (1) is written without specifying Itô vs Stratonovich convention, while the text says Stratonovich is used throughout. Please state the convention at the definition or immediately after it, especially because Eq. (1) is also used for Itô integrals in Appendix A.2.
  3. [Sec. 6.1.2] The formula log n/(2 log(d+1)) with d+1=3 gives the heuristic K=3 for a single day's signature, but the actual feature vector concatenates 22 daily signatures and has 858 features. The heuristic should be stated as applying to the per-day dictionary only, otherwise the reader may infer a feature dimension much smaller than 858.
  4. [Fig. 1] The horizontal axis is labeled "full signature depth K (log scale)" but the axis is discrete (2 to 9). A linear depth axis or a clear note that only integer depths are shown would be clearer.

Circularity Check

0 steps flagged

No significant circularity; the main approximation rate and the statistical consistency theorems are modular, with the truncation-rate condition supplied by an independent proof rather than fitted or assumed from self-citation.

full rationale

The central derivation is self-contained. Theorem 1 proves the L2 truncation bound E|ξ_X^K(Y)|^2 ≤ C H_γ(h)^2 K^{-2γ} for the smooth diffusion-functional class U_{γ,R} using localization, Jackson polynomial approximation, the shuffle identity, and BDG/Itô-isometry estimates; the minimax lower bound is obtained from an explicit Brownian first-chaos subclass, not from the upper bound or from any fitted quantity. The statistical theorems are stated under explicit assumptions: Assumption 1(ii) posits E[ξ^2]=O(K^{-Q}), and the paper identifies this rate as supplied by Theorem 1 for its class (‘The second condition is the approximation-rate input supplied by Section 3’), which is a modular input, not a circular one. The OLS, LASSO, and logistic proofs carry the projection residual through standard concentration and optimization arguments; no fitted parameter is renamed as a prediction. The self-citations involving overlapping authors (Guo et al. 2025 for irrepresentability of signature dictionaries, Bayer et al. 2026 for background) are not load-bearing: the theorems treat irrepresentability and spectral nondegeneracy as assumptions rather than deriving them from those citations. The reviewer-flagged issue about K-dependent sub-Weibull constants in Lemma A.3 is a possible correctness/rate-condition gap, not a circular reduction, and therefore does not affect the circularity score.

Axiom & Free-Parameter Ledger

3 free parameters · 7 axioms · 0 invented entities

The theory introduces no fitted constants: the exponent Q=2γ is a theorem output, and the inputs are standard inequalities, the stated smoothness class, and well-conditioning/irrepresentability assumptions common in sieve and high-dimensional statistics. The empirical section adds three data-dependent choices (K, λ, EEG threshold) that affect the reported gains but not the theorems. The main ledger burden is Assumption 1: it is the entry point through which all three consistency theorems become conditional on unproved spectral properties of real signature dictionaries, plus the deferral of the LASSO irrepresentability condition to a self-cited prior paper (Guo et al. 2025).

free parameters (3)
  • EEG decision threshold = 0.10
    Section 6.3.2 and Appendix C.4: the operating threshold is chosen from the test-set threshold sweep, where patient-averaged balanced accuracy peaks at 0.10 (Table A.7). All threshold-dependent metrics (ACC, B-ACC, SENS, SPEC) are reported at this value.
  • LASSO/elastic-net regularization strength λ = 5-fold CV (FX); not fully specified (battery, EEG)
    The feature-selection penalty is tuned on data, affecting which signature coordinates enter the reported models. It is an algorithmic tuning parameter rather than a constant in the theorems, but it is data-fitted and central to the empirical claims.
  • Application truncation levels K = K=3 (FX, battery), K=2 (EEG)
    Chosen via the log n/(2 log(d+1)) heuristic (Sec. 4.1) and dimensionality constraints; battery sensitivity analysis (Table A.4) shows K=3 optimal on the test set. A modeling choice, not an ad hoc constant inside the theorems.
axioms (7)
  • domain assumption b and σ globally Lipschitz and uniformly bounded, Eq. (9); X solves the Itô diffusion SDE, Eq. (8)
    Scope restriction of U_{γ,R} (Def. 2). Excludes common models with unbounded coefficients (GBM, CIR). Needed for the exponential-martingale tail bound (A.16) and BDG arguments in the proof of Theorem 1.
  • domain assumption Functional class (Eq. 10): Y = Σ_a ∫ h_a ∘ dX̃^a with smooth h; mixed smoothness norm H_γ (Eq. 11) requiring 2γ+1 / 2γ+2 spatial derivatives
    The rate K^{-2γ} is proved only for this class, which the paper states is narrower than all continuous path functionals (Sec. 3.2). Barriers, stopping rules, and occupation times are excluded (Sec. 7).
  • standard math Jackson polynomial approximation bounds on [0,T]×[-m,m]^d with explicit scaling, and classical L2 polynomial approximation lower bounds for Sobolev functions
    Used in Steps 1 and 5 of the proof of Theorem 1 (App. A.3, Eqs. A.17-A.18 and A.38).
  • standard math Burkholder-Davis-Gundy inequality (Schachermayer-Stebegg) and sub-Weibull tail bounds for signature coordinates (Lemma A.4)
    Controls moment growth of iterated integrals and drives the LASSO concentration argument in App. B.1 (Eq. A.80 and Lemma A.4).
  • standard math Wiener-Itô chaos expansion and Lemmas A.1/A.2 (Itô-Stratonovich span equality; signature-chaos polynomial identification)
    The Brownian benchmark and the minimax lower bound reduce the first-chaos problem to polynomial approximation of chaos kernels (Eq. A.8, A.37).
  • domain assumption LASSO irrepresentability condition, Assumption 2(i), inherited from Zhao-Yu (2006)/Wainwright (2009); applicability for Brownian signatures deferred to Guo et al. (2025)
    Load-bearing support for Theorem 3. Guo et al. is co-authored by two authors of this paper; the condition is assumed, not re-derived, for the diffusion case.
  • domain assumption Spectral uniformity: c < λ_min(Σ_K) ≤ λ_max(Σ_K) < C (Assumption 1(i)) and local logistic Hessian non-degeneracy (Assumption 3(i))
    Not proved for any concrete signature dictionary; the paper notes the full signature violates it via pure-time coordinates (Remark after Assumption 1).

pith-pipeline@v1.3.0-alltime-deepseek · 51623 in / 22461 out tokens · 180557 ms · 2026-08-01T16:46:28.550356+00:00 · methodology

0 comments
read the original abstract

Many prediction and decision-making problems in operations research involve path-valued covariates -- data that evolve over time -- for which path signatures have become a canonical feature representation. Their use is justified by a universal approximation theorem, but this is an existence result: it guarantees that a finite-level signature can approximate any continuous path functional, without quantifying how fast the approximation error decreases as the truncation level grows. This paper develops approximation and statistical theory for signature-based path regression. We establish an \(L^2\) approximation rate for smooth functionals of It\^{o} diffusions and show that it is minimax optimal. We then propagate the truncation error through three statistical learning procedures -- Signature-OLS, Signature-LASSO, and Signature-Logistic -- and establish their consistency. Three real-data applications show that signatures provide informative finite-dimensional representations of path-valued covariates and can improve prediction relative to handcrafted features, in the context of finance -- foreign exchange realized volatility forecasting from intraday price paths; energy -- battery end-of-life prediction from early diagnostic current-voltage pulse paths; and medicine -- epileptic seizure detection from short electroencephalogram windows.

Figures

Figures reproduced from arXiv: 2607.17865 by Binnan Wang, Blanka Horvath, Ruixun Zhang, Wen Su, Wu Su.

Figure 1
Figure 1. Figure 1: shows that the projection error decreases as the signature depth K increases. The smoother target, indexed by γ = 2, is approximated faster than the rougher target, indexed by γ = 1, in both the Brownian and OU benchmarks. The Brownian experiment gives the cleanest view of the polynomial-kernel approximation mechanism, while the OU experiment shows the same ordering with more visible pre-asymptotic effects… view at source ↗
Figure 2
Figure 2. Figure 2: illustrates the finite-sample approximation–estimation tradeoff. Small truncation levels are stable but underfit, whereas deeper signatures reduce approximation bias only at the cost of estimating many more coefficients. In the one-dimensional time-augmented case, the full nonconstant dictionary has dK = 2K+1 − 2 coordinates, so K = 8 already corresponds to 510 regressors. In both the OLS and logistic expe… view at source ↗
Figure 3
Figure 3. Figure 3: Signature-LASSO under an OU-driven integral target. Panel (a) reports out-of-sample prediction error. Panel (b) reports selection precision, defined as the proportion of selected coordi￾nates that belong to the finite-level OU-relevant family. Results are averaged over 100 independent replications. In summary, the simulations provide a concise numerical check of the three theoretical messages. Signature pr… view at source ↗
Figure 4
Figure 4. Figure 4: Forecasting advantage of Signature-OLS over benchmark models across currency pairs. Left panel (a): For each method, dots show the ratio of its RMSE to the Signature-OLS RMSE across all pairs; the horizontal dashed line marks parity (ratio = 1). Right panel (b): Percentage reduction in RMSE of Signature-OLS relative to the best-performing non-signature baseline for each currency pair. Positive values indic… view at source ↗
Figure 5
Figure 5. Figure 5: Illustration of the battery end-of-life definition and the retained HPPC pulse segment used for signature construction [PITH_FULL_IMAGE:figures/full_fig_p025_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Observed versus predicted battery end of life for the six main feature–estimator combina￾tions. Each panel includes the 45◦ reference line and distinguishes training and test observations. 6.3 Epileptic Seizure Detection 6.3.1 Empirical Design We finally consider a binary-response application: epileptic seizure detection from short EEG win￾dows. The dataset is the CHB-MIT scalp EEG database (Shoeb and Gutt… view at source ↗
Figure 7
Figure 7. Figure 7: Representative EEG windows from the CHB-MIT dataset. The figure illustrates typical waveform differences between non-seizure and seizure segments [PITH_FULL_IMAGE:figures/full_fig_p029_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Diagnostic results for seizure detection on the CHB-MIT dataset. Panel (a) reports patient-averaged balanced accuracy across decision thresholds. Panel (b) reports cross-subject mean ROC curves. Panel (c) compares subject-wise balanced accuracy of Signature-Logistic (EN) with the handcrafted random forest [PITH_FULL_IMAGE:figures/full_fig_p032_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 9 canonical work pages · 1 internal anchor

  1. [3]

    Let bAdenote the selected support

    Results are averaged over 100 independent replications. Let bAdenote the selected support. Prediction is evaluated by independent test MSE. The selection metric in the main text is precision for the OU-relevant family, Precision = | bA∩A OU K | | bA| ,(A.141) with the convention that precision is zero when bA=∅. For the selection diagnostic, the LASSO pen...

  2. [4]

    We first prove several technical lemmas

    Throughout the appendix, we maintain the standing assumption that the underlying path is an Itˆ o diffusion satisfying the conditions in A-14 Section 3, and that the target variableYis generated from the pair (X,h)∈ U γ,R. We first prove several technical lemmas. Lemma A.3 establishes the convergence rate of the sample covariance matrix and the convergenc...

  3. [5]

    doi: 10.1111/1467-9868.00336. O. E. Barndorff-Nielsen and N. Shephard. Power and bipower variation with stochastic volatility and jumps.Journal of Financial Econometrics, 2(1):1–37,

  4. [10]

    1109/TNSRE.2015.2505238. M. Zabihi, S. Kiranyaz, V. J ¨antti, T. Lipping, and M. Gabbouj. Patient-specific seizure detection using nonlinear dynamics and nullclines.IEEE Journal of Biomedical and Health Informatics, 24(2):543–555,

  5. [13]

    doi: 10.48550/arXiv.1603.03788. I. Chevyrev and T. Lyons. Characteristic functions of measures on geometric rough paths.The Annals of Probability, 44(6):4049–4082,

  6. [14]

    doi: 10.1214/15-AOP1068. I. Chevyrev and H. Oberhauser. Signature moments to characterize laws of stochastic processes. Journal of Machine Learning Research, 23(176):1–42,

  7. [17]

    doi: 10.1214/11-AOP721. F. Corsi. A simple approximate long-memory model of realized volatility.Journal of Financial Econometrics, 7(2):174–196,

  8. [21]

    doi: 10.1016/j.jmva.2022.105031. G. Flint, B. Hambly, and T. Lyons. Discretely sampled signals and the rough hoff process.Stochastic Processes and their Applications, 126(9):2593–2614,

  9. [22]

    doi: 10.1016/j.spa.2016.02.011. P. K. Friz and N. B. Victoir.Multidimensional Stochastic Processes as Rough Paths: Theory and Applications. Cambridge University Press,

  10. [23]

    doi: 10.1017/CBO9780511845079. M. Fujita, N. Sugiura, and S. Kouketsu. Prediction of atmospheric profiles with machine learn- ing using the signature method.Geophysical Research Letters, 51(6),

  11. [24]

    doi: 10.1287/opre.2024.1133. B. Hambly and T. Lyons. Uniqueness for the signature of a path of bounded variation and the reduced path group.Annals of Mathematics, 171(1):109–167,

  12. [25]

    doi: 10.4007/annals.2010. 171.109. S. H¨ormann and P. Kokoszka. Weakly dependent functional data.The Annals of Statistics, 38(3): 1845–1884,

  13. [26]

    doi: 10.1214/09-AOS768. R. Ibraheem, P. Dechent, and G. dos Reis. Path signature-based life prognostics of li-ion battery using pulse test data.Applied Energy, 378:124820,

  14. [27]

    doi: 10.1016/j.apenergy.2024.124820. 36 P. Kidger, P. Bonnier, I. Perez Arribas, C. Salvi, and T. Lyons. Deep signature transforms.Advances in Neural Information Processing Systems, 32,

  15. [28]

    doi: 10.48550/arXiv.1905.08494. F. J. Kir´ aly and H. Oberhauser. Kernels for sequentially ordered data.Journal of Machine Learning Research, 20(31):1–45,

  16. [29]

    doi: 10.5555/3322706.3361972. S. Kiranyaz, T. Ince, M. Zabihi, and D. Ince. Automated patient-specific classification of long-term electroencephalography.Journal of Biomedical Informatics, 49:16–31,

  17. [31]

    doi: 10.48550/arXiv.1309

  18. [33]

    doi: 10.1080/1350486X.2021.1891555. T. J. Lyons. Differential equations driven by rough signals.Revista Matem´ atica Iberoamericana, 14(2):215–310,

  19. [38]

    doi: 10.1016/j.jpowsour.2022.231127. 37 W. Schachermayer and F. Stebegg. The sharp constant for the burkholder–davis–gundy inequality and non-smooth pasting.Bernoulli, 24(4A):3032–3051,

  20. [40]

    doi: 10.1007/s10916-019-1234-4. K. A. Severson, P. M. Attia, N. Jin, N. Perkins, B. Jiang, Z. Yang, M. H. Chen, M. Aykol, P. K. Herring, D. Fraggedakis, M. Z. Bazant, S. J. Harris, W. C. Chueh, and R. D. Braatz. Data-driven prediction of battery cycle life before capacity degradation.Nature Energy, 4(5):383–391,

  21. [41]

    doi: 10.1038/s41560-019-0356-8. A. H. Shoeb and J. V. Guttag. Application of machine learning to epileptic seizure detection. In Proceedings of the 27th International Conference on Machine Learning, pages 975–982,

  22. [42]

    doi: 10.5555/3104322.3104446. S. A. van de Geer. High-dimensional generalized linear models and the lasso.The Annals of Statistics, 36(2):614–645,

  23. [44]

    doi: 10.1109/TIT.2009.2016018. M. Zabihi, S. Kiranyaz, A. B. Rad, A. K. Katsaggelos, M. Gabbouj, and T. Ince. Analysis of High- Dimensional Phase Space via Poincar´ e Section for Patient-Specific Seizure Detection.IEEE Transactions on Neural Systems and Rehabilitation Engineering, 24(3):386–398,

  24. [46]

    doi: 10.1109/JBHI.2019.2906400. H. Zhang and S. X. Chen. Concentration inequalities for statistical inference.Communications in Mathematical Research, 37(1):1–85,

  25. [47]

    doi: 10.4208/cmr.2020-0041. P. Zhao and B. Yu. On model selection consistency of lasso.Journal of Machine Learning Research, 7:2541–2563,

  26. [1957]

    doi: 10.2307/1969671. I. Chevyrev and A. Kormilitzin. A primer on the signature method in machine learning.arXiv preprint arXiv:1603.03788,

  27. [1998]

    doi: 10.4171/RMI/240. J. Morrill, C. Salvi, P. Kidger, and J. Foster. Neural rough differential equations for long time series. InProceedings of the 38th International Conference on Machine Learning, pages 7829–7838,

  28. [2001]

    doi: 10.1093/rfs/14.1.113. T. Lyons, S. Nejad, and I. Perez Arribas. Non-parametric pricing and hedging of exotic derivatives. Applied Mathematical Finance, 27(6):457–494,

  29. [2002]

    doi: 10.1111/1468-0262.00274. O. E. Barndorff-Nielsen and N. Shephard. Econometric analysis of realised volatility and its use 34 in estimating stochastic volatility models.Journal of the Royal Statistical Society: Series B (Statistical Methodology), 64(2):253–280,

  30. [2003]

    doi: 10.1111/1468-0262.00418. T. G. Andersen, T. Bollerslev, and F. X. Diebold. Roughing it up: Including jump components in the measurement, modeling, and forecasting of return volatility.The Review of Economics and Statistics, 89(4):701–720,

  31. [2004]

    doi: 10.1093/jjfinec/nbh001. C. Bayer, L. Pelizzari, and J. Schoenmakers. Primal and dual optimal stopping with signatures. Finance and Stochastics, 29:981–1014,

  32. [2006]

    38 Supplementary Appendices (Electronic Companion) Appendix Contents A Theoretical Details for Section 3 A-1 A.1 Technical Lemmas

    doi: 10.5555/1248547.1248637. 38 Supplementary Appendices (Electronic Companion) Appendix Contents A Theoretical Details for Section 3 A-1 A.1 Technical Lemmas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . A-2 A.2 Signature-Chaos Relation and Brownian Signature Approximation Rate . . . . . . . A-3 A.3 Proof of Theorem 1 . . . ....

  33. [2007]

    doi: 10.1162/rest.89.4.701. P. M. Attia, A. Grover, N. Jin, K. A. Severson, T. M. Markov, Y.-H. Liao, M. H. Chen, B. Cheong, N. Perkins, Z. Yang, P. K. Herring, M. Aykol, S. J. Harris, R. D. Braatz, S. Ermon, and W. C. Chueh. Closed-loop optimization of fast-charging protocols for batteries with machine learning. Nature, 578:397–402,

  34. [2008]

    doi: 10.1214/009053607000000929. M. J. Wainwright. Sharp thresholds for high-dimensional and noisy sparsity recovery usingℓ 1- constrained quadratic programming.IEEE Transactions on Information Theory, 55(5):2183– 2202,

  35. [2009]

    doi: 10.1093/jjfinec/nbp001. C. Cuchiero, G. Gazzani, and S. Svaluto-Ferro. Signature-based models: Theory and calibration. SIAM Journal on Financial Mathematics, 14(3):910–957,

  36. [2010]

    doi: 10.1016/j.jfa.2010.04.017. 35 R. Cont and D.-A. Fourni´ e. Functional itˆ o calculus and stochastic integral representation of mar- tingales.The Annals of Probability, 41(1):109–133,

  37. [2011]

    doi: 10.1016/j.jeconom.2010.03.034. A. J. Patton and K. Sheppard. Good volatility, bad volatility: Signed jumps and the persistence of volatility.The Review of Economics and Statistics, 97(3):683–697,

  38. [2013]

    doi: 10.3150/11-BEJ410. H. Boedihardjo, X. Geng, T. Lyons, and D. Yang. The signature of a rough path: Uniqueness. Advances in Mathematics, 293:720–737,

  39. [2014]

    2014.02.005

    doi: 10.1016/j.jbi. 2014.02.005. D. Levin, T. Lyons, and H. Ni. Learning from the past, predicting the statistics for the future, learning an evolving system.arXiv preprint arXiv:1309.0260,

  40. [2015]

    doi: 10.1162/REST a 00503. N. H. Paulson, J. Kubal, L. Ward, S. Saxena, W. Lu, and S. J. Babinec. Feature engineering for machine learning enabled early prediction of battery lifetime.Journal of Power Sources, 527: 231127,

  41. [2016]

    doi: 10.1016/j.aim.2016.02.011. K.-T. Chen. Integration of paths, geometric invariants and a generalized baker-hausdorff formula. Annals of Mathematics, 65(1):163–178,

  42. [2018]

    doi: 10.3150/17-BEJ935. R. S. Selvakumari, M. Mahalakshmi, and P. Prashalee. Patient-Specific Seizure Detection Method using Hybrid Classifier with Optimized Electrodes.Journal of Medical Systems, 43(5):121,

  43. [2019]

    doi: 10.1080/ 14697688.2019.1575974. A. Fermanian. Functional linear regression with truncated signatures.Journal of Multivariate Analysis, 192:105031,

  44. [2020]

    doi: 10.1038/s41586-020-1994-5. Y. A¨ıt-Sahalia. Maximum likelihood estimation of discretely sampled diffusions: a closed-form approximation approach.Econometrica, 70(1):223–262,

  45. [2021]

    doi: 10.48550/arXiv.2009.08295. A. J. Patton. Volatility forecast comparison using imperfect volatility proxies.Journal of Econo- metrics, 160(1):246–256,

  46. [2022]

    doi: 10.48550/arXiv.1810.10971. R. Cont and D.-A. Fourni´ e. Change of variable formulas for non-anticipative functionals on path space.Journal of Functional Analysis, 259(4):1043–1072,

  47. [2023]

    doi: 10.1137/22M1512338. B. Dupire. Functional itˆ o calculus.Quantitative Finance, 19(5):721–729,

  48. [2024]

    doi: 10.1137/23M1571563. A. Belloni and V. Chernozhukov. Least squares after model selection in high-dimensional sparse models.Bernoulli, 19(2):521–547,

  49. [2025]

    doi: 10.1007/s00780-025-00570-8. C. Bayer, G. dos Reis, B. Horvath, and H. Oberhauser, editors.Signature Methods in Finance: An Introduction with Computational Applications. Springer Finance Lecture Notes. Springer,

  50. [2026]

    doi: 10.1007/978-3-031-97239-3. E. Bayraktar, Q. Feng, and Z. Zhang. Deep signature algorithm for multidimensional path- dependent options.SIAM Journal on Financial Mathematics, 15(1):194–214,