Pith. sign in

REVIEW 2 major objections 4 minor 52 references

The optimal rule for quickest detection on rough-path space is the first time a linear functional of the path signature crosses a threshold.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 04:02 UTC pith:JAGR4BSK

load-bearing objection Concatenation model and BHR adaptation are solid, but §4's geometric detector has a load-bearing measurability gap; the numerics still stand but the statistical claims need rework. the 2 major comments →

arxiv 2607.22958 v1 pith:JAGR4BSK submitted 2026-07-24 math.ST cs.ITmath.ITmath.PRstat.TH

Quickest Detection with Rough Path Signatures

classification math.ST cs.ITmath.ITmath.PRstat.TH MSC 62L1560G4062C20
keywords quickest detectionchange-point detectionrough pathssignaturesoptimal stoppingnon-Markovianrobust detectionfractional Brownian motion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper aims to show that quickest detection of a change in the law of an irregular, non-Markovian signal can be solved optimally by a single linear functional of the path's rough-path signature: the rule is to alarm the first time |⟨l, signature⟩| reaches one. If true, this reduces a path-space optimal stopping problem to learning one coefficient vector, with no parametric model and no Markovian sufficiency required. The paper also constructs a geometrically motivated detector from the p-variation distance to the pre-change path, proves finite-sample bounds on delay and false alarm probability, extends to repeated experiments and distributionally robust settings, and demonstrates numerically that the learned rules match or beat classical baselines in Brownian and fractional-Brownian settings.

Core claim

The central claim, Proposition 3.5, is that the infimum over all (F^X_t)-stopping times of E[Y_{τ∧T}] equals the infimum over linear functionals l of E[Y_{τ_l∧T}], where τ_l is the first hitting time of the half-space |⟨l, X^{<∞}_{0,t}⟩| ≥ 1 and X is the concatenation of the pre- and post-change rough paths at the unknown change-point θ. This holds for both Bayesian loss objectives considered (false alarm plus delay penalty, and symmetric early/late penalties). The paper argues the same structural form emerges independently from the intrinsic geometry of the observed path, namely from the p-variation distance to the pre-change path, and derives statistical guarantees: with probability at lea

What carries the argument

The signature of a rough path — the sequence of iterated integrals encoding the entire path history — serves as a universal, model-free feature map. The paper's load-bearing mechanism is the half-space hitting time τ_l for a linear functional l of the signature, together with a density theorem (Lemma 2.6) stating that any continuous stopping policy can be uniformly approximated by such linear signature functionals on a compact set of probability arbitrarily close to one. Chen's identity lets concatenation at the change-point be represented as tensor product of signatures, and the homogeneous p-variation distance is used both to measure post-change divergence and to define the geometric detec

Load-bearing premise

The load-bearing premise is that the geometric detector f—the p-variation distance from the observed path to the unobserved pre-change path X^{(1)}—can be uniformly approximated by a single linear signature functional on a set of probability close to one, even though f depends on the post-change continuation of X^{(1)} that the observed path does not reveal.

What would settle it

Simulate the Brownian disorder model twice with the same post-change path X^{(2)} and same observed path up to θ, but different latent pre-change paths X^{(1)} that agree on [0,θ] and differ on (θ,T]. If the geometric detector f differs between the two runs while the observed path is identical, no signature functional of the observed path alone can approximate f uniformly, disproving Proposition 4.1's approximation claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Quickest detection becomes a finite-dimensional learning problem: estimate one linear functional from data, then monitor its absolute value against a threshold.
  • The method applies to non-Markovian, non-semimartingale signals such as fractional Brownian motion, where classical sufficient statistics like likelihood ratios are unavailable.
  • Explicit probabilistic bounds tie detection delay to the inverse of the separation function g and the approximation tolerance, giving a clear trade-off between speed and accuracy.
  • With R independent paths sharing a change point, both the expected delay and false-alarm probability shrink toward their lower bounds exponentially in R.
  • Under the least-favorable-pair assumption, the distributionally robust problem is solved by the same signature half-space rule, extending the approach to adversarial perturbations.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If Proposition 3.5 holds for all rough-path laws, it suggests the signature hitting time is a universal architecture for change-point detection, with the truncation level N as the only tuning parameter — a testable claim beyond the paper's specific examples.
  • The geometric-detector argument would be falsifiable in practice: check whether a single signature functional can uniformly approximate the p-variation distance when the latent pre-change path is only observed up to θ; the paper's Lemma 2.6 justification does not obviously cover dependence on the unobserved post-θ continuation.
  • The robust formulation hints that adversarial training of the signature coefficient could serve as a general-purpose robustification layer for any sequential decision rule, not just quickest detection.
  • The Brownian example suggests that the learned coefficient l should approximately recover the classical likelihood-ratio statistic; verifying that link could connect the rough-path approach to classical theory in a precise way.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes a rough-path signature framework for Bayesian quickest detection. The observed signal is modeled as the concatenation of two independent geometric p-rough paths at an unknown change-point θ. The authors show (Prop. 3.5) that the optimal stopping rule for the natural loss processes Y^1 and Y^2 coincides with the first hitting time of a half-plane by a linear functional of the time-augmented signature, extending the signature optimal-stopping theory of [3]. A 'geometric detector' based on the p-variation distance to the pre-change path is introduced (§4), and statistical guarantees on delay and false-alarm probability are derived (Props. 4.1–4.4), along with a repeated-experiment aggregation rule and a distributionally robust minimax extension (§5). Numerical experiments compare the learned signature rules with CUSUM/Shiryaev in Brownian models and with Page–Hinkley in fractional Brownian models, and demonstrate robustness to adversarial total-variation perturbations.

Significance. If Prop. 3.5 is correct, the paper makes a useful contribution by showing that signature half-space hitting times are universal for a class of non-Markovian change-point problems; the proof in Appendix B is a substantial part of the paper. The numerical evidence is promising and the robust formulation is sensible. However, the geometric-detector argument in §4—one of the advertised main contributions—is not valid as written because the detector is not a functional of the observed path. The statistical guarantees in §4 therefore do not follow from the supplied proof. The paper has enough independent value in Prop. 3.5 and the experiments to warrant a major revision, but the theoretical claims in §4 must be either corrected or substantially weakened.

major comments (2)
  1. [§4.1, Eq. (22) and proof of Prop. 4.1] The function f(\hat Z)=f(Z,t)=d_{p-var;[0,t]}(Z,X^{(1)}) is not a deterministic continuous functional of the observed path. In the model of §3, X^{(1)} is latent and independent of X^{(2)}; for t>θ the observed segment X_{[θ,t]} equals X^{(2)}_{[θ,t]} and contains no information about X^{(1)} on [θ,t]. Hence f(X,t) is not σ(\hat X_s:0≤s≤t)-measurable. Lemma 2.6 applies only to elements of T=C(Λ_T,R), i.e., functions that map each observed path segment to a real number. The proof's assertion 'Consider f∈T' is therefore unjustified; the uniform approximation (26), inequalities (24)–(25), and Propositions 4.2–4.4 are not established. The inequality E[f(X,t)]≥g((t-θ)^+) in (23) concerns a true but unobserved distance and does not make f an observable path functional.
  2. [Contribution (ii) and §7] The paper repeatedly claims that the same structural form 'arises independently from the intrinsic geometry of the observed path' (Abstract, §4, §7). Since Prop. 4.1 is the only support for this claim and it fails, the claim is unsupported. The numerical experiments in §6 train l directly via zeroth-order optimization and do not rely on Prop. 4.1; Prop. 3.5 is also independent. The authors should either (a) redefine the geometric detector so that it is a function of the observed path alone (e.g., distance to a known nominal pre-change path if one is available), or (b) present §4 as heuristic and remove or rework the statistical guarantees that depend on Prop. 4.1.
minor comments (4)
  1. [Abstract and Table 1] The abstract says the signature rules 'outperform' classical methods. In the Brownian experiments (Table 1), Shiryaev has lower E[Y^1] and shorter delay than the signature rule; the outperformance is really against Page–Hinkley in the fractional Brownian setting. Please qualify the claim.
  2. [§5.2, Example 5.3] The text says Assumption 5.1 holds for the Wasserstein ambiguity sets because of weak compactness and upper semicontinuity, but no proof is supplied. Either prove the statement for Example 5.3 or present it as a condition to be verified.
  3. [Corollary 4.3] The coefficient l* is defined as an arginf over T((R^{1+d})*). Existence of a minimizer is not established; Prop. 3.5 states equality of infima, not attainment. Please rephrase using an ε-optimal coefficient or prove existence.
  4. [§6.7] Restricting the adversary to perturbations with ||w||_{TV}=C_TV is described as 'without loss of generality'. For a general non-concave objective the supremum over the TV ball need not be attained on the boundary; please justify this reduction or soften the claim.

Circularity Check

1 steps flagged

Prop 4.1's geometric detector assumes the signature approximation it purports to prove: f in eq. (22) depends on the latent X^(1), so Lemma 2.6 does not apply to the observed path.

specific steps
  1. other [Section 4.1, eq. (22) and proof of Proposition 4.1; Lemma 2.6]
    "More precisely, for any Z∈Ω^p_T, define f( bZ):=f(Z,t)=d_p−var;[0,t](Z,X^{(1)}). (22) Observe that f is a continuous function... Proof. Consider f∈T defined as in (22). By Lemma 2.6, for any ε>0, there exists a compact set K⊂Ω̂^p_T with P(K)≥1−ε ..."

    In Section 3, X^(1) is latent (P=μ1⊗μ2⊗ν) and is not observed on [θ,t]; eq. (22) does not define a deterministic functional of the observed path. Lemma 2.6 applies only to T=C(Λ_T,R), functions of the observed time-augmented path. The proof's 'Consider f∈T' is the point to be proved: it assumes the geometric distance to the hidden pre-change path is an observed-path functional, after which the density theorem mechanically outputs a signature linear functional. The claimed 'same structural form arises independently from geometry' is therefore installed by hypothesis, not derived. Also, K has high probability under the observed-path law, not under the law of X^(1), so bound (24) is not established. Props 4.2-4.4 inherit this gap.

full rationale

The central optimal-stopping result (Prop 3.5) is not circular: it is an adaptation of the external theorem [3] (Bayer et al.) and the proof in Appendix B reduces to that external density/continuity machinery, not to this paper's own conclusions. No load-bearing self-citation chain is present; the authors' prior work [48] is only an optimization subroutine for the numerics. The main circularity-adjacent defect is in Section 4: the 'independent geometric detector' f in eq. (22) is defined through the latent pre-change rough path X^(1), so it is not a member of the observed-path function space T to which Lemma 2.6 applies. The proof of Prop 4.1 says 'Consider f∈T' and invokes Lemma 2.6; this is an unsupported reduction rather than a demonstration, and it makes the claimed uniform signature approximation (and the derived delay/false-alarm guarantees in Props 4.2-4.4) rest on the very signature-approximation conclusion they advertise. Because this defect is a false premise in a secondary 'prediction' (the geometric detector), not a circularity in the main optimal-stopping theorem, the overall circularity score is moderate (4), not higher. No evidence was found that the numerical experiments or the robust formulation introduce circular reasoning; their limitations are empirical rather than definitional.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 0 invented entities

The paper's reduction of quickest detection to signature half-space hitting times rests on the density theorem for signature functionals (cited, not reproved) and on Assumption 3.1's detectability condition. The geometric-detector part additionally needs the detector to be a genuine functional of the observed path, which the paper does not establish.

free parameters (6)
  • Learned signature coefficient l = not reported (esGS training)
    The stopping rule τ_l = inf{|⟨l, X̂^{<∞}⟩| ≥ 1} is trained by zeroth-order optimization on simulated paths; theoretical guarantees assume an l^{(1)} with separation properties exists, but experiments fit l to empirical risk without guarantees (Section 6.3-6.4).
  • Truncation level N = 4
    Signature truncated at order N=4 in all experiments; the theory works with the full signature.
  • Loss parameters c, a, b = c∈{0.25,...,2}, a=1, b∈{0.25,...,2}
    These define the objectives (4)-(5) and are varied to trace the delay-false alarm trade-off.
  • Adversarial budget C_TV = {0.5,1,2,4}
    Total-variation budget for adversarial perturbations in Section 6.7.
  • Dirichlet parameters α_m, α_t = not reported
    Parametrize the adversarial perturbation class (jump sizes and times); updated during adversarial training, values not given.
  • Page-Hinkley threshold (baseline) = calibrated to match P(τ<θ)
    Baseline threshold calibrated to the signature rule's false alarm; exact value not reported.
axioms (5)
  • standard math Lemma 2.6: continuous stopping policies can be uniformly approximated by linear functionals of the truncated signature on compact sets of high probability
    Cited from [3, Lemma 5.2] and [18, Lemma B.3] (signature density from Stone-Weierstrass); the entire reduction to signature hitting times rests on it (used in Propositions 3.5, 4.1, 4.2).
  • standard math Lyons' extension theorem and the signature of a geometric p-rough path exists, is multiplicative (Chen), and is group-like
    Standard rough path theory (Friz-Victoir [15], Lyons et al. [25]); used at the definition of the signature and in the shuffle-product step of Prop 3.5.
  • domain assumption Assumption 3.1: X^{(1)}, X^{(2)} are geometric p-rough paths with finite q-moments of the 1/p-Hölder norm, and their expected p-variation distance grows like g(|t−s|)
    This is the detectability/S-N condition; the delay bounds in Props 4.1-4.2 and 4.4 are expressed through g^{-1} and the moment bound M.
  • ad hoc to paper Assumption 5.1: a least favorable pair (P*_1, P*_2) exists for the robust problem
    Assumed to reduce the minimax problem to an ordinary stopping problem; the paper notes in Remark 5.2 that if it fails, the robust rule need not be a signature half-space hitting time.
  • ad hoc to paper The geometric detector f(Z,t)=d_{p-var}(Z, X^{(1)}) is a well-defined continuous functional of the observed path alone, approximable by one signature functional on a high-probability set
    Unstated and questionable: X^{(1)} is latent, so f is not determined by the observed path; Prop 4.1's proof invokes Lemma 2.6 on f without addressing this dependence (Section 4.1).

pith-pipeline@v1.3.0-alltime-deepseek · 34553 in / 25514 out tokens · 210476 ms · 2026-08-01T04:02:10.294220+00:00 · methodology

0 comments
read the original abstract

We develop a framework for quickest detection of distributional change in signals modeled as rough paths. By representing the pre-change and post-change dynamics as two independent rough paths and modeling the observed signal as their concatenation at an unknown random change-point, we formulate quickest detection as an optimal stopping problem using rough paths. We show that the optimal stopping rule takes the form of the first hitting time to a half-space by a linear functional of the rough path signature, and establish that the same structural form arises independently from the intrinsic geometry of the observed path. Statistical guarantees on detection delay and false alarm probability are derived, and the framework is extended to a distributionally robust formulation in which both the pre-change and post-change models are uncertain. The proposed rules are implemented via zeroth-order stochastic approximation over truncated-signature coefficients and evaluated numerically under Brownian and fractional Brownian dynamics, achieving performance comparable to optimal methods in the Brownian setting while outperforming them and remaining robust to adversarial path perturbations in the fractional Brownian setting.

Figures

Figures reproduced from arXiv: 2607.22958 by Mingrui Wang, Prakash Chakraborty.

Figure 1
Figure 1. Figure 1: An example of the hidden and observed paths in R2 3.2. The observed process. Having specified the probabilistic setup, we now describe the obser￾vation model. The observer has access neither to X(1) and X(2) individually nor to the change-point θ, but only to the path obtained by switching from one regime to the other at time θ. In our setting of the quickest detection problem, the individual processes X (… view at source ↗
Figure 2
Figure 2. Figure 2: Path examples with θ ∼ Exp(0.5) 6.5. Quickest detection for rough paths. Classical methods such as CUSUM and Shiryaev are most tractable when the likelihood ratio is explicitly available, as in Brownian diffusion models with known parameters. Extending quickest detection methods to rough path and general change time distributions is substantially more challenging. This work provides a way to address this g… view at source ↗
Figure 3
Figure 3. Figure 3: Change-point θ Exp and payoff Y 1 [PITH_FULL_IMAGE:figures/full_fig_p023_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Change-point θWei and payoff Y 1 . From [PITH_FULL_IMAGE:figures/full_fig_p023_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Change-point θ Exp and payoff Y 2 [PITH_FULL_IMAGE:figures/full_fig_p023_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Change-point θWei and payoff Y 2 . Since CUSUM and Shiryaev require specific structure of the model which is not available in this example, in [PITH_FULL_IMAGE:figures/full_fig_p024_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Page-Hinkley quickest detection experiment with change-point θ Exp [PITH_FULL_IMAGE:figures/full_fig_p024_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Page-Hinkley quickest detection experiment with change-point θWei . For change-point θ Exp, we consider the two loss functions Y 1 and Y 2 defined in (4) and (5), and examine how tuning the parameters in the loss affects the trade-off between expected detection delay and false alarm probability. For the loss Y 1 , we vary c ∈ {0.25, 0.5, 1, 1.5, 2}, retraining the stopping rule for each value. The results … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 1 linked inside Pith

  1. [1]

    The Pontryagin maximum principle andQ-functions in rough environments.arXiv preprint arXiv:2601.05354, 2026

    Estepan Ashkarian, Prakash Chakraborty, Harsha Honnappa, and Samy Tindel. The Pontryagin maximum principle andQ-functions in rough environments.arXiv preprint arXiv:2601.05354, 2026

  2. [2]

    Pricing under rough volatility.Quantitative Finance, 16(6):887– 904, 2016

    Christian Bayer, Peter Friz, and Jim Gatheral. Pricing under rough volatility.Quantitative Finance, 16(6):887– 904, 2016. 30 MINGRUI WANG AND PRAKASH CHAKRABORTY

  3. [3]

    Hager, Sebastian Riedel, and John Schoenmakers

    Christian Bayer, Paul P. Hager, Sebastian Riedel, and John Schoenmakers. Optimal stopping with signatures. The Annals of Applied Probability, 33(1):238–273, 2023

  4. [4]

    Primal and dual optimal stopping with signatures

    Christian Bayer, Luca Pelizzari, and John Schoenmakers. Primal and dual optimal stopping with signatures. Finance and Stochastics, 29:981–1014, 2025

  5. [5]

    Fractional processes as models in stochastic finance

    Christian Bender, Tommi Sottinen, and Esko Valkeila. Fractional processes as models in stochastic finance. In Advanced Mathematical Methods for Finance, pages 75–103. Springer Berlin Heidelberg, 2011

  6. [6]

    Pakkanen

    Mikkel Bennedsen, Asger Lunde, and Mikko S. Pakkanen. Decoupling the short- and long-term behavior of stochastic volatility.Journal of Financial Econometrics, 20(5):961–1006, 2022

  7. [7]

    Pathwise relaxed optimal control of rough differential equations.arXiv preprint arXiv:2402.17900, 2024

    Prakash Chakraborty, Harsha Honnappa, and Samy Tindel. Pathwise relaxed optimal control of rough differential equations.arXiv preprint arXiv:2402.17900, 2024

  8. [8]

    A primer on the signature method in machine learning

    Ilya Chevyrev and Andrey Kormilitzin. A primer on the signature method in machine learning. InSignature Methods in Finance, pages 3–64. Springer Nature Switzerland, 2026

  9. [9]

    Shanbhag, and Farzad Yousefian

    Shisheng Cui, Uday V. Shanbhag, and Farzad Yousefian. Complexity guarantees for an implicit smoothing- enabled method for stochastic MPECs.Mathematical Programming, 198(2):1153–1225, 2023

  10. [10]

    Friz, and Paul Gassiat

    Joscha Diehl, Peter K. Friz, and Paul Gassiat. Stochastic control with rough paths.Applied Mathematics & Optimization, 75(2):285–315, 2017

  11. [11]

    Embedding and learning with signatures.Computational Statistics & Data Analysis, 157:107148, 2021

    Adeline Fermanian. Embedding and learning with signatures.Computational Statistics & Data Analysis, 157:107148, 2021

  12. [12]

    Functional linear regression with truncated signatures.Journal of Multivariate Analysis, 192:105031, 2022

    Adeline Fermanian. Functional linear regression with truncated signatures.Journal of Multivariate Analysis, 192:105031, 2022

  13. [13]

    A note on the notion of geometric rough paths.Probability Theory and Related Fields, 136(3):395–416, 2006

    Peter Friz and Nicolas Victoir. A note on the notion of geometric rough paths.Probability Theory and Related Fields, 136(3):395–416, 2006

  14. [14]

    Friz and Martin Hairer.A Course on Rough Paths: With an Introduction to Regularity Structures

    Peter K. Friz and Martin Hairer.A Course on Rough Paths: With an Introduction to Regularity Structures. Universitext. Springer, Cham, 2020

  15. [15]

    Friz and Nicolas B

    Peter K. Friz and Nicolas B. Victoir.Multidimensional Stochastic Processes as Rough Paths: Theory and Appli- cations. Cambridge University Press, Cambridge, 2010

  16. [16]

    Volatility is rough.Quantitative Finance, 18(6):933– 949, March 2018

    Jim Gatheral, Thibault Jaisson, and Mathieu Rosenbaum. Volatility is rough.Quantitative Finance, 18(6):933– 949, March 2018

  17. [17]

    The limits of min-max optimization algorithms: Convergence to spurious non-critical sets

    Ya-Ping Hsieh, Panayotis Mertikopoulos, and Volkan Cevher. The limits of min-max optimization algorithms: Convergence to spurious non-critical sets. In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 4337–4348. PMLR, 2021

  18. [18]

    Optimal execution with rough path signatures.SIAM Journal on Financial Mathematics, 11(2):470–493, 2020

    Jasdeep Kalsi, Terry Lyons, and Imanol Perez Arribas. Optimal execution with rough path signatures.SIAM Journal on Financial Mathematics, 11(2):470–493, 2020

  19. [19]

    Neural controlled differential equations for irregular time series

    Patrick Kidger, James Morrill, James Foster, and Terry Lyons. Neural controlled differential equations for irregular time series. InAdvances in Neural Information Processing Systems, volume 33, pages 6696–6707, 2020

  20. [20]

    Goodfellow, and Samy Bengio

    Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. Adversarial machine learning at scale. InInternational Conference on Learning Representations, 2017

  21. [21]

    Sequential changepoint detection in quality control and dynamical systems.Journal of the Royal Statistical Society: Series B (Methodological), 57(4):613–644, 1995

    Tze Leung Lai. Sequential changepoint detection in quality control and dynamical systems.Journal of the Royal Statistical Society: Series B (Methodological), 57(4):613–644, 1995

  22. [22]

    Procedures for reacting to a change in distribution.The Annals of Mathematical Statistics, 42(6):1897–1908, 1971

    Gary Lorden. Procedures for reacting to a change in distribution.The Annals of Mathematical Statistics, 42(6):1897–1908, 1971

  23. [23]

    Oxford University Press, Oxford, 2002

    Terry Lyons and Zhongmin Qian.System Control and Rough Paths. Oxford University Press, Oxford, 2002

  24. [24]

    Terry J. Lyons. Differential equations driven by rough signals.Revista Matemática Iberoamericana, 14(2):215– 310, 1998

  25. [25]

    Terry J. Lyons, Michael Caruana, and Thierry Lévy.Differential Equations Driven by Rough Paths: École d’Été de Probabilités de Saint-Flour XXXIV - 2004, volume 1908 ofLecture Notes in Mathematics. Springer, Berlin, Heidelberg, 2007

  26. [26]

    Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu

    Aleksander Madry, Aleksandar A. Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. InInternational Conference on Learning Representations, 2018

  27. [27]

    Mandelbrot and John W

    Benoit B. Mandelbrot and John W. Van Ness. Fractional Brownian motions, fractional noises and applications. SIAM Review, 10(4):422–437, October 1968

  28. [28]

    Robust score-based quickest change detection.IEEE Transactions on Information Theory, 71(7):5539–5555, 2025

    Sean Moushegian, Suya Wu, Enmao Diao, Jie Ding, Taposh Banerjee, and Vahid Tarokh. Robust score-based quickest change detection.IEEE Transactions on Information Theory, 71(7):5539–5555, 2025

  29. [29]

    Moustakides

    George V. Moustakides. Optimal stopping times for detecting changes in distributions.The Annals of Statistics, 14(4):1379–1387, 1986

  30. [30]

    Moustakides

    George V. Moustakides. Optimality of the CUSUM procedure in continuous time.The Annals of Statistics, 32(1):302–315, 2004. QUICKEST DETECTION WITH SIGNATURES 31

  31. [31]

    Random gradient-free minimization of convex functions.Foundations of Computational Mathematics, 17(2):527–566, 2017

    Yurii Nesterov and Vladimir Spokoiny. Random gradient-free minimization of convex functions.Foundations of Computational Mathematics, 17(2):527–566, 2017

  32. [32]

    E. S. Page. Continuous inspection schemes.Biometrika, 41(1–2):100–115, 1954

  33. [33]

    Adversarial machine learning: a review of methods, tools, and critical industry sectors.Artificial Intelligence Review, 58(8):226, May 2025

    Sotiris Pelekis, Thanos Koutroubas, Afroditi Blika, Anastasis Berdelis, Evangelos Karakolis, Christos Ntanos, Evangelos Spiliotis, and Dimitris Askounis. Adversarial machine learning: a review of methods, tools, and critical industry sectors.Artificial Intelligence Review, 58(8):226, May 2025

  34. [34]

    Birkhäuser, Basel, 2006

    Goran Peskir and Albert Shiryaev.Optimal Stopping and Free-Boundary Problems. Birkhäuser, Basel, 2006

  35. [35]

    Optimal detection of a change in distribution.The Annals of Statistics, 13(1):206–227, 1985

    Moshe Pollak. Optimal detection of a change in distribution.The Annals of Statistics, 13(1):206–227, 1985

  36. [36]

    Reizenstein and Benjamin Graham

    Jeremy F. Reizenstein and Benjamin Graham. Algorithm 1004: The iisignature library: Efficient calculation of iterated-integral signatures and log signatures.ACM Transactions on Mathematical Software, 46(1):1–21, March

  37. [37]

    A. N. Shiryaev. On optimum methods in quickest detection problems.Theory of Probability & Its Applications, 8(1):22–46, 1963

  38. [38]

    Shiryaev.Optimal Stopping Rules

    Albert N. Shiryaev.Optimal Stopping Rules. Springer-Verlag, New York, 1978

  39. [39]

    Shiryaev

    Albert N. Shiryaev. Minimax optimality of the method of cumulative sums (CUSUM) in the case of continuous time.Russian Mathematical Surveys, 51(4):750–751, 1996

  40. [40]

    Shiryaev

    Albert N. Shiryaev. Quickest detection problems in the technical analysis of the financial data. InMathematical Finance—Bachelier Congress 2000: Selected Papers from the First World Congress of the Bachelier Finance Society, Paris, June 29–July 1, 2000, pages 487–521. Springer, Berlin, Heidelberg, 2002

  41. [41]

    Shiryaev

    Albert N. Shiryaev. Quickest detection problems: Fifty years later.Sequential Analysis, 29(4):345–385, 2010

  42. [42]

    David Siegmund and E. S. Venkatraman. Using the generalized likelihood ratio statistic for sequential detection of a change-point.The Annals of Statistics, 23(1):255–271, 1995

  43. [43]

    J. C. Spall. Multivariate stochastic approximation using a simultaneous perturbation gradient approximation. IEEE Transactions on Automatic Control, 37(3):332–341, 1992

  44. [44]

    Tartakovsky, Igor V

    Alexander G. Tartakovsky, Igor V. Nikiforov, and Michèle Basseville.Sequential Analysis: Hypothesis Testing and Changepoint Detection. Chapman and Hall/CRC, Boca Raton, FL, 2014

  45. [45]

    Tartakovsky and Venugopal V

    Alexander G. Tartakovsky and Venugopal V. Veeravalli. General asymptotic Bayesian theory of quickest change detection.Theory of Probability and Its Applications, 49(3):458–497, 2005

  46. [46]

    Veeravalli, and Sean Meyn

    Jayakrishnan Unnikrishnan, Venugopal V. Veeravalli, and Sean Meyn. Minimax robust quickest change detection. IEEE Transactions on Information Theory, 57(3):1604–1614, 2011

  47. [47]

    Veeravalli and Taposh Banerjee

    Venugopal V. Veeravalli and Taposh Banerjee. Quickest change detection. InAcademic Press Library in Signal Processing, volume 3, pages 209–255. Elsevier, 2014

  48. [48]

    Shanbhag

    Mingrui Wang, Prakash Chakraborty, and Uday V. Shanbhag. Improving dimension dependence in complexity guarantees for zeroth-order methods via exponentially-shifted Gaussian smoothing. In2024 Winter Simulation Conference (WSC), pages 3193–3204. IEEE, 2024

  49. [49]

    The structure of measurable mappings on metric spaces.Proceedings of the American Mathematical Society, 122(1):147–150, 1994

    Andrzej Wiśniewski. The structure of measurable mappings on metric spaces.Proceedings of the American Mathematical Society, 122(1):147–150, 1994

  50. [50]

    Robust quickest change detection for unnormalized models

    Suya Wu, Enmao Diao, Taposh Banerjee, Jie Ding, and Vahid Tarokh. Robust quickest change detection for unnormalized models. InProceedings of the 39th Conference on Uncertainty in Artificial Intelligence, volume 216 ofProceedings of Machine Learning Research, pages 2314–2323. PMLR, 2023

  51. [51]

    Veeravalli

    Liyan Xie, Yuchen Liang, and Venugopal V. Veeravalli. Distributionally robust quickest change detection using Wasserstein uncertainty sets. InProceedings of The 27th International Conference on Artificial Intelligence and Statistics, volume 238 ofProceedings of Machine Learning Research, pages 6491–6499. PMLR, 2024. AppendixA.Additional preliminaries A.1....

  52. [52]

    For a continuous stopping policyϕ∈ T, we define therandomized stopping timeby τ r ϕ := inf t≥0 : Z t∧T 0 ϕ(bX|[0,s])2 ds≥Z whereinf∅= +∞

    = 0. For a continuous stopping policyϕ∈ T, we define therandomized stopping timeby τ r ϕ := inf t≥0 : Z t∧T 0 ϕ(bX|[0,s])2 ds≥Z whereinf∅= +∞. Next we prove that stopping times can be approximated by randomized stopping times based on continuous stopping policies. Proposition B.3.LetYbe right continuous intandE[∥Y∥ ∞]<∞for any finite interval. For every s...