{"id":"101da9e0-a8e6-4d03-a2ab-a5cfb7ea4457","arxiv_id":"2506.17634","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Path signatures can be embedded in Gaussian process, deep learning, kernel, and graph diffusion models to match or beat established baselines on time series and graph benchmarks.","lead":"This mathematics thesis bundles five previously published algorithms that use path signatures, a way of encoding the shape of time-dependent data, to make scalable machine learning models for time series and graphs. A generalist might read it to see a coherent toolkit for signature-based learning, though many proofs live in the author's earlier papers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Scalability of the signature models depends on low-rank stacking preserving universality, but the thesis defers the proof to [263, App. C] and openly admits only partial results exist (Ch. 4 conclusion).","rationale":"The reader's weakest assumption is precisely that low-rank and random-feature truncations preserve signature expressiveness, and the thesis itself flags that the low-rank iteration results are only partial. The most load-bearing condition for the thesis's headline is therefore not a mathematical contradiction in the full-signature theory, but the transfer of universality from exact signatures to the scalable approximations. The thesis contains explicit evidence of the gap: Theorem 3.1 proves universality for unrestricted functionals, while the scalable Seq2Tens and Graph2Tens models use rank-1 truncated functionals; the proof that stacking recovers the lost expressiveness is deferred to an external appendix that is not included in the arXiv version, and Chapter 4's conclusion states only partial results exist. This is a genuine soft spot because the empirical claims and the scalability story would remain heuristic if the approximation theorem cannot be established. The order-1 time-coordinate issue is secondary but reinforces the concern. These are correctness risks, not fraud or inconsistency: the underlying published papers may contain the missing proofs, but the thesis as presented does not. Given that the reader's verdict is already CONDITIONAL with medium correctness risk, the appropriate recommendation is to keep that verdict unchanged rather than escalate to rejection. A conditional acceptance requiring the missing low-rank approximation proof to be supplied or precisely cited would resolve the concern.","tokens_in":61511,"tokens_out":4476,"duration_ms":51504,"concrete_test":"Independently re-derive the stacking approximation result promised in [263, App. C] from the definitions in Chapter 3 (Proposition 3.3, Theorem 3.1), and state its hypotheses and error bounds explicitly. In particular, check whether the approximation requires a time-augmented path or an unbounded number of layers or rank, and whether the error bound is uniform over a compact class of sequence functions. If the proof cannot be reconstructed from the thesis, or if the bound degrades with sequence length or dimension, the scalability claim loses its theoretical backing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that path signatures can be embedded into Gaussian processes, deep networks, kernels, and graph diffusions while remaining scalable and retaining the theoretical guarantees of signatures. The scalability mechanism in Chapters 3 and 4 replaces arbitrary linear functionals ℓ ∈ T((V)) by low-rank, truncated functionals and then stacks such layers to 'recover expressiveness.' Theorem 3.1 establishes universality only for unrestricted ℓ; the restriction to rank-1 functionals explicitly narrows the hypothesis class (Section 3.3). The claim that stacking restores approximation power is not proved in this thesis: Section 3.4 says a rigorous quantitative statement is provided in [263, App. C], but the Declarations state that Appendices A–C were deferred to the cited article; and Chapter 4's Conclusion states that 'for the iterations of low-rank approximations only partial results exist.' Thus the load-bearing assumption—that the scalable low-rank models retain the universality and expressiveness of full signatures—is only partially supported in the presented text. A related gap is that order-1 discretized signatures are universal on sequences only after adding a time coordinate (Section 1.2.6), while Algorithms 1 and 2 make time augmentation optional; if experimental configurations omit it, the stated theory does not apply.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The thesis integrates path signatures with scalable machine learning pipelines, presenting six chapters: an introduction to signatures, Gaussian processes with signature covariances (GPSig), the Seq2Tens framework for deep sequence modelling, graph models based on hypo-elliptic diffusions (G2TN), Random Fourier Signature Features, and Recurrent Sparse Spectrum Signature Gaussian Processes. Each chapter is largely self-contained and adapts previously published peer-reviewed papers. The core claim is that signature-based features can be embedded into GPs, neural networks, kernels, and graph diffusion models while retaining the theoretical guarantees of signatures and achieving scalability through low-rank approximations, random features, and sparse variational inference, leading to state-of-the-art empirical performance on several benchmarks.","tokens_in":61753,"tokens_out":2665,"duration_ms":30097,"significance":"If the central claim holds, the thesis would make a substantial contribution: signature-based models would be credible competitors to RNN, transformer, and GNN baselines for sequential and structured data, with theoretical backing that is rare in this area. The thesis has real strengths: it ships concrete algorithms with complexity analyses, provides external benchmark comparisons and ablations on multiple datasets, and presents some self-contained theoretical results, notably the regularity theorem in Chapter 2 and the concentration results in Chapter 5. The algebraic extension of signatures to graphs is a genuine innovation that connects random-walk diffusion to tensor-valued operators. However, the decisive theoretical statements that would justify the scalability claim are not fully contained in the manuscript: the universality of Seq2Tens is stated with a deferred proof, the recovery of expressiveness through stacking low-rank layers is explicitly deferred and conceded to be only partially resolved, and the characterization theorem for graphs is deferred to a separate article.","major_comments":[{"comment":"The universality theorem for Seq2Tens is stated only informally: Theorem 3.1 refers to a universal map φ with \"a lift that satisfies some mild constraints,\" and the proof is deferred to [263, App. B]. The restriction to rank-1 functionals in Section 3.3 explicitly narrows the hypothesis class, and the claim that stacking low-rank sequence-to-sequence transforms recovers expressiveness is not proved in the thesis: Section 3.4 says a rigorous quantitative statement is provided in [263, App. C], while the Declarations state that Appendices A–C were omitted and deferred to that article. Because the scalability of the entire model family rests on this step, the manuscript should either include a precise statement and proof of the stacking result or clearly mark this as an open conjecture.","section":"§3.3, Theorem 3.1 and §3.4"},{"comment":"The Chapter 4 conclusion admits that \"for the iterations of low-rank approximations only partial results exist.\" This is load-bearing for the graph contribution: Theorem 4.3 gives an efficient algorithm for rank-1 functionals, but the claim that composing such layers achieves the expressiveness of general high-degree functionals is used to motivate the G2TN architecture and is not established here. The characterization result for graphs (Theorem 4.2) is also deferred to [265, App. E]. The text should either supply the missing proof or explicitly frame the expressiveness of stacked low-rank layers as empirical rather than theoretical.","section":"§4.6, Conclusion"},{"comment":"The universality of order-1 discretized signatures is stated to require a time coordinate: Section 1.2.6 says the p=1 case is universal on sequences \"given the existence a time coordinate which encodes the position within the sequence.\" However, Algorithms 1 and 2 make time augmentation optional, and the experiments in Chapters 2, 3, and 6 do not consistently state whether this coordinate was included. If any experiment omitted the time coordinate, the stated theoretical guarantees (universality, and the characterization results that rely on it for graphs) do not apply to that configuration. The manuscript should state explicitly, for each experimental setup, whether time augmentation was used.","section":"§1.2.6 and Algorithms 1–2"}],"minor_comments":[{"comment":"In Definition 1.2, the sentence \"U× V is unique up to isomorphism\" should refer to the tensor product U⊗V rather than the product set U× V.","section":"§1.1.1"},{"comment":"Algorithms 1 and 2 contain duplicated line numbers (e.g., multiple lines labelled 9 and 10 in Algorithm 1), which makes the listings hard to follow.","section":"§1.3.2"},{"comment":"There is a typo in \"polynomail complexity\" that should read \"polynomial complexity.\"","section":"§1.3.2"},{"comment":"In Section 2.2.1, \"nuiscance function\" should be \"nuisance function.\"","section":"§2.2.1"},{"comment":"In Section 2.4.1, \"choosen\" should be \"chosen,\" and later in the same section \"maximising\" is inconsistently spelled.","section":"§2.4.1"},{"comment":"The thesis would benefit from a table summarizing the computational complexity of all proposed methods in one place; currently the complexity statements are scattered across Chapters 2–6.","section":"§5 and Chapter 6"}],"recommendation":"major_revision","confidential_remarks":"The thesis is a compilation of the author's published papers, which likely explains why several central proofs are deferred to prior publications. For a journal submission, I would weigh this more heavily than for a thesis defense: the arXiv version should either include the missing proofs in appendices or explicitly state which results are only available in the referenced works. The reader should also be told whether the reported experiments used time augmentation, since the stated theory depends on it for the order-1 features."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a PhD thesis, not a new research announcement. Its six chapters reprint five of the author's peer-reviewed papers (ICML, ICLR, NeurIPS, SIAM MDS, AISTATS), plus an introductory chapter and one combined algorithm. If you are looking for a new central theorem, you won't find it here. If you need a readable and well-organized reference for signature-based machine learning, this is a good one.\n\nThe thesis does several things well. Chapter 1 gives a genuinely useful introduction to tensor algebra and path signatures, including the discrete-time details that are often skipped in papers. The algorithmic chapters are self-contained enough that a motivated graduate student could implement the core ideas, and the author points to working code for the GP and kernel parts (GPSig, KSig). The experimental chapters report real benchmarks against serious baselines, and the author is honest in several places about what is not known—for example, Chapter 4's conclusion admits that only partial results exist for low-rank approximations. That honesty is worth crediting.\n\nThe main soft spot is the load-bearing claim about scalability preserving expressiveness. The thesis argues that low-rank functionals stacked in layers recover the universality of unrestricted signature features, but the proof is deferred to an appendix of a previous paper, and the thesis itself concedes the results are partial. Similarly, order-1 discretized signatures are universal on sequences only with a time coordinate, yet the provided algorithms make time augmentation optional. A careful reader should not take the scaling story as fully proven—it is a plausible conjecture backed by strong empirical evidence, but the theory is incomplete. The phrase 'often leading to outperforming state-of-the-art methods' overstates the benchmark results a little; in several tables the proposed models are competitive rather than uniformly superior.\n\nThere is also the obvious point that this is a compilation, so novelty relative to the published literature is low. A referee should not expect new results.\n\nWho is this for? Graduate students and researchers who want an accessible entry point into signature methods for time series and graphs, and practitioners who want a survey of the models. As a teaching or reference document it has real value. If it were submitted to a journal as a survey or monograph, it deserves a proper review. As a research paper claiming new results, it would need to be repositioned.\n\nMy recommendation: engage with it as a reference, but do not treat the scalability proofs as settled. It deserves peer review in an appropriate venue, provided expectations are set to 'review of a thesis or monograph' rather than 'novel research.'","headline":"A well-written thesis compiling five published papers, with a clear intro to path signatures but no new central result; the scalability theory is honestly flagged as partial.","tokens_in":62286,"tokens_out":4063,"would_cite":true,"duration_ms":40048,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60L10","68T07","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Path signatures, once too costly for machine learning, can be embedded into scalable Gaussian processes, deep networks, and graph diffusions while retaining their universal approximation guarantees.","keywords":["path signatures","iterated integrals","tensor algebra","Gaussian processes","random Fourier features","graph neural networks","time series forecasting","low-rank approximation"],"falsifier":"Take a sequence classification task whose labels change under time reparameterization and evaluate order-1 discretized signature features without the added time coordinate: if predictions stay identical for all reparameterized inputs, the claimed universality of those features fails. Alternatively, on a large-scale sequence dataset, compare the Random Fourier Signature Feature kernel with the exact signature kernel: if increasing the number of random features does not drive the approximation error toward zero as the concentration results predict, the central scalability claim is undermined.","tokens_in":61288,"feed_emoji":"📈","tokens_out":4806,"duration_ms":52968,"temperature":0.7,"pith_summary":"This thesis argues that path signatures—iterated integrals that encode the full history of a path—can be turned from a theoretically attractive but computationally prohibitive feature map into a practical foundation for scalable machine learning. The author embeds signatures into Gaussian processes, low-rank deep sequence layers, graph diffusion models, random Fourier features, and a recurrent sparse-spectrum Gaussian process forecaster. If the central claim is right, signature-based models become a serious alternative to recurrent, transformer, and graph-neural-network baselines for sequential and structured data, while keeping properties such as universality and reparameterization invariance. The unifying bet is that the combinatorial cost of signatures can be bypassed with low-rank and random-feature truncations without giving up their expressive power.","feed_headline":"Path signatures scale to GPs, deep nets, and graphs","feed_subtitle":"One universal feature map gives calibrated time-series models, long-range graph features, and fast kernel approximations.","key_machinery":"The central object is the path signature $S(x) = (1, S^1(x), S^2(x), \\ldots)$, the sequence of iterated integrals of a path, together with the tensor algebra $T((V))$ whose non-commutative product stitches together increments. In the discrete setting the order-$p$ signature features sum over subsequences of increments, generalizing string kernels. This machinery does three kinds of work: it gives a universal feature map for sequences, it defines a kernel by inner products in the tensor algebra, and it supports cheap rank-1 linear functionals that make the feature map computable without ever forming the full signature tensor.","core_discovery":"The central claim is that one mathematical object, the path signature, can serve as a common algebraic backbone for several scalable machine-learning pipelines. Signature inner products define Gaussian-process covariances; iterating rank-1 tensor functionals of signatures defines the deep sequence layer Seq2Tens; a tensor-valued hypo-elliptic Laplacian turns graph diffusion into a mechanism that summarizes random-walk histories; random Fourier projections approximate the signature kernel with concentration guarantees; and a decay parameter in the same feature space gives a forgetting mechanism for probabilistic forecasting. The thesis reports that these signature-based models are consistently competitive with strong baselines across time-series classification, mortality prediction, generative imputation, long-range graph tasks, and multi-horizon forecasting, and often outperform them.","pith_inferences":["The order-1 discretized signature is universal only when a time coordinate is added; a practical recipe that follows from the thesis is to always include such a coordinate, otherwise the theoretical guarantees do not transfer to low-rank or random-feature models.","If the low-rank and random-feature truncations prove as expressive as the full signature in further tests, signature-based features could become a default first step for sequence and graph data, analogous to how polynomial features are used in classical regression.","The hypo-elliptic diffusion view connects graph learning to sub-Riemannian geometry, suggesting that over-squashing could be studied through the geometry of lifted random walks rather than only through message-passing depth, a direction the thesis leaves open.","Random Fourier Signature Features could be used beyond prediction, for example to scale signature-based maximum mean discrepancy tests for the distribution of paths to large sample sizes, which would extend the thesis's claims without requiring new theory."],"forward_implications":["Signature covariances give Gaussian processes calibrated uncertainty on time-series classification, with the signature GP ranking ahead of other GP baselines and competitive with frequentist classifiers on accuracy.","Low-rank signature layers (Seq2Tens) can be grafted onto existing convolutional and variational-autoencoder models, improving accuracy, mortality prediction, and missing-data imputation.","The hypo-elliptic graph Laplacian yields graph and node features that characterize random-walk history, improving long-range graph classification without global attention or quadratic node interactions.","Random Fourier Signature Features reduce the signature kernel's quadratic cost in sequence length and sample size, extending signature methods to datasets with millions of sequences.","Recurrent Sparse Spectrum Signature Gaussian Processes provide scalable, probabilistic multi-horizon forecasts with an adaptive context length that outperforms plain Gaussian processes and competes with deep-learning forecasters."],"supporting_citations":[{"why":"Introduces the kernel trick for signature inner products, the computational foundation for signature kernels and signature covariances.","marker":"[146]"},{"why":"Seq2Tens paper; proves universality of discretized order-1 signature features with a time coordinate and motivates low-rank iterations.","marker":"[263]"},{"why":"Formalizes hypo-elliptic diffusion on graphs and the connection between expected signatures and random-walk histories.","marker":"[265]"},{"why":"Develops Random Fourier Signature Features with concentration guarantees, the basis of the scalable kernel approximation.","marker":"[267]"},{"why":"Introduces Recurrent Sparse Spectrum Signature Gaussian Processes with a forgetting mechanism for forecasting.","marker":"[262]"},{"why":"Seminal result on the hypo-elliptic Laplacian for path-dependent Brownian motion, the mathematical inspiration for the graph diffusion.","marker":"[93]"},{"why":"Survey of rough path theory supplying the signature's analytic properties: factorial decay, Chen identity, and universality.","marker":"[182]"},{"why":"PDE-based algorithm for the untruncated signature kernel, an alternative to dynamic programming that the thesis builds on.","marker":"[228]"},{"why":"Provides proofs of universality for signature features and the metric/statistical properties used to characterize paths and graphs.","marker":"[44]"}],"fun_headline_variants":["One signature feature map powers GP, deep, and graph models","Path signatures: scalable kernel, deep, and graph learning","Universal path signature features scale to GP and deep nets","Signature-based models scale to time series, graphs, and forecasting","From rough paths to scalable ML: one feature map for all"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole scalability story rests on the assumption that low-rank and random-feature truncations of the signature preserve enough of its universal expressive power; the thesis itself concedes that only partial results exist for iterations of low-rank approximations, and that order-1 discretized signatures need a time coordinate to be universal.","fun_headline_variants_meta":{"raw":{"variants":["One signature feature map powers GP, deep, and graph models","Path signatures: scalable kernel, deep, and graph learning","Universal path signature features scale to GP and deep nets","Signature-based models scale to time series, graphs, and forecasting","From rough paths to scalable ML: one feature map for all"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1304,"prompt_tokens":961,"completion_tokens":343,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":260}},"tokens_in":577,"tokens_out":343,"duration_ms":3944,"temperature":1.0,"reasoning_tokens":260,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:05:15.356979+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sequence classification task whose labels change under time reparameterization and evaluate order-1 discretized signature features without the added time coordinate: if predictions stay identical for all reparameterized inputs, the claimed universality of those features fails. Alternatively, on a large-scale sequence dataset, compare the Random Fourier Signature Feature kernel with the exact signature kernel: if increasing the number of random features does not drive the approximation error toward zero as the concentration results predict, the central scalability claim is undermined.","supporting_citations":[{"cited_title":"Seq2Tens: An efficient representa- tion of sequences by low-rank tensor projections","cited_arxiv_id":null,"evidence_quote":"Seq2Tens paper; proves universality of discretized order-1 signature features with a time coordinate and motivates low-rank iterations."},{"cited_title":"Capturing graphs with hypo-elliptic diffusions","cited_arxiv_id":null,"evidence_quote":"Formalizes hypo-elliptic diffusion on graphs and the connection between expected signatures and random-walk histories."}],"review_version":2}