Pith. sign in

REVIEW 3 major objections 6 minor 298 references

Scalable Machine Learning Algorithms using Path Signatures

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Path signatures, once too costly for machine learning, can be embedded into scalable Gaussian processes, deep networks, and graph diffusions while retaining their universal approximation guarantees.

desk verdict A well-written thesis compiling five published papers, with a clear intro to path signatures but no new central result; the scalability theory is honestly flagged as partial. read the letter →

arxiv 2506.17634 v2 pith:5DKNQZSR submitted 2025-06-21 stat.ML cs.LGmath.PR

classification stat.MLcs.LGmath.PR MSC 60L1068T0768T05
keywords pathsignaturesiteratedintegralstensoralgebraGaussianprocessesrandomFourierfeaturesgraphneuralnetworkstimeseriesforecastinglow-rankapproximation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This thesis argues that path signatures—iterated integrals that encode the full history of a path—can be turned from a theoretically attractive but computationally prohibitive feature map into a practical foundation for scalable machine learning. The author embeds signatures into Gaussian processes, low-rank deep sequence layers, graph diffusion models, random Fourier features, and a recurrent sparse-spectrum Gaussian process forecaster. If the central claim is right, signature-based models become a serious alternative to recurrent, transformer, and graph-neural-network baselines for sequential and structured data, while keeping properties such as universality and reparameterization invariance. The unifying bet is that the combinatorial cost of signatures can be bypassed with low-rank and random-feature truncations without giving up their expressive power.

What carries the argument

The central object is the path signature $S(x) = (1, S^1(x), S^2(x), \ldots)$, the sequence of iterated integrals of a path, together with the tensor algebra $T((V))$ whose non-commutative product stitches together increments. In the discrete setting the order-$p$ signature features sum over subsequences of increments, generalizing string kernels. This machinery does three kinds of work: it gives a universal feature map for sequences, it defines a kernel by inner products in the tensor algebra, and it supports cheap rank-1 linear functionals that make the feature map computable without ever forming the full signature tensor.

What would settle it

Take a sequence classification task whose labels change under time reparameterization and evaluate order-1 discretized signature features without the added time coordinate: if predictions stay identical for all reparameterized inputs, the claimed universality of those features fails. Alternatively, on a large-scale sequence dataset, compare the Random Fourier Signature Feature kernel with the exact signature kernel: if increasing the number of random features does not drive the approximation error toward zero as the concentration results predict, the central scalability claim is undermined.

Watch

Extended reading notes

Core claim

The central claim is that one mathematical object, the path signature, can serve as a common algebraic backbone for several scalable machine-learning pipelines. Signature inner products define Gaussian-process covariances; iterating rank-1 tensor functionals of signatures defines the deep sequence layer Seq2Tens; a tensor-valued hypo-elliptic Laplacian turns graph diffusion into a mechanism that summarizes random-walk histories; random Fourier projections approximate the signature kernel with concentration guarantees; and a decay parameter in the same feature space gives a forgetting mechanism for probabilistic forecasting. The thesis reports that these signature-based models are consistently competitive with strong baselines across time-series classification, mortality prediction, generative imputation, long-range graph tasks, and multi-horizon forecasting, and often outperform them.

Load-bearing premise

The whole scalability story rests on the assumption that low-rank and random-feature truncations of the signature preserve enough of its universal expressive power; the thesis itself concedes that only partial results exist for iterations of low-rank approximations, and that order-1 discretized signatures need a time coordinate to be universal.

Editorial extensions

If this is right

  • Signature covariances give Gaussian processes calibrated uncertainty on time-series classification, with the signature GP ranking ahead of other GP baselines and competitive with frequentist classifiers on accuracy.
  • Low-rank signature layers (Seq2Tens) can be grafted onto existing convolutional and variational-autoencoder models, improving accuracy, mortality prediction, and missing-data imputation.
  • The hypo-elliptic graph Laplacian yields graph and node features that characterize random-walk history, improving long-range graph classification without global attention or quadratic node interactions.
  • Random Fourier Signature Features reduce the signature kernel's quadratic cost in sequence length and sample size, extending signature methods to datasets with millions of sequences.
  • Recurrent Sparse Spectrum Signature Gaussian Processes provide scalable, probabilistic multi-horizon forecasts with an adaptive context length that outperforms plain Gaussian processes and competes with deep-learning forecasters.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The order-1 discretized signature is universal only when a time coordinate is added; a practical recipe that follows from the thesis is to always include such a coordinate, otherwise the theoretical guarantees do not transfer to low-rank or random-feature models.
  • If the low-rank and random-feature truncations prove as expressive as the full signature in further tests, signature-based features could become a default first step for sequence and graph data, analogous to how polynomial features are used in classical regression.
  • The hypo-elliptic diffusion view connects graph learning to sub-Riemannian geometry, suggesting that over-squashing could be studied through the geometry of lifted random walks rather than only through message-passing depth, a direction the thesis leaves open.
  • Random Fourier Signature Features could be used beyond prediction, for example to scale signature-based maximum mean discrepancy tests for the distribution of paths to large sample sizes, which would extend the thesis's claims without requiring new theory.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The thesis integrates path signatures with scalable machine learning pipelines, presenting six chapters: an introduction to signatures, Gaussian processes with signature covariances (GPSig), the Seq2Tens framework for deep sequence modelling, graph models based on hypo-elliptic diffusions (G2TN), Random Fourier Signature Features, and Recurrent Sparse Spectrum Signature Gaussian Processes. Each chapter is largely self-contained and adapts previously published peer-reviewed papers. The core claim is that signature-based features can be embedded into GPs, neural networks, kernels, and graph diffusion models while retaining the theoretical guarantees of signatures and achieving scalability through low-rank approximations, random features, and sparse variational inference, leading to state-of-the-art empirical performance on several benchmarks.

Significance. If the central claim holds, the thesis would make a substantial contribution: signature-based models would be credible competitors to RNN, transformer, and GNN baselines for sequential and structured data, with theoretical backing that is rare in this area. The thesis has real strengths: it ships concrete algorithms with complexity analyses, provides external benchmark comparisons and ablations on multiple datasets, and presents some self-contained theoretical results, notably the regularity theorem in Chapter 2 and the concentration results in Chapter 5. The algebraic extension of signatures to graphs is a genuine innovation that connects random-walk diffusion to tensor-valued operators. However, the decisive theoretical statements that would justify the scalability claim are not fully contained in the manuscript: the universality of Seq2Tens is stated with a deferred proof, the recovery of expressiveness through stacking low-rank layers is explicitly deferred and conceded to be only partially resolved, and the characterization theorem for graphs is deferred to a separate article.

major comments (3)
  1. [§3.3, Theorem 3.1 and §3.4] The universality theorem for Seq2Tens is stated only informally: Theorem 3.1 refers to a universal map φ with "a lift that satisfies some mild constraints," and the proof is deferred to [263, App. B]. The restriction to rank-1 functionals in Section 3.3 explicitly narrows the hypothesis class, and the claim that stacking low-rank sequence-to-sequence transforms recovers expressiveness is not proved in the thesis: Section 3.4 says a rigorous quantitative statement is provided in [263, App. C], while the Declarations state that Appendices A–C were omitted and deferred to that article. Because the scalability of the entire model family rests on this step, the manuscript should either include a precise statement and proof of the stacking result or clearly mark this as an open conjecture.
  2. [§4.6, Conclusion] The Chapter 4 conclusion admits that "for the iterations of low-rank approximations only partial results exist." This is load-bearing for the graph contribution: Theorem 4.3 gives an efficient algorithm for rank-1 functionals, but the claim that composing such layers achieves the expressiveness of general high-degree functionals is used to motivate the G2TN architecture and is not established here. The characterization result for graphs (Theorem 4.2) is also deferred to [265, App. E]. The text should either supply the missing proof or explicitly frame the expressiveness of stacked low-rank layers as empirical rather than theoretical.
  3. [§1.2.6 and Algorithms 1–2] The universality of order-1 discretized signatures is stated to require a time coordinate: Section 1.2.6 says the p=1 case is universal on sequences "given the existence a time coordinate which encodes the position within the sequence." However, Algorithms 1 and 2 make time augmentation optional, and the experiments in Chapters 2, 3, and 6 do not consistently state whether this coordinate was included. If any experiment omitted the time coordinate, the stated theoretical guarantees (universality, and the characterization results that rely on it for graphs) do not apply to that configuration. The manuscript should state explicitly, for each experimental setup, whether time augmentation was used.
minor comments (6)
  1. [§1.1.1] In Definition 1.2, the sentence "U× V is unique up to isomorphism" should refer to the tensor product U⊗V rather than the product set U× V.
  2. [§1.3.2] Algorithms 1 and 2 contain duplicated line numbers (e.g., multiple lines labelled 9 and 10 in Algorithm 1), which makes the listings hard to follow.
  3. [§1.3.2] There is a typo in "polynomail complexity" that should read "polynomial complexity."
  4. [§2.2.1] In Section 2.2.1, "nuiscance function" should be "nuisance function."
  5. [§2.4.1] In Section 2.4.1, "choosen" should be "chosen," and later in the same section "maximising" is inconsistently spelled.
  6. [§5 and Chapter 6] The thesis would benefit from a table summarizing the computational complexity of all proposed methods in one place; currently the complexity statements are scattered across Chapters 2–6.

Circularity Check

1 steps flagged · score 4.0 of 10

Scalability guarantee for stacked low-rank signature layers is deferred to a coauthored appendix and admitted to be partial; empirical benchmarks remain external.

  1. self citation load bearing [Section 3.4, with supporting statements in the Declarations and Chapter 4 Conclusion]
    "Making precise how the stacking of such low-rank seq2seq transformations approximates general functions requires more tools from algebra, and we provide a rigorous quantitative statement in [263, App. C]. Here, we just appeal to the analogy made with adding depth in neural networks mentioned earlier and empirically validate this in our experiments in Section 3.4."

    The thesis's scalable signature models depend on the claim that stacking rank-1/low-rank functionals recovers the universal expressiveness that is explicitly lost by the rank restriction in Section 3.3. The only proof support offered for this load-bearing claim is [263, App. C], a paper written by the thesis author and collaborators; the Declarations state that Appendices A-C were written by Patric and 'omitted here and deferred to the article,' so the proof is not independently presented or verified in this thesis.

full rationale

The core empirical claims are tested against external benchmarks (UCR time series, NCI graphs, PhysioNet), and no fitted parameter or learned hyperparameter is renamed as a prediction; the benchmark results are not constructed from the target quantities. However, the theoretical premise that low-rank stacking preserves universal expressiveness is load-bearing for the scalability story and is not proven in the thesis: Section 3.3 restricts the hypothesis class to rank-1 functionals, Section 3.4 defers the quantitative stacking statement to [263, App. C], the Declarations state that the relevant appendices were written by a collaborator and omitted from the thesis, and the Chapter 4 Conclusion admits only partial results exist for iterations of low-rank approximations. Because [263] is the author's own coauthored ICLR paper, this is a load-bearing self-citation rather than an independently verified mathematical fact. That warrants a score of 4: partial circularity in the theoretical foundation, but the central empirical demonstrations remain self-contained and externally evaluated.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The ledger lists the main hyperparameters and background assumptions the central claims rely on: signature-level weights, time-augmentation scale, kernel bandwidth, truncation degree, random feature counts, and the assumption that low-rank truncations preserve expressiveness. These are not hidden physical entities; they are modeling choices that should be made explicit.

free parameters (5)
  • signature level weights sigma_m^2 = learned
    Ch 2 Eq (2.6) and (2.7) weight each tensor level in the covariance; these are optimized during GP training, not derived.
  • time-augmentation scale tau = learned via ARD
    Ch 2 Sec 2.3.1 treats tau as a hyperparameter controlling parameterization invariance; its value is inferred from data.
  • static kernel bandwidth = chosen by validation
    RBF bandwidth for the kernel lift is a validation-selected hyperparameter throughout the thesis.
  • tensor truncation degree M = M=4 in Chapter 2, 2-4 in Chapter 3, 2 in Chapter 4
    The truncation degree is a hand or validation-selected hyperparameter; Chapter 2 adopted M=4 as an optimal trade-off.
  • number of random features = not specified in available text
    Random Fourier Signature Features (Ch 5) and RS3GP (Ch 6) depend on the number of random features, which is chosen experimentally.
assumptions (6)
  • standard math Chen's identity and shuffle identity for path signatures
    Used throughout Chapters 1-4 to factor signatures of concatenated and reparameterized paths.
  • standard math Universality of signature linear functionals via Stone-Weierstrass
    Invoked in Chapter 1 Prop 1.19 and Chapter 3 Theorem 3.1.
  • domain assumption Order-1 discretized signature is universal only with an added time coordinate
    Stated in Chapter 1 Sec 1.2.6 and credited to [263]; many models use order-1 features.
  • standard math Kernel trick for signature kernels and its PDE representation
    Foundation for signature Gaussian processes and random feature approximations, following [146] and [228].
  • standard math Concentration inequalities for random Fourier features
    Chapter 5 relies on subexponential and alpha-exponential concentration; these are standard but not proved in full in the available text.
  • domain assumption Low-rank truncations preserve the approximation properties of signatures
    Chapter 4 conclusion admits only partial theoretical results for iterations of low-rank approximations; the scalability claim depends on this.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Scalable Machine Learning Algorithms using Path Signatures." pith.science (2026). https://pith.science/paper/5DKNQZSR

@misc{pith2026250617634,
  author       = {Pith},
  title        = {Pith review of: Scalable Machine Learning Algorithms using Path Signatures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5DKNQZSR}},
  note         = {Machine review of arXiv:2506.17634}
}
read the original abstract

The interface between stochastic analysis and machine learning is a rapidly evolving field, with path signatures - iterated integrals that provide faithful, hierarchical representations of paths - offering a principled and universal feature map for sequential and structured data. Rooted in rough path theory, path signatures are invariant to reparameterization and well-suited for modelling evolving dynamics, long-range dependencies, and irregular sampling - common challenges in real-world time series and graph data. This thesis investigates how to harness the expressive power of path signatures within scalable machine learning pipelines. It introduces a suite of models that combine theoretical robustness with computational efficiency, bridging rough path theory with probabilistic modelling, deep learning, and kernel methods. Key contributions include: Gaussian processes with signature kernel-based covariance functions for uncertainty-aware time series modelling; the Seq2Tens framework, which employs low-rank tensor structure in the weight space for scalable deep modelling of long-range dependencies; and graph-based models where expected signatures over graphs induce hypo-elliptic diffusion processes, offering expressive yet tractable alternatives to standard graph neural networks. Further developments include Random Fourier Signature Features, a scalable kernel approximation with theoretical guarantees, and Recurrent Sparse Spectrum Signature Gaussian Processes, which combine Gaussian processes, signature kernels, and random features with a principled forgetting mechanism for multi-horizon time series forecasting with adaptive context length. We hope this thesis serves as both a methodological toolkit and a conceptual bridge, and provides a useful reference for the current state of the art in scalable, signature-based learning for sequential and structured data.

Figures

Figures reproduced from arXiv: 2506.17634 by the authors.

Figure 2.1
Figure 2.1. Box-plots of negative log-predictive probabilities (left) and classification accuracies [PITH_FULL_IMAGE:figures/full_fig_p045_2_1.png] view at source ↗
Figure 2.4
Figure 2.4. Achieved ELBO (top), accuracy (middle), mean NLPP (bottom) after 300 epochs of training the variational parameters with random initialization and frozen pre-learnt kernel hyperparameters. Solid is the mean over 5 independent runs, dashed is the 1-std region. classes hardly look linearly separable; (ii) the tensors, however, seem to live on a completely separate subspace than the data. The algorithm achieves 92% accu… view at source ↗
Figure 2.5
Figure 2.5. A UMAP visualization of the features corresponding to data-points (coloured), and inducing tensors (black) in the feature space on the AUSLAN dataset. 2.6 Conclusion We used a classical object from stochastic analysis – signatures – to define a GP for sequential data. The GP inherits many of the theoretical guarantees that are known for signature features such as universality and parameterization invariance. To make… view at source ↗
Figures from the paper (8 more)
Figure 3.1
Figure 3.1. Figure 3.1: Box-plot of classification accuracies (left) and critical difference diagram (right). [PITH_FULL_IMAGE:figures/full_fig_p064_3_1.png]
Figure 4.1
Figure 4.1. Figure 4.1: Visualization of the architecture used for NCI1 and NCI109 described in Section 4.5. [PITH_FULL_IMAGE:figures/full_fig_p098_4_1.png]
Figure 5.1
Figure 5.1. Figure 5.1: Approximation error of random kernels against [PITH_FULL_IMAGE:figures/full_fig_p121_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: Critical difference diagram com￾parison on moderate datasets of considered ap￾proaches using two-tailed Nemenyi test [66] [PITH_FULL_IMAGE:figures/full_fig_p125_5_2.png]
Figure 6
Figure 6. Figure 6: qualitatively illustrates the predictive mean and uncertainty [PITH_FULL_IMAGE:figures/full_fig_p169_6.png]
Figure 6.1
Figure 6.1. Figure 6.1: Predictive mean and uncertainty on a toy dataset composed of multi-sinusoidal [PITH_FULL_IMAGE:figures/full_fig_p170_6_1.png]
Figure 6.2
Figure 6.2. Figure 6.2: Reparameterizing the beta distribution for various shape parameters given fixed [PITH_FULL_IMAGE:figures/full_fig_p175_6_2.png]
Figure 6.3
Figure 6.3. Figure 6.3: Computation time (left) and memory consumption (right) of [PITH_FULL_IMAGE:figures/full_fig_p177_6_3.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

298 extracted references · 48 canonical work pages

  1. [1]

    https://pubchem.ncbi.nlm.nih.gov/

    Pubchem. https://pubchem.ncbi.nlm.nih.gov/

  2. [2]

    Bayesian model selection of lithium-ion battery models via Bayesian quadrature

    Masaki Adachi, Yannick Kuhn, Birger Horstmann, Arnulf Latz, Michael A Osborne, and David A Howey . Bayesian model selection of lithium-ion battery models via Bayesian quadrature. IFAC-PapersOnLine, 56(2):10521–10526, 2023

  3. [3]

    Adler and J.E

    R.J. Adler and J.E. Taylor. Random Fields and Geometry. Springer Monographs in Math- ematics. Springer New York, 2009

  4. [4]

    Predicting battery end of life from solar off-grid system field data using machine learning

    Antti Aitio and David A Howey . Predicting battery end of life from solar off-grid system field data using machine learning. Joule, 5(12):3204–3220, 2021

  5. [5]

    Learning scalable deep kernels with recurrent structure

    Maruan Al-Shedivat, Andrew G Wilson, Yunus Saatchi, Zhiting Hu, and Eric P Xing. Learning scalable deep kernels with recurrent structure. Journal of Machine Learning Research (JMLR), 18(82):1–37, 2017

  6. [6]

    GluonTS: Probabilistic and neural time series mod- eling in Python

    Alexander Alexandrov, Konstantinos Benidis, Michael Bohlke-Schneider, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, Danielle C Maddix, Syama Rangapuram, David Salinas, Jasper Schulz, et al. GluonTS: Probabilistic and neural time series mod- eling in Python. Journal of Machine Learning Research (JMLR), 21(116):1–6, 2020

  7. [7]

    On the bottleneck of graph neural networks and its practical implications

    Uri Alon and Eran Yahav. On the bottleneck of graph neural networks and its practical implications. In International Conference on Learning Representations, 2021

  8. [8]

    Signature-based valida- tion of real-world economic scenarios.ASTIN Bulletin: The Journal of the IAA, 54(2):410– 440, 2024

    Hervé Andrès, Alexandre Boumezoued, and Benjamin Jourdain. Signature-based valida- tion of real-world economic scenarios.ASTIN Bulletin: The Journal of the IAA, 54(2):410– 440, 2024

Show all 298 references
  1. [9]

    UCI machine learning repository , 2007

    Arthur Asuncion, David Newman, et al. UCI machine learning repository , 2007

  2. [10]

    Random Fourier features for kernel ridge regression: Approximation 175 bounds and statistical guarantees

    Haim Avron, Michael Kapralov, Cameron Musco, Christopher Musco, Ameya Velingker, and Amir Zandieh. Random Fourier features for kernel ridge regression: Approximation 175 bounds and statistical guarantees. InInternational Conference on Machine Learning, pages 253–262, 2017

  3. [11]

    Layer normalization

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016

  4. [12]

    Sharp analysis of low-rank kernel matrix approximations

    Francis Bach. Sharp analysis of low-rank kernel matrix approximations. InConference on Learning Theory, pages 185–209, 2013

  5. [13]

    Neural machine translation by jointly learning to align and translate

    Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015, 2015

  6. [14]

    Ex- ploiting the past and the future in protein secondary structure prediction.Bioinformatics, 15(11):937–946, 1999

    Pierre Baldi, Søren Brunak, Paolo Frasconi, Giovanni Soda, and Gianluca Pollastri. Ex- ploiting the past and the future in protein secondary structure prediction.Bioinformatics, 15(11):937–946, 1999

  7. [15]

    Interaction networks for learning about objects, relations and physics.Advances in neural information processing systems, 29, 2016

    Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics.Advances in neural information processing systems, 29, 2016

  8. [16]

    Understanding prob- abilistic sparse Gaussian process approximations

    Matthias Bauer, Mark van der Wilk, and Carl Edward Rasmussen. Understanding prob- abilistic sparse Gaussian process approximations. In Advances in neural information pro- cessing systems, pages 1533–1541, 2016

  9. [17]

    Multivariate Time Series Classification Datasets

    Mustafa Baydogan. Multivariate Time Series Classification Datasets . http:// mustafabaydogan.com, 2015. [Accessed: 2020-02-05]

  10. [18]

    Learning a symbolic representation for multivariate time series classification

    Mustafa Gokce Baydogan and George Runger. Learning a symbolic representation for multivariate time series classification. Data Mining and Knowledge Discovery, 29(2):400– 422, 2015

  11. [19]

    Mustafa Gokce Baydogan and George C. Runger. Time series representation and simi- larity based on local autopatterns. Data Mining and Knowledge Discovery , 30:476–509, 2015

  12. [20]

    A Bayesian wilcoxon signed-rank test based on the Dirichlet process

    Alessio Benavoli, Giorgio Corani, Francesca Mangili, Marco Zaffalon, and Fabrizio Rug- geri. A Bayesian wilcoxon signed-rank test based on the Dirichlet process. InInternational conference on machine learning, pages 1026–1034, 2014

  13. [21]

    Reproducing kernel Hilbert spaces in proba- bility and statistics

    Alain Berlinet and Christine Thomas-Agnan. Reproducing kernel Hilbert spaces in proba- bility and statistics. Springer Science & Business Media, 2011. 176

  14. [22]

    Deep signature transforms

    Patric Bonnier, Patrick Kidger, Imanol Perez Arribas, Cristopher Salvi, and Terry Lyons. Deep signature transforms. 33rd Conference on Neural Information Processing Systems, NeurIPS, 2019

  15. [23]

    Graph kernels: State-of-the-art and future challenges

    Karsten Borgwardt, Elisabetta Ghisu, Felipe Llinares-López, Leslie O’Bray , and Bas- tian Rieck. Graph kernels: State-of-the-art and future challenges. arXiv preprint arXiv:2011.03854, 2020

  16. [24]

    Concentration Inequalities: A Nonasymptotic Theory of Independence

    Stéphane Boucheron, Gábor Lugosi, and Pascal Massart. Concentration Inequalities: A Nonasymptotic Theory of Independence. Oxford University Press, 2013

  17. [25]

    Geometric deep learning: Grids, groups, graphs, geodesics, and gauges.arXiv preprint arXiv:2104.13478, 2021

    Michael M Bronstein, Joan Bruna, Taco Cohen, and Petar Veli ˇckovi´c. Geometric deep learning: Grids, groups, graphs, geodesics, and gauges.arXiv preprint arXiv:2104.13478, 2021

  18. [26]

    Spectral networks and locally connected networks on graphs

    Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun. Spectral networks and locally connected networks on graphs. arXiv preprint arXiv:1312.6203, 2013

  19. [27]

    Generating financial markets with signatures

    H Buehler, B Horvath, T Lyons, I Perez, and B Wood. Generating financial markets with signatures. Risk.net, 2021

  20. [28]

    Time series forecasting for healthcare diagnosis and prognostics with the focus on cardiovascular diseases

    C Bui, N Pham, A Vo, A Tran, A Nguyen, and T Le. Time series forecasting for healthcare diagnosis and prognostics with the focus on cardiovascular diseases. In International Conference on the Development of Biomedical Engineering in Vietnam (BME) , pages 809–

  21. [29]

    Bui, Josiah Yan, and Richard E

    Thang D. Bui, Josiah Yan, and Richard E. Turner. A Unifying Framework for Gaussian Process Pseudo-Point Approximations using Power Expectation Propagation. Journal of Machine Learning Research, 18(104):1–72, 2017

  22. [30]

    Learning with SGD and random features

    Luigi Carratino, Alessandro Rudi, and Lorenzo Rosasco. Learning with SGD and random features. In Advances in Neural Information Processing Systems , pages 10192–10203, 2018

  23. [31]

    Eckart-young

    J Douglas Carroll and Jih-Jie Chang. Analysis of individual differences in multidimen- sional scaling via an n-way generalization of “Eckart-young” decomposition. Psychome- trika, 35(3):283–319, 1970

  24. [32]

    Weighted signature kernels

    Thomas Cass, Terry Lyons, and Xingcheng Xu. Weighted signature kernels. The Annals of Applied Probability, 34(1A):585–626, 2024. 177

  25. [33]

    Lecture notes on rough paths and applications to machine learning

    Thomas Cass and Cristopher Salvi. Lecture notes on rough paths and applications to machine learning. arXiv preprint arXiv:2404.06583, 2024

  26. [34]

    Variational multinomial logit Gaussian process

    Kian Ming A Chai. Variational multinomial logit Gaussian process. Journal of Machine Learning Research, 13(Jun):1745–1808, 2012

  27. [35]

    Orlicz random Fourier features

    Linda Chamakh, Emmanuel Gobet, and Zoltán Szabó. Orlicz random Fourier features. Journal of Machine Learning Research, 21(145):1–37, 2020

  28. [36]

    Bronstein

    Benjamin Paul Chamberlain, James Rowbottom, Davide Eynard, Francesco Di Giovanni, Xiaowen Dong, and Michael M. Bronstein. Beltrami flow and neural diffusion on graphs. CoRR, abs/2110.09443, 2021

  29. [37]

    Rowbottom, Maria I

    Benjamin Paul Chamberlain, James R. Rowbottom, Maria I. Gorinova, Stefan Webb, Emanuele Rossi, and Michael M. Bronstein. GRAND: Graph neural diffusion. In ICML, 2021

  30. [38]

    Re- current neural networks for multivariate time series with missing values

    Zhengping Che, Sanjay Purushotham, Kyunghyun Cho, David Sontag, and Yan Liu. Re- current neural networks for multivariate time series with missing values. Scientific re- ports, 8(1):6085, 2018

  31. [39]

    Convolutional kernel networks for graph-structured data

    Dexiong Chen, Laurent Jacob, and Julien Mairal. Convolutional kernel networks for graph-structured data. In International Conference on Machine Learning , pages 1576–

  32. [40]

    Integration of paths—a faithful representation of paths by non- commutative formal power series

    Kuo-Tsai Chen. Integration of paths—a faithful representation of paths by non- commutative formal power series. Trans. Amer. Math. Soc., 89:395–407, 1958

  33. [41]

    On the equivalence between graph isomorphism testing and function approximation with gnns

    Zhengdao Chen, Soledad Villar, Lei Chen, and Joan Bruna. On the equivalence between graph isomorphism testing and function approximation with gnns. Advances in neural information processing systems, 32, 2019

  34. [42]

    A primer on the signature method in machine learning

    Ilya Chevyrev and Andrey Kormilitzin. A primer on the signature method in machine learning. arXiv preprint arXiv:1603.03788, 2016

  35. [43]

    Persistence Paths and Signature Features in Topological Data Analysis

    Ilya Chevyrev, Vidit Nanda, and Harald Oberhauser. Persistence Paths and Signature Features in Topological Data Analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–1, 2018

  36. [44]

    Signature moments to characterize laws of stochastic processes

    Ilya Chevyrev and Harald Oberhauser. Signature moments to characterize laws of stochastic processes. Journal of Machine Learning Research, 23(176):1–42, 2022. 178

  37. [45]

    Learning phrase representations using rnn encoder-decoder for statistical machine translation

    Kyunghyun Cho, B van Merrienboer, Caglar Gulcehre, F Bougares, H Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. In Conference on Empirical Methods in Natural Language Processing (EMNLP 2014), 2014

  38. [46]

    Francois Chollet et al. Keras. https://github.com/fchollet/keras, 2015

  39. [47]

    The geometry of random features

    Krzysztof Choromanski, Mark Rowland, Tamas Sarlos, Vikas Sindhwani, Richard Turner, and Adrian Weller. The geometry of random features. In International Conference on Artificial Intelligence and Statistics, pages 1–9, 2018

  40. [48]

    The unreasonable effective- ness of structured random orthogonal embeddings

    Krzysztof Choromanski, Mark Rowland, and Adrian Weller. The unreasonable effective- ness of structured random orthogonal embeddings. InInternational Conference on Neural Information Processing Systems, pages 218–227, 2017

  41. [49]

    Recycling randomness with structure for sublinear time kernel expansions

    Krzysztof Choromanski and Vikas Sindhwani. Recycling randomness with structure for sublinear time kernel expansions. InInternational Conference on Machine Learning, pages 2502–2510, 2016

  42. [50]

    Hybrid random features

    Krzysztof Marcin Choromanski, Han Lin, Haoxian Chen, Arijit Sehanobish, Yuanzhe Ma, Deepali Jain, Jake Varley , Andy Zeng, Michael S Ryoo, Valerii Likhosherstov, Dmitry Kalashnikov, Vikas Sindhwani, and Adrian Weller. Hybrid random features. In Inter- national Conference on Le...

  43. [51]

    Tensor networks for dimensionality reduction and large-scale optimization: Part 1 low-rank tensor decompositions

    Andrzej Cichocki, Namgil Lee, Ivan Oseledets, Anh-Huy Phan, Qibin Zhao, and Danilo P Mandic. Tensor networks for dimensionality reduction and large-scale optimization: Part 1 low-rank tensor decompositions. Foundations and Trends® in Machine Learning, 9(4- 5):249–429, 2016

  44. [52]

    Sk-tree: a systematic malware detection algorithm on streaming trees via the signature kernel

    Thomas Cochrane, Peter Foster, Varun Chhabra, Maud Lemercier, Terry Lyons, and Cristo- pher Salvi. Sk-tree: a systematic malware detection algorithm on streaming trees via the signature kernel. In 2021 IEEE International Conference on Cyber Security and Resilience (CSR), pages...

  45. [53]

    On the expressive power of deep learning: A tensor analysis

    Nadav Cohen, Or Sharir, and Amnon Shashua. On the expressive power of deep learning: A tensor analysis. In Conference on learning theory, pages 698–728, 2016

  46. [54]

    Inference suboptimality in variational autoencoders

    Chris Cremer, Xuechen Li, and David Duvenaud. Inference suboptimality in variational autoencoders. In Proceedings of the 35th International Conference on Machine Learning , pages 1078–1086, 2018. 179

  47. [55]

    Discrete-time signatures and randomness in reservoir computing

    Christa Cuchiero, Lukas Gonon, Lyudmila Grigoryeva, Juan-Pablo Ortega, and Josef Te- ichmann. Discrete-time signatures and randomness in reservoir computing. IEEE Trans- actions on Neural Networks and Learning Systems, 33(11):6321–6330, 2022

  48. [56]

    On the mathematical foundations of learning

    Felipe Cucker and Steve Smale. On the mathematical foundations of learning. Bulletin of the American mathematical society, 39(1):1–49, 2002

  49. [57]

    Bonilla, Pietro Michiardi, and Maurizio Filippone

    Kurt Cutajar, Edwin V . Bonilla, Pietro Michiardi, and Maurizio Filippone. Random feature expansions for deep Gaussian processes. InInternational Conference on Machine Learning (ICML), pages 884–893, 2017

  50. [58]

    Fast global alignment kernels

    Marco Cuturi. Fast global alignment kernels. In International Conference on Machine Learning, pages 929–936, 2011

  51. [60]

    Autoregressive kernels for time series

    Marco Cuturi and Arnaud Doucet. Autoregressive kernels for time series. arXiv preprint arXiv:1101.0673, 2011

  52. [61]

    Deep Gaussian processes

    Andreas Damianou and Neil Lawrence. Deep Gaussian processes. InArtificial Intelligence and Statistics, pages 207–215, 2013

  53. [62]

    Gaussian quadrature for kernel fea- tures

    Tri Dao, Christopher M De Sa, and Christopher Ré. Gaussian quadrature for kernel fea- tures. Advances in Neural Information Processing Systems (NeurIPS), 30, 2017

  54. [63]

    The UCR time series archive

    Hoang Anh Dau, Anthony Bagnall, Kaveh Kamgar, Chin-Chia Michael Yeh, Yan Zhu, Shaghayegh Gharghabi, Chotirat Ann Ratanamahatana, and Eamonn Keogh. The UCR time series archive. IEEE/CAA Journal of Automatica Sinica, 6(6):1293–1305, 2019

  55. [64]

    Alexander G. de G. Matthews, Mark van der Wilk, Tom Nickson, Keisuke Fujii, Alexis Boukouvalas, Pablo León-Villagrá, Zoubin Ghahramani, and James Hensman. Gpflow: A Gaussian process library using TensorFlow. Journal of Machine Learning Research , 18:40:1–40:6, 2017

  56. [65]

    Convolutional neural networks on graphs with fast localized spectral filtering

    Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. Advances in neural information processing systems, 29, 2016

  57. [66]

    Statistical comparisons of classifiers over multiple data sets

    Janez Demšar. Statistical comparisons of classifiers over multiple data sets. Journal of Machine learning research, 7(Jan):1–30, 2006. 180

  58. [67]

    Group representations in probability and statistics

    Persi Diaconis. Group representations in probability and statistics. Lecture notes- monograph series, 11:i–192, 1988

  59. [68]

    Time-warping invariants of multidimensional time series

    Joscha Diehl, Kurusch Ebrahimi-Fard, and Nikolas Tapia. Time-warping invariants of multidimensional time series. Acta Applicandae Mathematicae, 170(1):265–290, 2020

  60. [69]

    Generalized iterated-sums sig- natures

    Joscha Diehl, Kurusch Ebrahimi-Fard, and Nikolas Tapia. Generalized iterated-sums sig- natures. Journal of Algebra, 632:801–824, 2023

  61. [70]

    Probabilistic recurrent state-space models

    Andreas Doerr, Christian Daniel, Martin Schiegg, Nguyen-Tuong Duy , Stefan Schaal, Marc Toussaint, and Trimpe Sebastian. Probabilistic recurrent state-space models. In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Ma- chine Learni...

  62. [71]

    Probabilistic recurrent state-space models

    Andreas Doerr, Christian Daniel, Martin Schiegg, Nguyen-Tuong Duy , Stefan Schaal, Marc Toussaint, and Trimpe Sebastian. Probabilistic recurrent state-space models. In Interna- tional Conference on Machine Learning (ICML), pages 1280–1289. PMLR, 2018

  63. [72]

    Incorporating Nesterov Momentum into Adam

    Timothy Dozat. Incorporating Nesterov Momentum into Adam. In International Confer- ence on Learning Representations, 2015

  64. [73]

    UCI machine learning repository , 2017

    Dheeru Dua and Casey Graff. UCI machine learning repository , 2017

  65. [74]

    The sizes of compact subsets of Hilbert space and continuity of Gaus- sian processes

    Richard M Dudley . The sizes of compact subsets of Hilbert space and continuity of Gaus- sian processes. Journal of Functional Analysis, 1(3):290–330, 1967

  66. [75]

    Cambridge University Press, 2002

    Richard M Dudley .Real analysis and probability, volume 74. Cambridge University Press, 2002

  67. [76]

    Approximate Bayesian computation with path signatures

    Joel Dyer, Patrick Cannon, and Sebastian M Schmon. Approximate Bayesian computation with path signatures. InThe 40th Conference on Uncertainty in Artificial Intelligence, 2024

  68. [77]

    Cannon, and Sebastian M

    Joel Dyer, Patrick W . Cannon, and Sebastian M. Schmon. Amortised likelihood-free in- ference for expensive time-series simulators with signatured ratio estimation. InInterna- tional Conference on Artificial Intelligence and Statistics, pages 11131–11144, 2022

  69. [78]

    Identifi- cation of Gaussian process state space models.Advances in Neural Information Processing Systems (NeurIPS), 30, 2017

    Stefanos Eleftheriadis, Tom Nicholson, Marc Deisenroth, and James Hensman. Identifi- cation of Gaussian process state space models.Advances in Neural Information Processing Systems (NeurIPS), 30, 2017. 181

  70. [79]

    Ahmed A. A. Elhag, Gabriele Corso, Hannes Stärk, and Michael M. Bronstein. Graph anisotropic diffusion. In ICLR 2022 Workshop on Geometrical and Topological Representa- tion Learning, 2022

  71. [80]

    A fair comparison of graph neural networks for graph classification

    Federico Errica, Marco Podda, Davide Bacciu, and Alessio Micheli. A fair comparison of graph neural networks for graph classification. In International Conference on Learning Representations, 2019

  72. [81]

    Random feature mapping with signed circulant matrix projection

    Chang Feng, Qinghua Hu, and Shizhong Liao. Random feature mapping with signed circulant matrix projection. In International Joint Conference on Artificial Intelligence , page 3490–3496, 2015

  73. [82]

    Framing RNN as a kernel method: A neural ODE approach

    Adeline Fermanian, Pierre Marion, Jean-Philippe Vert, and Gérard Biau. Framing RNN as a kernel method: A neural ODE approach. In Advances in Neural Information Processing Systems, pages 3121–3134, 2021

  74. [83]

    GP-VAE: Deep probabilistic time series imputation

    Vincent Fortuin, Dmitry Baranchuk, Gunnar Rätsch, and Stephan Mandt. GP-VAE: Deep probabilistic time series imputation. In International Conference on Artificial Intelligence and Statistics, pages 1651–1661. PMLR, 2020

  75. [84]

    Variational Gaussian process state-space models

    Roger Frigola, Yutian Chen, and Carl Edward Rasmussen. Variational Gaussian process state-space models. Advances in Neural Information Processing Systems (NeurIPS) , 27, 2014

  76. [85]

    Bayesian inference and learning in Gaussian process state-space models with particle MCMC

    Roger Frigola, Fredrik Lindsten, Thomas B Schön, and Carl Edward Rasmussen. Bayesian inference and learning in Gaussian process state-space models with particle MCMC. In C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger, editors,Ad- vances in Neural I...

  77. [86]

    Bayesian inference and learning in Gaussian process state-space models with particle MCMC

    Roger Frigola, Fredrik Lindsten, Thomas B Schön, and Carl Edward Rasmussen. Bayesian inference and learning in Gaussian process state-space models with particle MCMC. Ad- vances in Neural Information Processing Systems (NeurIPS), 26, 2013

  78. [87]

    A course on rough paths

    Peter K Friz and Martin Hairer. A course on rough paths. Springer, 2020

  79. [88]

    Sriperumbudur

    Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Bharath K. Sriperumbudur. Characteristic kernels on groups and semigroups. In Advances in Neural Information Pro- cessing Systems, page 473–480, 2008. 182

  80. [89]

    Improving the Gaussian process sparse spectrum approx- imation by representing uncertainty in frequency inputs

    Yarin Gal and Richard Turner. Improving the Gaussian process sparse spectrum approx- imation by representing uncertainty in frequency inputs. In International Conference on Machine Learning (ICML), pages 655–664, 2015

  81. [90]

    Rank and symmetries of signature tensors

    Francesco Galuppi and Pierpaola Santarsiero. Rank and symmetries of signature tensors. arXiv preprint arXiv:2407.20405, 2024

  82. [91]

    GPytorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration

    Jacob Gardner, Geoff Pleiss, Kilian Q Weinberger, David Bindel, and Andrew G Wilson. GPytorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration. Advances in Neural Information Processing Systems (NeurIPS), 31, 2018

  83. [92]

    Probabilistic forecasting with spline quantile function rnns

    Jan Gasthaus, Konstantinos Benidis, Yuyang Wang, Syama Sundar Rangapuram, David Salinas, Valentin Flunkert, and Tim Januschowski. Probabilistic forecasting with spline quantile function rnns. In International Conference on Artificial Intelligence and Statistics (AISTAS), pages...

  84. [93]

    Principe de moindre action, propagation de la chaleur et estimees sous elliptiques sur certains groupes nilpotents

    Bernard Gaveau. Principe de moindre action, propagation de la chaleur et estimees sous elliptiques sur certains groupes nilpotents. Acta Mathematica, 139(none):95–153, January 1977

  85. [94]

    Bayesian non-parametrics and the probabilistic approach to mod- elling

    Zoubin Ghahramani. Bayesian non-parametrics and the probabilistic approach to mod- elling. Philosophical transactions. Series A, Mathematical, physical, and engineering sci- ences, 371:20110553, 02 2013

  86. [95]

    Neural message passing for quantum chemistry

    Justin Gilmer, Samuel S Schoenholz, Patrick F Riley , Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry . In International conference on machine learning, pages 1263–1272. PMLR, 2017

  87. [96]

    Signatures, lipschitz-free spaces, and paths of persistence diagrams

    Chad Giusti and Darrick Lee. Signatures, lipschitz-free spaces, and paths of persistence diagrams. SIAM Journal on Applied Algebra and Geometry, 7(4):828–866, 2023

  88. [97]

    Strictly proper scoring rules, prediction, and estimation

    Tilmann Gneiting and Adrian E Raftery . Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007

  89. [98]

    Webb, Rob Hyndman, and Pablo Montero-Manso

    Rakshitha Wathsadini Godahewa, Christoph Bergmeir, Geoffrey I. Webb, Rob Hyndman, and Pablo Montero-Manso. Monash time series forecasting archive. In Thirty-fifth Con- ference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021

  90. [99]

    Components of a new research resource for complex physiologic signals

    AL Goldberger, LAN Amaral, L Glass, JM Hausdorff, P Ch Ivanov, RG Mark, JE Mietus, GB Moody , CK Peng, and HE Stanley . Components of a new research resource for complex physiologic signals. PhysioBank, PhysioToolkit, and Physionet, 2000. 183

  91. [100]

    Concentration inequalities for poly- nomials inα-sub-exponential random variables

    Friedrich Götze, Holger Sambale, and Arthur Sinulis. Concentration inequalities for poly- nomials inα-sub-exponential random variables. Electronic Journal of Probability, 26:1 – 22, 2021

  92. [101]

    Sparse arrays of signatures for online character recognition

    Benjamin Graham. Sparse arrays of signatures for online character recognition. arXiv preprint arXiv:1308.0371, 2013

  93. [102]

    An introduction to long-memory time series models and fractional differencing

    Clive WJ Granger and Roselyne Joyeux. An introduction to long-memory time series models and fractional differencing. Journal of time series analysis, 1(1):15–29, 1980

  94. [103]

    Graph neural networks in tensorflow and keras with spektral [application notes]

    Daniele Grattarola and Cesare Alippi. Graph neural networks in tensorflow and keras with spektral [application notes]. IEEE Computational Intelligence Magazine, 16(1):99– 106, 2021

  95. [104]

    Hybrid speech recognition with deep bidirectional lstm

    Alex Graves, Navdeep Jaitly , and Abdel-rahman Mohamed. Hybrid speech recognition with deep bidirectional lstm. In 2013 IEEE workshop on automatic speech recognition and understanding, pages 273–278. IEEE, 2013

  96. [105]

    Speech recognition with deep recurrent neural networks

    Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. Speech recognition with deep recurrent neural networks. In2013 IEEE international conference on acoustics, speech and signal processing, pages 6645–6649. IEEE, 2013

  97. [106]

    Framewise phoneme classification with bidirec- tional lstm and other neural network architectures

    Alex Graves and Jürgen Schmidhuber. Framewise phoneme classification with bidirec- tional lstm and other neural network architectures. Neural Networks, 18(5):602–610, 2005

  98. [107]

    A kernel two-sample test

    Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola. A kernel two-sample test. Journal of Machine Learning Research, 13(Mar):723– 773, 2012

  99. [108]

    Heat kernel and analysis on manifolds, volume 47

    Alexander Grigoryan. Heat kernel and analysis on manifolds, volume 47. American Math- ematical Soc., 2009

  100. [109]

    node2vec: Scalable feature learning for networks

    Aditya Grover and Jure Leskovec. node2vec: Scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining, pages 855–864, 2016

  101. [110]

    Uniqueness for the Signature of a path of bounded variation and the reduced path group

    Ben Hambly and Terry Lyons. Uniqueness for the Signature of a path of bounded variation and the reduced path group. Ann. of Math. (2), 171(1):109–167, 2010

  102. [111]

    Hambly and Terry J

    Ben M. Hambly and Terry J. Lyons. Uniqueness for the signature of a path of bounded variation and the reduced path group. Annals of Mathematics, 171(1):109–167, 2010. 184

  103. [112]

    Inductive representation learning on large graphs

    Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. Advances in neural information processing systems, 30, 2017

  104. [113]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  105. [114]

    Heather and Benjamin Chain

    James M. Heather and Benjamin Chain. The sequence of sequencers: The history of sequencing dna. Genomics, 107(1):1–8, 2016

  106. [115]

    James Hensman, Alexander G. de G. Matthews, and Zoubin Ghahramani. Scalable varia- tional gaussian process classification. In Guy Lebanon and S. V . N. Vishwanathan, editors, AISTATS, volume 38 of JMLR Workshop and Conference Proceedings. JMLR.org, 2015

  107. [116]

    Variational fourier features for Gaus- sian processes

    James Hensman, Nicolas Durrande, and Arno Solin. Variational fourier features for Gaus- sian processes. Journal of Machine Learning Research, 18(151):1–52, 2018

  108. [117]

    Lawrence

    James Hensman, Nicoló Fusi, and Neil D. Lawrence. Gaussian processes for big data. CoRR, abs/1309.6835, 2013

  109. [118]

    Gaussian processes for big data

    James Hensman, Nicolò Fusi, and Neil D Lawrence. Gaussian processes for big data. In Uncertainty in Artificial Intelligence (UAI), 2013

  110. [119]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems (NeurIPS), 33:6840–6851, 2020

  111. [120]

    Long short-term memory .Neural computation, 9(8):1735–1780, 1997

    Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory .Neural computation, 9(8):1735–1780, 1997

  112. [121]

    Long short-term memory

    Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory . Neural Computa- tion, 9(8):1735–1780, 1997

  113. [122]

    Hypoelliptic second order differential equations

    Lars Hörmander. Hypoelliptic second order differential equations. Acta Mathematica, 119:147–171, 1967

  114. [123]

    Set functions for time series

    Max Horn, Michael Moor, Christian Bock, Bastian Rieck, and Karsten Borgwardt. Set functions for time series. In ICML, 2020

  115. [124]

    Approximation capabilities of multilayer feedforward networks

    Kurt Hornik. Approximation capabilities of multilayer feedforward networks. Neural networks, 4(2):251–257, 1991. 185

  116. [125]

    Deep multimodal multilinear fusion with high-order polynomial pooling

    Ming Hou, Jiajia Tang, Jianhai Zhang, Wanzeng Kong, and Qibin Zhao. Deep multimodal multilinear fusion with high-order polynomial pooling. InAdvances in Neural Information Processing Systems, pages 12136–12145, 2019

  117. [126]

    Forecasting: principles and practice

    RJ Hyndman. Forecasting: principles and practice. OTexts, 2018

  118. [127]

    Overcoming mean-field approximations in recurrent Gaussian process models

    Alessandro Davide Ialongo, Mark Van Der Wilk, James Hensman, and Carl Edward Ras- mussen. Overcoming mean-field approximations in recurrent Gaussian process models. In International conference on machine learning (ICML), pages 2931–2940. PMLR, 2019

  119. [128]

    Deep learning for time series classification: a review

    Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhassane Idoumghar, and Pierre-Alain Muller. Deep learning for time series classification: a review. Data Min- ing and Knowledge Discovery, 33(4):917–963, July 2019

  120. [129]

    Non-parametric online market regime detection and regime clustering for multidimensional and path-dependent data structures

    Zacharia Issa and Blanka Horvath. Non-parametric online market regime detection and regime clustering for multidimensional and path-dependent data structures. arXiv preprint arXiv:2306.15835, 2023

  121. [130]

    Non-adversarial training of neural sdes with signature kernel scores

    Zacharia Issa, Blanka Horvath, Maud Lemercier, and Cristopher Salvi. Non-adversarial training of neural sdes with signature kernel scores. In Advances in Neural Information Processing Systems, volume 36, pages 11102–11126, 2023

  122. [131]

    On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables

    Leon Isserlis. On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables. Biometrika, 12(1/2):134–139, 1918

  123. [132]

    Neural tangent kernel: Convergence and generalization in neural networks.Advances in neural information processing systems, 31, 2018

    Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks.Advances in neural information processing systems, 31, 2018

  124. [133]

    Parametric Gaussian process regres- sors

    Martin Jankowiak, Geoff Pleiss, and Jacob Gardner. Parametric Gaussian process regres- sors. In International conference on machine learning (ICML), 2020

  125. [134]

    Gaussian Hilbert Spaces

    Svante Janson. Gaussian Hilbert Spaces. Cambridge University Press, 1997

  126. [135]

    Extensions of Lips- chitz maps into Banach spaces

    William B Johnson, Joram Lindenstrauss, and Gideon Schechtman. Extensions of Lips- chitz maps into Banach spaces. Israel Journal of Mathematics, 54(2):129–138, 1986

  127. [136]

    Multivariate lstm-fcns for time series classification

    Fazle Karim, Somshubra Majumdar, Houshang Darabi, and Samuel Harford. Multivariate lstm-fcns for time series classification. Neural Networks, 116:237–245, 2019

  128. [137]

    Generalized random shapelet forests

    Isak Karlsson, Panagiotis Papapetrou, and Henrik Boström. Generalized random shapelet forests. Data Min. Knowl. Discov., 30(5):1053–1085, September 2016. 186

  129. [138]

    Multiparameter processes: an introduction to random fields

    Davar Khoshnevisan. Multiparameter processes: an introduction to random fields. Springer Science & Business Media, 2002

  130. [139]

    Generalized tensor models for recurrent neural networks

    Valentin Khrulkov, Oleksii Hrinchuk, and Ivan Oseledets. Generalized tensor models for recurrent neural networks. arXiv preprint arXiv:1901.10801, 2019

  131. [140]

    Deep signature transforms

    Patrick Kidger, Patric Bonnier, Imanol Perez Arribas, Cristopher Salvi, and Terry Lyons. Deep signature transforms. In H. Wallach, H. Larochelle, A. Beygelzimer, F . d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32 , pages 3099...

  132. [141]

    Neural SDEs as infinite-dimensional GANs

    Patrick Kidger, James Foster, Xuechen Li, Harald Oberhauser, and Terry J Lyons. Neural SDEs as infinite-dimensional GANs. In International Conference on Machine Learning , pages 5453–5463, 2021

  133. [142]

    Signatory: differentiable computations of the signa- ture and logsignature transforms, on both CPU and GPU

    Patrick Kidger and Terry Lyons. Signatory: differentiable computations of the signa- ture and logsignature transforms, on both CPU and GPU. In International Conference on Learning Representations, 2021

  134. [143]

    Adam: A method for stochastic optimization

    Diederik P Kingma. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014

  135. [144]

    Auto-encoding variational Bayes

    Diederik P Kingma and Max Welling. Auto-encoding variational Bayes. In International Conference on Learning Representations (ICLR), 2014

  136. [145]

    Semi-Supervised Classification with Graph Convolutional Networks

    Thomas Kipf and Max Welling. Semi-Supervised Classification with Graph Convolutional Networks. In International Conference on Learning Representations, 2017

  137. [146]

    Kernels for sequentially ordered data

    Franz J Király and Harald Oberhauser. Kernels for sequentially ordered data. Journal of Machine Learning Research, 20(31):1–45, 2019

  138. [147]

    Diffusion improves graph learning

    Johannes Klicpera, Stefan Weissenberger, and Stephan Günnemann. Diffusion improves graph learning. ArXiv, abs/1911.05485, 2019

  139. [148]

    Tensor decompositions and applications

    Tamara G Kolda and Brett W Bader. Tensor decompositions and applications. SIAM review, 51(3):455–500, 2009

  140. [149]

    Predict, refine, synthesize: Self-guiding diffusion mod- els for probabilistic time series forecasting

    Marcel Kollovieh, Abdul Fatir Ansari, Michael Bohlke-Schneider, Jasper Zschiegner, Hao Wang, and Yuyang Bernie Wang. Predict, refine, synthesize: Self-guiding diffusion mod- els for probabilistic time series forecasting. Advances in Neural Information Processing Systems (NeurI...

  141. [150]

    Diffusion kernels on graphs and other discrete structures

    Risi Kondor. Diffusion kernels on graphs and other discrete structures. In ICML 2002, 2002

  142. [151]

    Tensor regression networks

    Jean Kossaifi, Zachary C Lipton, Aran Khanna, Tommaso Furlanello, and Anima Anand- kumar. Tensor regression networks. arXiv preprint arXiv:1707.08308, 2017

  143. [152]

    Gaussian sketching yields a JL lemma in RKHS

    Samory Kpotufe and Bharath Sriperumbudur. Gaussian sketching yields a JL lemma in RKHS. In International Conference on Artificial Intelligence and Statistics , pages 3928– 3937, 2020

  144. [153]

    Moving beyond sub-Gaussianity in high-dimensional statistics: Applications in covariance estimation and linear regression

    Arun Kumar Kuchibhotla and Abhishek Chakrabortty . Moving beyond sub-Gaussianity in high-dimensional statistics: Applications in covariance estimation and linear regression. Information and Inference: A Journal of the IMA, 11(4):1389–1456, 2022

  145. [154]

    Modeling long-and short- term temporal patterns with deep neural networks

    Guokun Lai, Wei-Cheng Chang, Yiming Yang, and Hanxiao Liu. Modeling long-and short- term temporal patterns with deep neural networks. In The 41st international ACM SIGIR conference on research & development in information retrieval, pages 95–104, 2018

  146. [155]

    Serge Lang. Algebra. Springer, 2002

  147. [156]

    Samuel Lanthaler and Nicholas H. Nelsen. Error bounds for learning with vector-valued random features. In Advances in Neural Information Processing Systems, 2023

  148. [157]

    Inter-domain Gaussian processes for sparse inference using inducing features

    Miguel Lázaro-Gredilla and Anibal Figueiras-Vidal. Inter-domain Gaussian processes for sparse inference using inducing features. In Advances in Neural Information Processing Systems, pages 1087–1095, 2009

  149. [158]

    Sparse spectrum Gaussian process regression

    Miguel Lázaro-Gredilla, Joaquin Quinonero-Candela, Carl Edward Rasmussen, and Aníbal R Figueiras-Vidal. Sparse spectrum Gaussian process regression. Journal of Ma- chine Learning Research (JMLR), 11:1865–1881, 2010

  150. [159]

    Fastfood-approximating kernel expansions in loglinear time

    Quoc Le, Tamás Sarlós, Alex Smola, et al. Fastfood-approximating kernel expansions in loglinear time. In International Conference on Machine Learning, page 8, 2013

  151. [160]

    Convolutional networks for images, speech, and time series

    Yann LeCun, Yoshua Bengio, et al. Convolutional networks for images, speech, and time series. The handbook of brain theory and neural networks, 3361(10):1995, 1995

  152. [161]

    Gradient-based learning applied to document recognition

    Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998

  153. [162]

    Path signatures on lie groups

    Darrick Lee and Robert Ghrist. Path signatures on lie groups. arXiv preprint arXiv:2007.06633, 2020. 188

  154. [163]

    The signature kernel

    Darrick Lee and Harald Oberhauser. The signature kernel. arXiv preprint arXiv:2305.04625, 2023

  155. [164]

    Self-attention graph pooling

    Junhyun Lee, Inyeop Lee, and Jaewoo Kang. Self-attention graph pooling. In Interna- tional conference on machine learning, pages 3734–3743. PMLR, 2019

  156. [165]

    Bonilla, Theodoros Damoulas, and Terry J Lyons

    Maud Lemercier, Cristopher Salvi, Thomas Cass, Edwin V . Bonilla, Theodoros Damoulas, and Terry J Lyons. SigGPDE: Scaling sparse Gaussian processes on sequential data. In International Conference on Machine Learning, pages 6233–6242, 2021

  157. [166]

    SigGPDE: Scaling sparse Gaussian processes on sequential data

    Maud Lemercier, Cristopher Salvi, Thomas Cass, Edwin V Bonilla, Theodoros Damoulas, and Terry J Lyons. SigGPDE: Scaling sparse Gaussian processes on sequential data. In International conference on machine learning (ICML), pages 6233–6242. PMLR, 2021

  158. [167]

    Trigonometric quadrature Fourier features for scalable Gaussian process regression

    Kevin Li, Max Balakirsky , and Simon Mak. Trigonometric quadrature Fourier features for scalable Gaussian process regression. In International Conference on Artificial Intelligence and Statistics (AISTAS), pages 3484–3492. PMLR, 2024

  159. [168]

    Disentangled sequential autoencoder

    Yingzhen Li and Stephan Mandt. Disentangled sequential autoencoder. arXiv preprint arXiv:1803.02991, 2018

  160. [169]

    Gated graph sequence neural networks

    Yujia Li, Richard Zemel, Marc Brockschmidt, and Daniel Tarlow. Gated graph sequence neural networks. In Proceedings of ICLR’16, 2016

  161. [170]

    Towards a unified analysis of random Fourier features

    Zhu Li, Jean-Francois Ton, Dino Oglic, and Dino Sejdinovic. Towards a unified analysis of random Fourier features. In International Conference on Machine Learning, pages 3905– 3914, 2019

  162. [171]

    Learning representations from imperfect time series data via tensor rank regularization

    Paul Pu Liang, Zhun Liu, Yao-Hung Hubert Tsai, Qibin Zhao, Ruslan Salakhutdinov, and Louis-Philippe Morency . Learning representations from imperfect time series data via tensor rank regularization. arXiv preprint arXiv:1907.01011, 2019

  163. [172]

    A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent

    Zhenyu Liao, Romain Couillet, and Michael W Mahoney . A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent. In Advances in Neural Information Processing Systems, pages 13939–13950, 2020

  164. [173]

    Random features for kernel approximation: A survey on algorithms, theory , and beyond.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):7128–7148, 2021

    Fanghui Liu, Xiaolin Huang, Yudong Chen, and Johan AK Suykens. Random features for kernel approximation: A survey on algorithms, theory , and beyond.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):7128–7148, 2021. 189

  165. [174]

    An intriguing failing of convolutional neural networks and the coord- conv solution

    Rosanne Liu, Joel Lehman, Piero Molino, Felipe Petroski Such, Eric Frank, Alex Sergeev, and Jason Yosinski. An intriguing failing of convolutional neural networks and the coord- conv solution. In Proceedings of the 32nd International Conference on Neural Information Processing...

  166. [175]

    Efficient low-rank multimodal fusion with modality-specific factors

    Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency . Efficient low-rank multimodal fusion with modality-specific factors. arXiv preprint arXiv:1806.00064, 2018

  167. [176]

    Text classification using string kernels

    Huma Lodhi, Craig Saunders, John Shawe-Taylor, Nello Cristianini, and Chris Watkins. Text classification using string kernels. Journal of Machine Learning Research (JMLR) , 2(Feb):419–444, 2002

  168. [177]

    Sgdr: Stochastic gradient descent with warm restarts

    Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In ICLR, 2017

  169. [178]

    Rough paths, signatures and the modelling of functions on streams

    Terry Lyons. Rough paths, signatures and the modelling of functions on streams. In Proceedings of the International Congress of Mathematicians, 2014

  170. [179]

    Terry Lyons, Michael Caruana, and Thierry Lévy .Differential Equations Driven by Rough Paths. Éc. Été Probab. St.-Flour. Springer-Verlag, Berlin Heidelberg, 2007

  171. [180]

    Sketching the order of events

    Terry Lyons and Harald Oberhauser. Sketching the order of events. arXiv preprint arXiv:1708.09708, 2017

  172. [181]

    Inversion of signature for paths of bounded variation

    Terry Lyons and Weijun Xu. Inversion of signature for paths of bounded variation. arXiv preprint arXiv:1112.0452, 2011

  173. [182]

    Springer, 2007

    Terry J Lyons, Michael Caruana, and Thierry Lévy .Differential equations driven by rough paths. Springer, 2007

  174. [183]

    The M4 competi- tion: 100,000 time series and 61 forecasting methods

    Spyros Makridakis, Evangelos Spiliotis, and Vassilios Assimakopoulos. The M4 competi- tion: 100,000 time series and 61 forecasting methods. International Journal of Forecast- ing, 36(1):54–74, 2020

  175. [184]

    Provably powerful graph networks

    Haggai Maron, Heli Ben-Hamu, Hadar Serviansky , and Yaron Lipman. Provably powerful graph networks. Advances in neural information processing systems, 32, 2019

  176. [185]

    Scalable Gaussian process inference using variational meth- ods

    Alexander G de G Matthews. Scalable Gaussian process inference using variational meth- ods. PhD thesis, Cambridge University , 2017. 190

  177. [186]

    On sparse variational methods and the Kullback-Leibler divergence between stochastic processes

    Alexander G de G Matthews, James Hensman, Richard Turner, and Zoubin Ghahramani. On sparse variational methods and the Kullback-Leibler divergence between stochastic processes. Journal of Machine Learning Research, 51:231–239, 2016

  178. [187]

    Mattos, Zhenwen Dai, Andreas Damianou, Jeremy Forth, Guilherme A

    César Lincoln C. Mattos, Zhenwen Dai, Andreas Damianou, Jeremy Forth, Guilherme A. Barreto, and Neil D. Lawrence. Recurrent Gaussian processes. InInternational Conference on Learning Representations (ICLR), volume 3, 2016

  179. [188]

    Umap: Uniform manifold approximation and projection.The Journal of Open Source Software, 3(29):861, 2018

    Leland McInnes, John Healy , Nathaniel Saul, and Lukas Grossberger. Umap: Uniform manifold approximation and projection.The Journal of Open Source Software, 3(29):861, 2018

  180. [189]

    Universal kernels

    Charles A Micchelli, Yuesheng Xu, and Haizhang Zhang. Universal kernels. Journal of Machine Learning Research, 7(Dec):2651–2667, 2006

  181. [190]

    Shen-Orr, Shalev Itzkovitz, Nadav Kashtan, Dmitri B

    Ron Milo, Shai S. Shen-Orr, Shalev Itzkovitz, Nadav Kashtan, Dmitri B. Chklovskii, and Uri Alon. Network motifs: simple building blocks of complex networks. Science, 298 5594:824–7, 2002

  182. [191]

    Geometric deep learning on graphs and manifolds using mixture model cnns

    Federico Monti, Davide Boscaini, Jonathan Masci, Emanuele Rodola, Jan Svoboda, and Michael M Bronstein. Geometric deep learning on graphs and manifolds using mixture model cnns. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 5115–5124, 2017

  183. [192]

    A generalised signature method for multivariate time series feature extraction

    James Morrill, Adeline Fermanian, Patrick Kidger, and Terry Lyons. A generalised signature method for multivariate time series feature extraction. arXiv preprint arXiv:2006.00873, 2020

  184. [193]

    Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe

    Christopher Morris, Martin Ritzert, Matthias Fey , William L. Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and Leman Go Neural: Higher-Order Graph Neural Networks. Proceedings of the AAAI Conference on Artificial Intelligence , 33(01):4602–4609, July 2019

  185. [194]

    On rings of operators

    Francis J Murray and J v Neumann. On rings of operators. Annals of Mathematics, pages 116–229, 1936

  186. [195]

    Reparameterization gradients through acceptance-rejection sampling algorithms

    Christian Naesseth, Francisco Ruiz, Scott Linderman, and David Blei. Reparameterization gradients through acceptance-rejection sampling algorithms. In International Conference on Artificial Intelligence and Statistics (AISTAS), pages 489–498, 2017. 191

  187. [196]

    Query-driven active surveying for collective classification

    Galileo Namata, Ben London, Lise Getoor, Bert Huang, and U Edu. Query-driven active surveying for collective classification. In 10th International Workshop on Mining and Learning with Graphs, volume 8, page 1, 2012

  188. [197]

    Handling in- complete heterogeneous data using vaes

    Alfredo Nazabal, Pablo M Olmos, Zoubin Ghahramani, and Isabel Valera. Handling in- complete heterogeneous data using vaes. arXiv preprint arXiv:1807.03653, 2018

  189. [198]

    Nemenyi.Distribution-free Multiple Comparisons

    P . Nemenyi.Distribution-free Multiple Comparisons. Princeton University , 1963

  190. [199]

    Sig-Wasserstein GANs for time series generation

    Hao Ni, Lukasz Szpruch, Marc Sabate-Vidales, Baoren Xiao, Magnus Wiese, and Shujian Liao. Sig-Wasserstein GANs for time series generation. In International Conference on AI in Finance, pages 1–8, 2022

  191. [200]

    Random walk graph neural networks

    Giannis Nikolentzos and Michalis Vazirgiannis. Random walk graph neural networks. Advances in Neural Information Processing Systems, 33:16211–16222, 2020

  192. [201]

    Wavenet: A genera- tive model for raw audio

    Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. Wavenet: A genera- tive model for raw audio. arXiv preprint arXiv:1609.03499, 2016

  193. [202]

    Tensor-train decomposition

    Ivan V Oseledets. Tensor-train decomposition. SIAM Journal on Scientific Computing , 33(5):2295–2317, 2011

  194. [203]

    PyTorch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury , Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. PyTorch: An imperative style, high-performance deep learning library . Advances in Neural Infor- mation Processing Systems ...

  195. [204]

    Pedregosa, G

    F . Pedregosa, G. Varoquaux, A. Gramfort, V . Michel, B. Thirion, O. Grisel, M. Blondel, P . Prettenhofer, R. Weiss, V . Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay . Scikit-learn: Machine learning in Python. Journal of Ma- chine L...

  196. [205]

    Glove: Global vectors for word representation

    Jeffrey Pennington, Richard Socher, and Christopher Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages 1532–1543, Doha, Qatar, October 2014. Association for Computati...

  197. [206]

    Deepwalk: Online learning of so- cial representations

    Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. Deepwalk: Online learning of so- cial representations. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 701–710, 2014. 192

  198. [207]

    Satellite image time series analy- sis under time warping

    Francois Petitjean, Jordi Inglada, and Pierre Gancarski. Satellite image time series analy- sis under time warping. IEEE transactions on geoscience and remote sensing, 50(8):3081– 3095, 2012

  199. [208]

    A unifying view of sparse approximate Gaussian process regression

    Joaquin Quiñonero-Candela and Carl Edward Rasmussen. A unifying view of sparse approximate Gaussian process regression. Journal of Machine Learning Research , 6(Dec):1939–1959, 2005

  200. [209]

    Random features for large-scale kernel machines

    Ali Rahimi and Benjamin Recht. Random features for large-scale kernel machines. In Advances in Neural Information Processing Systems, pages 1177–1184, 2007

  201. [210]

    Uniform approximation of functions with random bases

    Ali Rahimi and Benjamin Recht. Uniform approximation of functions with random bases. In Allerton Conference on Communication, Control, and Computing, pages 555–561, 2008

  202. [211]

    Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning

    Ali Rahimi and Benjamin Recht. Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning. Advances in Neural Information Processing Systems, pages 1313–1320, 2008

  203. [212]

    Tensorized random projections

    Beheshteh Rakhshan and Guillaume Rabusseau. Tensorized random projections. In In- ternational Conference on Artificial Intelligence and Statistics, pages 3306–3316, 2020

  204. [213]

    Hierarchical graph neural nets can capture long-range interactions

    Ladislav Rampášek and Guy Wolf. Hierarchical graph neural nets can capture long-range interactions. In 2021 IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP), pages 1–6. IEEE, 2021

  205. [214]

    Deep state space models for time series forecasting

    Syama Sundar Rangapuram, Matthias W Seeger, Jan Gasthaus, Lorenzo Stella, Yuyang Wang, and Tim Januschowski. Deep state space models for time series forecasting. Ad- vances in Neural Information Processing Systems (NeurIPS), 31, 2018

  206. [215]

    Machine learning in Python: Main developments and technology trends in data science, machine learning, and artifi- cial intelligence

    Sebastian Raschka, Joshua Patterson, and Corey Nolet. Machine learning in Python: Main developments and technology trends in data science, machine learning, and artifi- cial intelligence. Information, 11(4), 2020

  207. [216]

    Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, 2006

  208. [217]

    Free Lie algebras

    Christophe Reutenauer. Free Lie algebras. Handbook of Algebra, 3:887–903, 2003

  209. [218]

    Linda Preiss Rothschild and Elias M. Stein. Hypoelliptic differential operators and nilpo- tent groups. Acta Mathematica, 137:247–320, 1976

  210. [219]

    Generalization properties of learning with random features

    Alessandro Rudi and Lorenzo Rosasco. Generalization properties of learning with random features. In Advances in Neural Information Processing Systems, pages 3218–3228, 2017. 193

  211. [220]

    Rudin.Fourier Analysis on Groups

    W . Rudin.Fourier Analysis on Groups. Dover Publications, 2017

  212. [221]

    Principles of mathematical analysis

    Walter Rudin. Principles of mathematical analysis. McGraw-hill New York, 1976

  213. [222]

    A survey on over- smoothing in graph neural networks

    T Konstantin Rusch, Michael M Bronstein, and Siddhartha Mishra. A survey on over- smoothing in graph neural networks. arXiv preprint arXiv:2303.10993, 2023

  214. [223]

    R.A. Ryan. Introduction to Tensor Products of Banach Spaces. Springer London, 2013

  215. [224]

    Spectral clustering of graphs with the bethe hessian

    Alaa Saade, Florent Krzakala, and Lenka Zdeborová. Spectral clustering of graphs with the bethe hessian. In NIPS, 2014

  216. [225]

    Sakoe and S

    H. Sakoe and S. Chiba. Dynamic programming algorithm optimization for spoken word recognition. IEEE Transactions on Acoustics, Speech, and Signal Processing, 26(1):43–49, February 1978

  217. [226]

    Deep Gaus- sian processes with importance-weighted variational inference

    Hugh Salimbeni, Vincent Dutordoir, James Hensman, and Marc Deisenroth. Deep Gaus- sian processes with importance-weighted variational inference. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volum...

  218. [227]

    DeepAR: Proba- bilistic forecasting with autoregressive recurrent networks

    David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski. DeepAR: Proba- bilistic forecasting with autoregressive recurrent networks. International journal of fore- casting, 36(3):1181–1191, 2020

  219. [228]

    The signature kernel is the solution of a Goursat PDE

    Cristopher Salvi, Thomas Cass, James Foster, Terry Lyons, and Weixin Yang. The signature kernel is the solution of a Goursat PDE. SIAM Journal on Mathematics of Data Science , 3(3):873–899, 2021

  220. [229]

    Some notes on concentration for α-subexponential random variables

    Holger Sambale. Some notes on concentration for α-subexponential random variables. In High Dimensional Probability IX: The Ethereal Volume, pages 167–192. Springer, 2023

  221. [230]

    Z. Sasvári. Multivariate Characteristic and Correlation Functions. De Gruyter, 2013

  222. [231]

    The graph neural network model

    Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model. IEEE transactions on neural networks , 20(1):61–80, 2008

  223. [232]

    Multivariate time series classification with weasel +muse

    Patrick Schäfer and Ulf Leser. Multivariate time series classification with weasel +muse. ArXiv, abs/1711.11343, 2017. 194

  224. [233]

    Modeling relational data with graph convolutional networks

    Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. Modeling relational data with graph convolutional networks. In European semantic web conference, pages 593–607. Springer, 2018

  225. [234]

    Learning with Kernels: Support Vector Ma- chines, Regularization, Optimization, and Beyond

    Bernhard Schölkopf and Alexander Smola. Learning with Kernels: Support Vector Ma- chines, Regularization, Optimization, and Beyond. MIT Press, 2002

  226. [235]

    Bidirectional recurrent neural networks.IEEE trans- actions on Signal Processing, 45(11):2673–2681, 1997

    Mike Schuster and Kuldip K Paliwal. Bidirectional recurrent neural networks.IEEE trans- actions on Signal Processing, 45(11):2673–2681, 1997

  227. [236]

    Motifs for processes on networks

    Alice C Schwarze and Mason A Porter. Motifs for processes on networks. SIAM Journal on Applied Dynamical Systems, 20(4):2516–2557, 2021

  228. [237]

    Collective classification in network data

    Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. Collective classification in network data. AI magazine, 29(3):93–93, 2008

  229. [238]

    Financial time series forecasting with deep learning: A systematic literature review: 2005–2019.Applied soft computing, 90:106181, 2020

    Omer Berat Sezer, Mehmet Ugur Gudelek, and Ahmet Murat Ozbayoglu. Financial time series forecasting with deep learning: A systematic literature review: 2005–2019.Applied soft computing, 90:106181, 2020

  230. [239]

    Cambridge University Press, 2004

    John Shawe-Taylor and Nello Cristianini.Kernel Methods for Pattern Analysis. Cambridge University Press, 2004

  231. [240]

    Pitfalls of graph neural network evaluation

    Oleksandr Shchur, Maximilian Mumme, Aleksandar Bojchevski, and Stephan Günne- mann. Pitfalls of graph neural network evaluation. arXiv preprint arXiv:1811.05868 , 2018

  232. [241]

    Sparse Gaussian processes using pseudo- inputs

    Edward Snelson and Zoubin Ghahramani. Sparse Gaussian processes using pseudo- inputs. In Advances in neural information processing systems, pages 1257–1264, 2006

  233. [242]

    Deep unsupervised learning using nonequilibrium thermodynamics

    Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. InInternational conference on machine learning (ICML), pages 2256–2265. PMLR, 2015

  234. [243]

    On the relation between universality , characteristic kernels and RKHS embedding of measures

    Bharath Sriperumbudur, Kenji Fukumizu, and Gert Lanckriet. On the relation between universality , characteristic kernels and RKHS embedding of measures. In International Conference on Artificial Intelligence and Statistics, pages 773–780, 2010

  235. [244]

    Optimal rates for random Fourier features

    Bharath Sriperumbudur and Zoltán Szabó. Optimal rates for random Fourier features. In Advances in Neural Information Processing Systems, pages 1144–1152, 2015. 195

  236. [245]

    Universality , charac- teristic kernels and rkhs embedding of measures

    Bharath K Sriperumbudur, Kenji Fukumizu, and Gert RG Lanckriet. Universality , charac- teristic kernels and rkhs embedding of measures. Journal of Machine Learning Research, 12(Jul):2389–2410, 2011

  237. [246]

    Approximate kernel PCA: Computational versus statistical trade-off

    Bharath K Sriperumbudur and Nicholas Sterge. Approximate kernel PCA: Computational versus statistical trade-off. The Annals of Statistics, 50(5):2713–2736, 2022

  238. [247]

    Dropout: a simple way to prevent neural networks from overfitting

    Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky , Ilya Sutskever, and Ruslan Salakhut- dinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014

  239. [248]

    Support Vector Machines

    Ingo Steinwart and Andreas Christmann. Support Vector Machines. Springer Science & Business Media, 2008

  240. [249]

    Sub-riemannian geometry

    Robert S Strichartz. Sub-riemannian geometry . Journal of Differential Geometry , 24(2):221–263, 1986

  241. [250]

    Tensor random projection for low memory dimension reduction

    Yiming Sun, Yang Guo, Joel A Tropp, and Madeleine Udell. Tensor random projection for low memory dimension reduction. arXiv preprint arXiv:2105.00105, 2021

  242. [251]

    But how does it work in theory? Linear SVM with random features

    Yitong Sun, Anna Gilbert, and Ambuj Tewari. But how does it work in theory? Linear SVM with random features. In Advances in Neural Information Processing Systems, pages 3379–3388, 2018

  243. [252]

    Translation modeling with bidirectional recurrent neural networks

    Martin Sundermeyer, Tamer Alkhouli, Joern Wuebker, and Hermann Ney . Translation modeling with bidirectional recurrent neural networks. InProceedings of the 2014 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), pages 14–25, 2014

  244. [253]

    On the error of random Fourier features

    Dougal J Sutherland and Jeff Schneider. On the error of random Fourier features. In Conference on Uncertainty in Artificial Intelligence, pages 862–871, 2015

  245. [254]

    Sequence to sequence learning with neural networks

    Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks. In Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 27, pages 3104–3112. Curran Associates, I...

  246. [255]

    On kernel derivative approximation with random Fourier features

    Zoltan Szabo and Bharath Sriperumbudur. On kernel derivative approximation with random Fourier features. InInternational Conference on Artificial Intelligence and Statistics (AISTAS), pages 827–836, 2019. 196

  247. [256]

    Fourier features let networks learn high frequency functions in low dimensional domains

    Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Ragha- van, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. In Advances in Neural Information...

  248. [257]

    General tensor dis- criminant analysis and gabor features for gait recognition

    Dacheng Tao, Xuelong Li, Xindong Wu, and Stephen J Maybank. General tensor dis- criminant analysis and gabor features for gait recognition. IEEE transactions on pattern analysis and machine intelligence, 29(10):1700–1715, 2007

  249. [258]

    GRAND ++: Graph neural diffusion with a source term

    Matthew Thorpe, Tan Minh Nguyen, Hedi Xia, Thomas Strohmer, Andrea Bertozzi, Stan- ley Osher, and Bao Wang. GRAND ++: Graph neural diffusion with a source term. In International Conference on Learning Representations, 2022

  250. [259]

    Variational learning of inducing variables in sparse Gaussian processes

    Michalis Titsias. Variational learning of inducing variables in sparse Gaussian processes. In David van Dyk and Max Welling, editors,Proceedings of the Twelth International Confer- ence on Artificial Intelligence and Statistics , volume 5 of Proceedings of Machine Learning Res...

  251. [261]

    Bronstein

    Jake Topping, Francesco Di Giovanni, Benjamin Paul Chamberlain, Xiaowen Dong, and Michael M. Bronstein. Understanding over-squashing and bottlenecks on graphs via cur- vature. In International Conference on Learning Representations, 2022

  252. [262]

    Learning to for- get: Bayesian time series forecasting using recurrent sparse spectrum signature Gaussian processes

    Csaba Tóth, Masaki Adachi, Michael A Osborne, and Harald Oberhauser. Learning to for- get: Bayesian time series forecasting using recurrent sparse spectrum signature Gaussian processes. arXiv preprint arXiv:2412.19727, 2024

  253. [263]

    Seq2Tens: An efficient representa- tion of sequences by low-rank tensor projections

    Csaba Tóth, Patric Bonnier, and Harald Oberhauser. Seq2Tens: An efficient representa- tion of sequences by low-rank tensor projections. InInternational Conference on Learning Representations, 2021

  254. [264]

    A user’s guide to KSig: GPU- accelerated computation of the signature kernel.arXiv preprint arXiv:2501.07145, 2025

    Csaba Tóth, Danilo Jr Dela Cruz, and Harald Oberhauser. A user’s guide to KSig: GPU- accelerated computation of the signature kernel.arXiv preprint arXiv:2501.07145, 2025

  255. [265]

    Capturing graphs with hypo-elliptic diffusions

    Csaba Tóth, Darrick Lee, Celia Hacker, and Harald Oberhauser. Capturing graphs with hypo-elliptic diffusions. In Advances in Neural Information Processing Systems , pages 38803–38817, 2022. 197

  256. [266]

    Bayesian learning from sequential data using Gaus- sian processes with signature covariances

    Csaba Tóth and Harald Oberhauser. Bayesian learning from sequential data using Gaus- sian processes with signature covariances. In International Conference on Machine Learn- ing, pages 9548–9560, 2020

  257. [267]

    Random Fourier signature features

    Csaba Tóth, Harald Oberhauser, and Zoltan Szabo. Random Fourier signature features. arXiv preprint arXiv:2311.12214, 2023

  258. [268]

    Some mathematical notes on three-mode factor analysis

    Ledyard R Tucker. Some mathematical notes on three-mode factor analysis. Psychome- trika, 31(3):279–311, 1966

  259. [269]

    Autoregressive forests for multivari- ate time series modeling

    Kerem Sinan Tuncel and Mustafa Gokce Baydogan. Autoregressive forests for multivari- ate time series modeling. Pattern Recognition, 73:202–215, 2018

  260. [270]

    Why are big data matrices approximately low rank? SIAM Journal on Mathematics of Data Science, 2019

    Madeleine Udell and Alex Townsend. Why are big data matrices approximately low rank? SIAM Journal on Mathematics of Data Science, 2019

  261. [271]

    Streaming kernel PCA with ˜O(pn) random features

    Enayat Ullah, Poorya Mianjy , Teodor Vanislavov Marinov, and Raman Arora. Streaming kernel PCA with ˜O(pn) random features. In Advances in Neural Information Processing Systems, 2018

  262. [272]

    Wellner.Weak Convergence and Empirical Processes: With Appli- cations to Statistics

    AW van der Vaart and J. Wellner.Weak Convergence and Empirical Processes: With Appli- cations to Statistics. Springer, 1996

  263. [273]

    Convolutional Gaus- sian processes

    Mark van der Wilk, Carl Edward Rasmussen, and James Hensman. Convolutional Gaus- sian processes. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vish- wanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems , volume 30, 2017

  264. [274]

    Analysis and geometry on groups

    N Th Varopoulos. Analysis and geometry on groups. Cambridge Tracts in Math. , 100, 1992

  265. [275]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017

  266. [276]

    Graph attention networks

    Petar Veliˇckovi´c, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In International Conference on Learning Rep- resentations, 2018

  267. [277]

    Vershynin

    R. Vershynin. High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press, 2018. 198

  268. [278]

    Order matters: Sequence to se- quence for sets

    Oriol Vinyals, Samy Bengio, and Manjunath Kudlur. Order matters: Sequence to se- quence for sets. In ICLR, 2016

  269. [279]

    Local random feature approximations of the Gaus- sian kernel

    Jonas Wacker and Maurizio Filippone. Local random feature approximations of the Gaus- sian kernel. Procedia Computer Science, 207:987–996, 2022

  270. [280]

    Improved random features for dot product kernels

    Jonas Wacker, Motonobu Kanagawa, and Maurizio Filippone. Improved random features for dot product kernels. arXiv preprint arXiv:2201.08712, 2022

  271. [281]

    Complex-to-real sketches for ten- sor products with applications to the polynomial kernel

    Jonas Wacker, Ruben Ohana, and Maurizio Filippone. Complex-to-real sketches for ten- sor products with applications to the polynomial kernel. In International Conference on Artificial Intelligence and Statistics, pages 5181–5212, 2023

  272. [282]

    Watson, and George Karypis

    Nikil Wale, Ian A. Watson, and George Karypis. Comparison of descriptor spaces for chemical compound retrieval and classification. Knowledge and Information Systems , 14(3):347–375, March 2008

  273. [283]

    A review of deep learning for renewable energy forecasting

    Huaizhi Wang, Zhenxing Lei, Xian Zhang, Bin Zhou, and Jianchun Peng. A review of deep learning for renewable energy forecasting. Energy Conversion and Management , 198:111799, 2019

  274. [284]

    Z. Wang, W . Yan, and T . Oates. Time series classification from scratch with deep neural networks: A strong baseline. In 2017 International Joint Conference on Neural Networks (IJCNN), pages 1578–1585, 2017

  275. [285]

    Transformers in time series: A survey .arXiv preprint arXiv:2202.07125, 2022

    Qingsong Wen, Tian Zhou, Chaoli Zhang, Weiqi Chen, Ziqing Ma, Junchi Yan, and Liang Sun. Transformers in time series: A survey .arXiv preprint arXiv:2202.07125, 2022

  276. [286]

    A multi- horizon quantile recurrent forecaster

    Ruofeng Wen, Kari Torkkola, Balakrishnan Narayanaswamy , and Dhruv Madeka. A multi- horizon quantile recurrent forecaster. arXiv preprint arXiv:1711.11053, 2017

  277. [287]

    Using the Nyström method to speed up kernel machines

    Christopher Williams and Matthias Seeger. Using the Nyström method to speed up kernel machines. In Advances in Neural Information Processing Systems, pages 682–688, 2000

  278. [288]

    Fast kernel learn- ing for multidimensional pattern extrapolation.Advances in Neural Information Processing Systems (NeurIPS), 27, 2014

    Andrew G Wilson, Elad Gilboa, Arye Nehorai, and John P Cunningham. Fast kernel learn- ing for multidimensional pattern extrapolation.Advances in Neural Information Processing Systems (NeurIPS), 27, 2014

  279. [289]

    Deep kernel learn- ing

    Andrew G Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P Xing. Deep kernel learn- ing. In Artificial intelligence and statistics (AISTATS), pages 370–378. PMLR, 2016. 199

  280. [290]

    Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P . Xing. Deep kernel learning. In Arthur Gretton and Christian C. Robert, editors, Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, volume 51 of Proceedings of Machine ...

  281. [291]

    Random Walks on Infinite Graphs and Groups

    Wolfgang Woess. Random Walks on Infinite Graphs and Groups . Cambridge Tracts in Mathematics. Cambridge University Press, Cambridge, 2000

  282. [292]

    Random warping series: A random features method for time-series embedding

    Lingfei Wu, Ian En-Hsu Yen, Jinfeng Yi, Fangli Xu, Qi Lei, and Michael Witbrock. Random warping series: A random features method for time-series embedding. In International Conference on Artificial Intelligence and Statistics, pages 793–802, 2018

  283. [293]

    Group normalization

    Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European confer- ence on computer vision (ECCV), pages 3–19, 2018

  284. [294]

    Representing long-range context for graph neural networks with global at- tention

    Zhanghao Wu, Paras Jain, Matthew Wright, Azalia Mirhoseini, Joseph E Gonzalez, and Ion Stoica. Representing long-range context for graph neural networks with global at- tention. Advances in Neural Information Processing Systems, 34:13266–13279, 2021

  285. [295]

    Louis-Pascal A. C. Xhonneux, Meng Qu, and Jian Tang. Continuous graph neural net- works. In Proceedings of the 37th International Conference on Machine Learning, ICML’20, pages 10432–10441. JMLR.org, July 2020

  286. [296]

    How powerful are graph neural networks? In International Conference on Learning Representations, 2019

    Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? In International Conference on Learning Representations, 2019

  287. [297]

    Representation learning on graphs with jumping knowledge networks

    Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka. Representation learning on graphs with jumping knowledge networks. In International Conference on Machine Learning, pages 5453–5462. PMLR, 2018

  288. [298]

    Hamilton, and Jure Leskovec

    Rex Ying, Jiaxuan You, Christopher Morris, Xiang Ren, William L. Hamilton, and Jure Leskovec. Hierarchical graph representation learning with differentiable pooling. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, pages 48...

  289. [299]

    Tensor Spaces and Exterior Algebra

    Takeo Yokonuma. Tensor Spaces and Exterior Algebra. American Mathematical Society , 1992

  290. [300]

    Orthogonal random features

    Felix Xinnan X Yu, Ananda Theertha Suresh, Krzysztof M Choromanski, Daniel N Holtmann-Rice, and Sanjiv Kumar. Orthogonal random features. In Advances in Neu- ral Information Processing Systems, pages 1975–1983, 2016. 200

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.