REVIEW 3 major objections 6 minor 298 references
Scalable Machine Learning Algorithms using Path Signatures
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Path signatures, once too costly for machine learning, can be embedded into scalable Gaussian processes, deep networks, and graph diffusions while retaining their universal approximation guarantees.
desk verdict A well-written thesis compiling five published papers, with a clear intro to path signatures but no new central result; the scalability theory is honestly flagged as partial. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the path signature $S(x) = (1, S^1(x), S^2(x), \ldots)$, the sequence of iterated integrals of a path, together with the tensor algebra $T((V))$ whose non-commutative product stitches together increments. In the discrete setting the order-$p$ signature features sum over subsequences of increments, generalizing string kernels. This machinery does three kinds of work: it gives a universal feature map for sequences, it defines a kernel by inner products in the tensor algebra, and it supports cheap rank-1 linear functionals that make the feature map computable without ever forming the full signature tensor.
What would settle it
Take a sequence classification task whose labels change under time reparameterization and evaluate order-1 discretized signature features without the added time coordinate: if predictions stay identical for all reparameterized inputs, the claimed universality of those features fails. Alternatively, on a large-scale sequence dataset, compare the Random Fourier Signature Feature kernel with the exact signature kernel: if increasing the number of random features does not drive the approximation error toward zero as the concentration results predict, the central scalability claim is undermined.
Extended reading notes
Core claim
The central claim is that one mathematical object, the path signature, can serve as a common algebraic backbone for several scalable machine-learning pipelines. Signature inner products define Gaussian-process covariances; iterating rank-1 tensor functionals of signatures defines the deep sequence layer Seq2Tens; a tensor-valued hypo-elliptic Laplacian turns graph diffusion into a mechanism that summarizes random-walk histories; random Fourier projections approximate the signature kernel with concentration guarantees; and a decay parameter in the same feature space gives a forgetting mechanism for probabilistic forecasting. The thesis reports that these signature-based models are consistently competitive with strong baselines across time-series classification, mortality prediction, generative imputation, long-range graph tasks, and multi-horizon forecasting, and often outperform them.
Load-bearing premise
The whole scalability story rests on the assumption that low-rank and random-feature truncations of the signature preserve enough of its universal expressive power; the thesis itself concedes that only partial results exist for iterations of low-rank approximations, and that order-1 discretized signatures need a time coordinate to be universal.
Editorial extensions
If this is right
- Signature covariances give Gaussian processes calibrated uncertainty on time-series classification, with the signature GP ranking ahead of other GP baselines and competitive with frequentist classifiers on accuracy.
- Low-rank signature layers (Seq2Tens) can be grafted onto existing convolutional and variational-autoencoder models, improving accuracy, mortality prediction, and missing-data imputation.
- The hypo-elliptic graph Laplacian yields graph and node features that characterize random-walk history, improving long-range graph classification without global attention or quadratic node interactions.
- Random Fourier Signature Features reduce the signature kernel's quadratic cost in sequence length and sample size, extending signature methods to datasets with millions of sequences.
- Recurrent Sparse Spectrum Signature Gaussian Processes provide scalable, probabilistic multi-horizon forecasts with an adaptive context length that outperforms plain Gaussian processes and competes with deep-learning forecasters.
Reading between the lines
- The order-1 discretized signature is universal only when a time coordinate is added; a practical recipe that follows from the thesis is to always include such a coordinate, otherwise the theoretical guarantees do not transfer to low-rank or random-feature models.
- If the low-rank and random-feature truncations prove as expressive as the full signature in further tests, signature-based features could become a default first step for sequence and graph data, analogous to how polynomial features are used in classical regression.
- The hypo-elliptic diffusion view connects graph learning to sub-Riemannian geometry, suggesting that over-squashing could be studied through the geometry of lifted random walks rather than only through message-passing depth, a direction the thesis leaves open.
- Random Fourier Signature Features could be used beyond prediction, for example to scale signature-based maximum mean discrepancy tests for the distribution of paths to large sample sizes, which would extend the thesis's claims without requiring new theory.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The thesis integrates path signatures with scalable machine learning pipelines, presenting six chapters: an introduction to signatures, Gaussian processes with signature covariances (GPSig), the Seq2Tens framework for deep sequence modelling, graph models based on hypo-elliptic diffusions (G2TN), Random Fourier Signature Features, and Recurrent Sparse Spectrum Signature Gaussian Processes. Each chapter is largely self-contained and adapts previously published peer-reviewed papers. The core claim is that signature-based features can be embedded into GPs, neural networks, kernels, and graph diffusion models while retaining the theoretical guarantees of signatures and achieving scalability through low-rank approximations, random features, and sparse variational inference, leading to state-of-the-art empirical performance on several benchmarks.
Significance. If the central claim holds, the thesis would make a substantial contribution: signature-based models would be credible competitors to RNN, transformer, and GNN baselines for sequential and structured data, with theoretical backing that is rare in this area. The thesis has real strengths: it ships concrete algorithms with complexity analyses, provides external benchmark comparisons and ablations on multiple datasets, and presents some self-contained theoretical results, notably the regularity theorem in Chapter 2 and the concentration results in Chapter 5. The algebraic extension of signatures to graphs is a genuine innovation that connects random-walk diffusion to tensor-valued operators. However, the decisive theoretical statements that would justify the scalability claim are not fully contained in the manuscript: the universality of Seq2Tens is stated with a deferred proof, the recovery of expressiveness through stacking low-rank layers is explicitly deferred and conceded to be only partially resolved, and the characterization theorem for graphs is deferred to a separate article.
major comments (3)
- [§3.3, Theorem 3.1 and §3.4] The universality theorem for Seq2Tens is stated only informally: Theorem 3.1 refers to a universal map φ with "a lift that satisfies some mild constraints," and the proof is deferred to [263, App. B]. The restriction to rank-1 functionals in Section 3.3 explicitly narrows the hypothesis class, and the claim that stacking low-rank sequence-to-sequence transforms recovers expressiveness is not proved in the thesis: Section 3.4 says a rigorous quantitative statement is provided in [263, App. C], while the Declarations state that Appendices A–C were omitted and deferred to that article. Because the scalability of the entire model family rests on this step, the manuscript should either include a precise statement and proof of the stacking result or clearly mark this as an open conjecture.
- [§4.6, Conclusion] The Chapter 4 conclusion admits that "for the iterations of low-rank approximations only partial results exist." This is load-bearing for the graph contribution: Theorem 4.3 gives an efficient algorithm for rank-1 functionals, but the claim that composing such layers achieves the expressiveness of general high-degree functionals is used to motivate the G2TN architecture and is not established here. The characterization result for graphs (Theorem 4.2) is also deferred to [265, App. E]. The text should either supply the missing proof or explicitly frame the expressiveness of stacked low-rank layers as empirical rather than theoretical.
- [§1.2.6 and Algorithms 1–2] The universality of order-1 discretized signatures is stated to require a time coordinate: Section 1.2.6 says the p=1 case is universal on sequences "given the existence a time coordinate which encodes the position within the sequence." However, Algorithms 1 and 2 make time augmentation optional, and the experiments in Chapters 2, 3, and 6 do not consistently state whether this coordinate was included. If any experiment omitted the time coordinate, the stated theoretical guarantees (universality, and the characterization results that rely on it for graphs) do not apply to that configuration. The manuscript should state explicitly, for each experimental setup, whether time augmentation was used.
minor comments (6)
- [§1.1.1] In Definition 1.2, the sentence "U× V is unique up to isomorphism" should refer to the tensor product U⊗V rather than the product set U× V.
- [§1.3.2] Algorithms 1 and 2 contain duplicated line numbers (e.g., multiple lines labelled 9 and 10 in Algorithm 1), which makes the listings hard to follow.
- [§1.3.2] There is a typo in "polynomail complexity" that should read "polynomial complexity."
- [§2.2.1] In Section 2.2.1, "nuiscance function" should be "nuisance function."
- [§2.4.1] In Section 2.4.1, "choosen" should be "chosen," and later in the same section "maximising" is inconsistently spelled.
- [§5 and Chapter 6] The thesis would benefit from a table summarizing the computational complexity of all proposed methods in one place; currently the complexity statements are scattered across Chapters 2–6.
Circularity Check
Scalability guarantee for stacked low-rank signature layers is deferred to a coauthored appendix and admitted to be partial; empirical benchmarks remain external.
-
self citation load bearing
[Section 3.4, with supporting statements in the Declarations and Chapter 4 Conclusion]
"Making precise how the stacking of such low-rank seq2seq transformations approximates general functions requires more tools from algebra, and we provide a rigorous quantitative statement in [263, App. C]. Here, we just appeal to the analogy made with adding depth in neural networks mentioned earlier and empirically validate this in our experiments in Section 3.4."
The thesis's scalable signature models depend on the claim that stacking rank-1/low-rank functionals recovers the universal expressiveness that is explicitly lost by the rank restriction in Section 3.3. The only proof support offered for this load-bearing claim is [263, App. C], a paper written by the thesis author and collaborators; the Declarations state that Appendices A-C were written by Patric and 'omitted here and deferred to the article,' so the proof is not independently presented or verified in this thesis.
full rationale
The core empirical claims are tested against external benchmarks (UCR time series, NCI graphs, PhysioNet), and no fitted parameter or learned hyperparameter is renamed as a prediction; the benchmark results are not constructed from the target quantities. However, the theoretical premise that low-rank stacking preserves universal expressiveness is load-bearing for the scalability story and is not proven in the thesis: Section 3.3 restricts the hypothesis class to rank-1 functionals, Section 3.4 defers the quantitative stacking statement to [263, App. C], the Declarations state that the relevant appendices were written by a collaborator and omitted from the thesis, and the Chapter 4 Conclusion admits only partial results exist for iterations of low-rank approximations. Because [263] is the author's own coauthored ICLR paper, this is a load-bearing self-citation rather than an independently verified mathematical fact. That warrants a score of 4: partial circularity in the theoretical foundation, but the central empirical demonstrations remain self-contained and externally evaluated.
Assumptions & free parameters
free parameters (5)
- signature level weights sigma_m^2 =
learned
- time-augmentation scale tau =
learned via ARD
- static kernel bandwidth =
chosen by validation
- tensor truncation degree M =
M=4 in Chapter 2, 2-4 in Chapter 3, 2 in Chapter 4
- number of random features =
not specified in available text
assumptions (6)
- standard math Chen's identity and shuffle identity for path signatures
- standard math Universality of signature linear functionals via Stone-Weierstrass
- domain assumption Order-1 discretized signature is universal only with an added time coordinate
- standard math Kernel trick for signature kernels and its PDE representation
- standard math Concentration inequalities for random Fourier features
- domain assumption Low-rank truncations preserve the approximation properties of signatures
Cite this review
Pith. "Pith review of Scalable Machine Learning Algorithms using Path Signatures." pith.science (2026). https://pith.science/paper/5DKNQZSR
@misc{pith2026250617634,
author = {Pith},
title = {Pith review of: Scalable Machine Learning Algorithms using Path Signatures},
year = {2026},
howpublished = {\url{https://pith.science/paper/5DKNQZSR}},
note = {Machine review of arXiv:2506.17634}
}
read the original abstract
The interface between stochastic analysis and machine learning is a rapidly evolving field, with path signatures - iterated integrals that provide faithful, hierarchical representations of paths - offering a principled and universal feature map for sequential and structured data. Rooted in rough path theory, path signatures are invariant to reparameterization and well-suited for modelling evolving dynamics, long-range dependencies, and irregular sampling - common challenges in real-world time series and graph data. This thesis investigates how to harness the expressive power of path signatures within scalable machine learning pipelines. It introduces a suite of models that combine theoretical robustness with computational efficiency, bridging rough path theory with probabilistic modelling, deep learning, and kernel methods. Key contributions include: Gaussian processes with signature kernel-based covariance functions for uncertainty-aware time series modelling; the Seq2Tens framework, which employs low-rank tensor structure in the weight space for scalable deep modelling of long-range dependencies; and graph-based models where expected signatures over graphs induce hypo-elliptic diffusion processes, offering expressive yet tractable alternatives to standard graph neural networks. Further developments include Random Fourier Signature Features, a scalable kernel approximation with theoretical guarantees, and Recurrent Sparse Spectrum Signature Gaussian Processes, which combine Gaussian processes, signature kernels, and random features with a principled forgetting mechanism for multi-horizon time series forecasting with adaptive context length. We hope this thesis serves as both a methodological toolkit and a conceptual bridge, and provides a useful reference for the current state of the art in scalable, signature-based learning for sequential and structured data.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
https://pubchem.ncbi.nlm.nih.gov/
Pubchem. https://pubchem.ncbi.nlm.nih.gov/
-
[2]
Bayesian model selection of lithium-ion battery models via Bayesian quadrature
Masaki Adachi, Yannick Kuhn, Birger Horstmann, Arnulf Latz, Michael A Osborne, and David A Howey . Bayesian model selection of lithium-ion battery models via Bayesian quadrature. IFAC-PapersOnLine, 56(2):10521–10526, 2023
2023
-
[3]
Adler and J.E
R.J. Adler and J.E. Taylor. Random Fields and Geometry. Springer Monographs in Math- ematics. Springer New York, 2009
2009
-
[4]
Predicting battery end of life from solar off-grid system field data using machine learning
Antti Aitio and David A Howey . Predicting battery end of life from solar off-grid system field data using machine learning. Joule, 5(12):3204–3220, 2021
2021
-
[5]
Learning scalable deep kernels with recurrent structure
Maruan Al-Shedivat, Andrew G Wilson, Yunus Saatchi, Zhiting Hu, and Eric P Xing. Learning scalable deep kernels with recurrent structure. Journal of Machine Learning Research (JMLR), 18(82):1–37, 2017
2017
-
[6]
GluonTS: Probabilistic and neural time series mod- eling in Python
Alexander Alexandrov, Konstantinos Benidis, Michael Bohlke-Schneider, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, Danielle C Maddix, Syama Rangapuram, David Salinas, Jasper Schulz, et al. GluonTS: Probabilistic and neural time series mod- eling in Python. Journal of Machine Learning Research (JMLR), 21(116):1–6, 2020
2020
-
[7]
On the bottleneck of graph neural networks and its practical implications
Uri Alon and Eran Yahav. On the bottleneck of graph neural networks and its practical implications. In International Conference on Learning Representations, 2021
2021
-
[8]
Signature-based valida- tion of real-world economic scenarios.ASTIN Bulletin: The Journal of the IAA, 54(2):410– 440, 2024
Hervé Andrès, Alexandre Boumezoued, and Benjamin Jourdain. Signature-based valida- tion of real-world economic scenarios.ASTIN Bulletin: The Journal of the IAA, 54(2):410– 440, 2024
2024
Show all 298 references
-
[9]
UCI machine learning repository , 2007
Arthur Asuncion, David Newman, et al. UCI machine learning repository , 2007
2007
-
[10]
Random Fourier features for kernel ridge regression: Approximation 175 bounds and statistical guarantees
Haim Avron, Michael Kapralov, Cameron Musco, Christopher Musco, Ameya Velingker, and Amir Zandieh. Random Fourier features for kernel ridge regression: Approximation 175 bounds and statistical guarantees. InInternational Conference on Machine Learning, pages 253–262, 2017
2017
-
[11]
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016
2016 arXiv
-
[12]
Sharp analysis of low-rank kernel matrix approximations
Francis Bach. Sharp analysis of low-rank kernel matrix approximations. InConference on Learning Theory, pages 185–209, 2013
2013
-
[13]
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015, 2015
2015
-
[14]
Ex- ploiting the past and the future in protein secondary structure prediction.Bioinformatics, 15(11):937–946, 1999
Pierre Baldi, Søren Brunak, Paolo Frasconi, Giovanni Soda, and Gianluca Pollastri. Ex- ploiting the past and the future in protein secondary structure prediction.Bioinformatics, 15(11):937–946, 1999
1999
-
[15]
Interaction networks for learning about objects, relations and physics.Advances in neural information processing systems, 29, 2016
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics.Advances in neural information processing systems, 29, 2016
2016
-
[16]
Understanding prob- abilistic sparse Gaussian process approximations
Matthias Bauer, Mark van der Wilk, and Carl Edward Rasmussen. Understanding prob- abilistic sparse Gaussian process approximations. In Advances in neural information pro- cessing systems, pages 1533–1541, 2016
2016
-
[17]
Multivariate Time Series Classification Datasets
Mustafa Baydogan. Multivariate Time Series Classification Datasets . http:// mustafabaydogan.com, 2015. [Accessed: 2020-02-05]
2015
-
[18]
Learning a symbolic representation for multivariate time series classification
Mustafa Gokce Baydogan and George Runger. Learning a symbolic representation for multivariate time series classification. Data Mining and Knowledge Discovery, 29(2):400– 422, 2015
2015
-
[19]
Mustafa Gokce Baydogan and George C. Runger. Time series representation and simi- larity based on local autopatterns. Data Mining and Knowledge Discovery , 30:476–509, 2015
2015
-
[20]
A Bayesian wilcoxon signed-rank test based on the Dirichlet process
Alessio Benavoli, Giorgio Corani, Francesca Mangili, Marco Zaffalon, and Fabrizio Rug- geri. A Bayesian wilcoxon signed-rank test based on the Dirichlet process. InInternational conference on machine learning, pages 1026–1034, 2014
2014
-
[21]
Reproducing kernel Hilbert spaces in proba- bility and statistics
Alain Berlinet and Christine Thomas-Agnan. Reproducing kernel Hilbert spaces in proba- bility and statistics. Springer Science & Business Media, 2011. 176
2011
-
[22]
Deep signature transforms
Patric Bonnier, Patrick Kidger, Imanol Perez Arribas, Cristopher Salvi, and Terry Lyons. Deep signature transforms. 33rd Conference on Neural Information Processing Systems, NeurIPS, 2019
2019
-
[23]
Graph kernels: State-of-the-art and future challenges
Karsten Borgwardt, Elisabetta Ghisu, Felipe Llinares-López, Leslie O’Bray , and Bas- tian Rieck. Graph kernels: State-of-the-art and future challenges. arXiv preprint arXiv:2011.03854, 2020
2011 arXiv
-
[24]
Concentration Inequalities: A Nonasymptotic Theory of Independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart. Concentration Inequalities: A Nonasymptotic Theory of Independence. Oxford University Press, 2013
2013
-
[25]
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges.arXiv preprint arXiv:2104.13478, 2021
Michael M Bronstein, Joan Bruna, Taco Cohen, and Petar Veli ˇckovi´c. Geometric deep learning: Grids, groups, graphs, geodesics, and gauges.arXiv preprint arXiv:2104.13478, 2021
2021 arXiv
-
[26]
Spectral networks and locally connected networks on graphs
Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun. Spectral networks and locally connected networks on graphs. arXiv preprint arXiv:1312.6203, 2013
2013 arXiv
-
[27]
Generating financial markets with signatures
H Buehler, B Horvath, T Lyons, I Perez, and B Wood. Generating financial markets with signatures. Risk.net, 2021
2021
-
[28]
Time series forecasting for healthcare diagnosis and prognostics with the focus on cardiovascular diseases
C Bui, N Pham, A Vo, A Tran, A Nguyen, and T Le. Time series forecasting for healthcare diagnosis and prognostics with the focus on cardiovascular diseases. In International Conference on the Development of Biomedical Engineering in Vietnam (BME) , pages 809–
-
[29]
Bui, Josiah Yan, and Richard E
Thang D. Bui, Josiah Yan, and Richard E. Turner. A Unifying Framework for Gaussian Process Pseudo-Point Approximations using Power Expectation Propagation. Journal of Machine Learning Research, 18(104):1–72, 2017
2017
-
[30]
Learning with SGD and random features
Luigi Carratino, Alessandro Rudi, and Lorenzo Rosasco. Learning with SGD and random features. In Advances in Neural Information Processing Systems , pages 10192–10203, 2018
2018
-
[31]
Eckart-young
J Douglas Carroll and Jih-Jie Chang. Analysis of individual differences in multidimen- sional scaling via an n-way generalization of “Eckart-young” decomposition. Psychome- trika, 35(3):283–319, 1970
1970
-
[32]
Weighted signature kernels
Thomas Cass, Terry Lyons, and Xingcheng Xu. Weighted signature kernels. The Annals of Applied Probability, 34(1A):585–626, 2024. 177
2024
-
[33]
Lecture notes on rough paths and applications to machine learning
Thomas Cass and Cristopher Salvi. Lecture notes on rough paths and applications to machine learning. arXiv preprint arXiv:2404.06583, 2024
2024 arXiv
-
[34]
Variational multinomial logit Gaussian process
Kian Ming A Chai. Variational multinomial logit Gaussian process. Journal of Machine Learning Research, 13(Jun):1745–1808, 2012
2012
-
[35]
Orlicz random Fourier features
Linda Chamakh, Emmanuel Gobet, and Zoltán Szabó. Orlicz random Fourier features. Journal of Machine Learning Research, 21(145):1–37, 2020
2020
-
[36]
Bronstein
Benjamin Paul Chamberlain, James Rowbottom, Davide Eynard, Francesco Di Giovanni, Xiaowen Dong, and Michael M. Bronstein. Beltrami flow and neural diffusion on graphs. CoRR, abs/2110.09443, 2021
2021 arXiv
-
[37]
Rowbottom, Maria I
Benjamin Paul Chamberlain, James R. Rowbottom, Maria I. Gorinova, Stefan Webb, Emanuele Rossi, and Michael M. Bronstein. GRAND: Graph neural diffusion. In ICML, 2021
2021
-
[38]
Re- current neural networks for multivariate time series with missing values
Zhengping Che, Sanjay Purushotham, Kyunghyun Cho, David Sontag, and Yan Liu. Re- current neural networks for multivariate time series with missing values. Scientific re- ports, 8(1):6085, 2018
2018
-
[39]
Convolutional kernel networks for graph-structured data
Dexiong Chen, Laurent Jacob, and Julien Mairal. Convolutional kernel networks for graph-structured data. In International Conference on Machine Learning , pages 1576–
-
[40]
Integration of paths—a faithful representation of paths by non- commutative formal power series
Kuo-Tsai Chen. Integration of paths—a faithful representation of paths by non- commutative formal power series. Trans. Amer. Math. Soc., 89:395–407, 1958
1958
-
[41]
On the equivalence between graph isomorphism testing and function approximation with gnns
Zhengdao Chen, Soledad Villar, Lei Chen, and Joan Bruna. On the equivalence between graph isomorphism testing and function approximation with gnns. Advances in neural information processing systems, 32, 2019
2019
-
[42]
A primer on the signature method in machine learning
Ilya Chevyrev and Andrey Kormilitzin. A primer on the signature method in machine learning. arXiv preprint arXiv:1603.03788, 2016
2016
-
[43]
Persistence Paths and Signature Features in Topological Data Analysis
Ilya Chevyrev, Vidit Nanda, and Harald Oberhauser. Persistence Paths and Signature Features in Topological Data Analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–1, 2018
2018
-
[44]
Signature moments to characterize laws of stochastic processes
Ilya Chevyrev and Harald Oberhauser. Signature moments to characterize laws of stochastic processes. Journal of Machine Learning Research, 23(176):1–42, 2022. 178
2022
-
[45]
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, B van Merrienboer, Caglar Gulcehre, F Bougares, H Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. In Conference on Empirical Methods in Natural Language Processing (EMNLP 2014), 2014
2014
-
[46]
Francois Chollet et al. Keras. https://github.com/fchollet/keras, 2015
2015
-
[47]
The geometry of random features
Krzysztof Choromanski, Mark Rowland, Tamas Sarlos, Vikas Sindhwani, Richard Turner, and Adrian Weller. The geometry of random features. In International Conference on Artificial Intelligence and Statistics, pages 1–9, 2018
2018
-
[48]
The unreasonable effective- ness of structured random orthogonal embeddings
Krzysztof Choromanski, Mark Rowland, and Adrian Weller. The unreasonable effective- ness of structured random orthogonal embeddings. InInternational Conference on Neural Information Processing Systems, pages 218–227, 2017
2017
-
[49]
Recycling randomness with structure for sublinear time kernel expansions
Krzysztof Choromanski and Vikas Sindhwani. Recycling randomness with structure for sublinear time kernel expansions. InInternational Conference on Machine Learning, pages 2502–2510, 2016
2016
-
[50]
Hybrid random features
Krzysztof Marcin Choromanski, Han Lin, Haoxian Chen, Arijit Sehanobish, Yuanzhe Ma, Deepali Jain, Jake Varley , Andy Zeng, Michael S Ryoo, Valerii Likhosherstov, Dmitry Kalashnikov, Vikas Sindhwani, and Adrian Weller. Hybrid random features. In Inter- national Conference on Le...
2022
-
[51]
Tensor networks for dimensionality reduction and large-scale optimization: Part 1 low-rank tensor decompositions
Andrzej Cichocki, Namgil Lee, Ivan Oseledets, Anh-Huy Phan, Qibin Zhao, and Danilo P Mandic. Tensor networks for dimensionality reduction and large-scale optimization: Part 1 low-rank tensor decompositions. Foundations and Trends® in Machine Learning, 9(4- 5):249–429, 2016
2016
-
[52]
Sk-tree: a systematic malware detection algorithm on streaming trees via the signature kernel
Thomas Cochrane, Peter Foster, Varun Chhabra, Maud Lemercier, Terry Lyons, and Cristo- pher Salvi. Sk-tree: a systematic malware detection algorithm on streaming trees via the signature kernel. In 2021 IEEE International Conference on Cyber Security and Resilience (CSR), pages...
2021
-
[53]
On the expressive power of deep learning: A tensor analysis
Nadav Cohen, Or Sharir, and Amnon Shashua. On the expressive power of deep learning: A tensor analysis. In Conference on learning theory, pages 698–728, 2016
2016
-
[54]
Inference suboptimality in variational autoencoders
Chris Cremer, Xuechen Li, and David Duvenaud. Inference suboptimality in variational autoencoders. In Proceedings of the 35th International Conference on Machine Learning , pages 1078–1086, 2018. 179
2018
-
[55]
Discrete-time signatures and randomness in reservoir computing
Christa Cuchiero, Lukas Gonon, Lyudmila Grigoryeva, Juan-Pablo Ortega, and Josef Te- ichmann. Discrete-time signatures and randomness in reservoir computing. IEEE Trans- actions on Neural Networks and Learning Systems, 33(11):6321–6330, 2022
2022
-
[56]
On the mathematical foundations of learning
Felipe Cucker and Steve Smale. On the mathematical foundations of learning. Bulletin of the American mathematical society, 39(1):1–49, 2002
2002
-
[57]
Bonilla, Pietro Michiardi, and Maurizio Filippone
Kurt Cutajar, Edwin V . Bonilla, Pietro Michiardi, and Maurizio Filippone. Random feature expansions for deep Gaussian processes. InInternational Conference on Machine Learning (ICML), pages 884–893, 2017
2017
-
[58]
Fast global alignment kernels
Marco Cuturi. Fast global alignment kernels. In International Conference on Machine Learning, pages 929–936, 2011
2011
-
[60]
Autoregressive kernels for time series
Marco Cuturi and Arnaud Doucet. Autoregressive kernels for time series. arXiv preprint arXiv:1101.0673, 2011
2011 arXiv
-
[61]
Deep Gaussian processes
Andreas Damianou and Neil Lawrence. Deep Gaussian processes. InArtificial Intelligence and Statistics, pages 207–215, 2013
2013
-
[62]
Gaussian quadrature for kernel fea- tures
Tri Dao, Christopher M De Sa, and Christopher Ré. Gaussian quadrature for kernel fea- tures. Advances in Neural Information Processing Systems (NeurIPS), 30, 2017
2017
-
[63]
The UCR time series archive
Hoang Anh Dau, Anthony Bagnall, Kaveh Kamgar, Chin-Chia Michael Yeh, Yan Zhu, Shaghayegh Gharghabi, Chotirat Ann Ratanamahatana, and Eamonn Keogh. The UCR time series archive. IEEE/CAA Journal of Automatica Sinica, 6(6):1293–1305, 2019
2019
-
[64]
Alexander G. de G. Matthews, Mark van der Wilk, Tom Nickson, Keisuke Fujii, Alexis Boukouvalas, Pablo León-Villagrá, Zoubin Ghahramani, and James Hensman. Gpflow: A Gaussian process library using TensorFlow. Journal of Machine Learning Research , 18:40:1–40:6, 2017
2017
-
[65]
Convolutional neural networks on graphs with fast localized spectral filtering
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. Advances in neural information processing systems, 29, 2016
2016
-
[66]
Statistical comparisons of classifiers over multiple data sets
Janez Demšar. Statistical comparisons of classifiers over multiple data sets. Journal of Machine learning research, 7(Jan):1–30, 2006. 180
2006
-
[67]
Group representations in probability and statistics
Persi Diaconis. Group representations in probability and statistics. Lecture notes- monograph series, 11:i–192, 1988
1988
-
[68]
Time-warping invariants of multidimensional time series
Joscha Diehl, Kurusch Ebrahimi-Fard, and Nikolas Tapia. Time-warping invariants of multidimensional time series. Acta Applicandae Mathematicae, 170(1):265–290, 2020
2020
-
[69]
Generalized iterated-sums sig- natures
Joscha Diehl, Kurusch Ebrahimi-Fard, and Nikolas Tapia. Generalized iterated-sums sig- natures. Journal of Algebra, 632:801–824, 2023
2023
-
[70]
Probabilistic recurrent state-space models
Andreas Doerr, Christian Daniel, Martin Schiegg, Nguyen-Tuong Duy , Stefan Schaal, Marc Toussaint, and Trimpe Sebastian. Probabilistic recurrent state-space models. In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Ma- chine Learni...
2018
-
[71]
Probabilistic recurrent state-space models
Andreas Doerr, Christian Daniel, Martin Schiegg, Nguyen-Tuong Duy , Stefan Schaal, Marc Toussaint, and Trimpe Sebastian. Probabilistic recurrent state-space models. In Interna- tional Conference on Machine Learning (ICML), pages 1280–1289. PMLR, 2018
2018
-
[72]
Incorporating Nesterov Momentum into Adam
Timothy Dozat. Incorporating Nesterov Momentum into Adam. In International Confer- ence on Learning Representations, 2015
2015
-
[73]
UCI machine learning repository , 2017
Dheeru Dua and Casey Graff. UCI machine learning repository , 2017
2017
-
[74]
The sizes of compact subsets of Hilbert space and continuity of Gaus- sian processes
Richard M Dudley . The sizes of compact subsets of Hilbert space and continuity of Gaus- sian processes. Journal of Functional Analysis, 1(3):290–330, 1967
1967
-
[75]
Cambridge University Press, 2002
Richard M Dudley .Real analysis and probability, volume 74. Cambridge University Press, 2002
2002
-
[76]
Approximate Bayesian computation with path signatures
Joel Dyer, Patrick Cannon, and Sebastian M Schmon. Approximate Bayesian computation with path signatures. InThe 40th Conference on Uncertainty in Artificial Intelligence, 2024
2024
-
[77]
Cannon, and Sebastian M
Joel Dyer, Patrick W . Cannon, and Sebastian M. Schmon. Amortised likelihood-free in- ference for expensive time-series simulators with signatured ratio estimation. InInterna- tional Conference on Artificial Intelligence and Statistics, pages 11131–11144, 2022
2022
-
[78]
Identifi- cation of Gaussian process state space models.Advances in Neural Information Processing Systems (NeurIPS), 30, 2017
Stefanos Eleftheriadis, Tom Nicholson, Marc Deisenroth, and James Hensman. Identifi- cation of Gaussian process state space models.Advances in Neural Information Processing Systems (NeurIPS), 30, 2017. 181
2017
-
[79]
Ahmed A. A. Elhag, Gabriele Corso, Hannes Stärk, and Michael M. Bronstein. Graph anisotropic diffusion. In ICLR 2022 Workshop on Geometrical and Topological Representa- tion Learning, 2022
2022
-
[80]
A fair comparison of graph neural networks for graph classification
Federico Errica, Marco Podda, Davide Bacciu, and Alessio Micheli. A fair comparison of graph neural networks for graph classification. In International Conference on Learning Representations, 2019
2019
-
[81]
Random feature mapping with signed circulant matrix projection
Chang Feng, Qinghua Hu, and Shizhong Liao. Random feature mapping with signed circulant matrix projection. In International Joint Conference on Artificial Intelligence , page 3490–3496, 2015
2015
-
[82]
Framing RNN as a kernel method: A neural ODE approach
Adeline Fermanian, Pierre Marion, Jean-Philippe Vert, and Gérard Biau. Framing RNN as a kernel method: A neural ODE approach. In Advances in Neural Information Processing Systems, pages 3121–3134, 2021
2021
-
[83]
GP-VAE: Deep probabilistic time series imputation
Vincent Fortuin, Dmitry Baranchuk, Gunnar Rätsch, and Stephan Mandt. GP-VAE: Deep probabilistic time series imputation. In International Conference on Artificial Intelligence and Statistics, pages 1651–1661. PMLR, 2020
2020
-
[84]
Variational Gaussian process state-space models
Roger Frigola, Yutian Chen, and Carl Edward Rasmussen. Variational Gaussian process state-space models. Advances in Neural Information Processing Systems (NeurIPS) , 27, 2014
2014
-
[85]
Bayesian inference and learning in Gaussian process state-space models with particle MCMC
Roger Frigola, Fredrik Lindsten, Thomas B Schön, and Carl Edward Rasmussen. Bayesian inference and learning in Gaussian process state-space models with particle MCMC. In C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger, editors,Ad- vances in Neural I...
2013
-
[86]
Bayesian inference and learning in Gaussian process state-space models with particle MCMC
Roger Frigola, Fredrik Lindsten, Thomas B Schön, and Carl Edward Rasmussen. Bayesian inference and learning in Gaussian process state-space models with particle MCMC. Ad- vances in Neural Information Processing Systems (NeurIPS), 26, 2013
2013
-
[87]
A course on rough paths
Peter K Friz and Martin Hairer. A course on rough paths. Springer, 2020
2020
-
[88]
Sriperumbudur
Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Bharath K. Sriperumbudur. Characteristic kernels on groups and semigroups. In Advances in Neural Information Pro- cessing Systems, page 473–480, 2008. 182
2008
-
[89]
Improving the Gaussian process sparse spectrum approx- imation by representing uncertainty in frequency inputs
Yarin Gal and Richard Turner. Improving the Gaussian process sparse spectrum approx- imation by representing uncertainty in frequency inputs. In International Conference on Machine Learning (ICML), pages 655–664, 2015
2015
-
[90]
Rank and symmetries of signature tensors
Francesco Galuppi and Pierpaola Santarsiero. Rank and symmetries of signature tensors. arXiv preprint arXiv:2407.20405, 2024
2024 arXiv
-
[91]
GPytorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration
Jacob Gardner, Geoff Pleiss, Kilian Q Weinberger, David Bindel, and Andrew G Wilson. GPytorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration. Advances in Neural Information Processing Systems (NeurIPS), 31, 2018
2018
-
[92]
Probabilistic forecasting with spline quantile function rnns
Jan Gasthaus, Konstantinos Benidis, Yuyang Wang, Syama Sundar Rangapuram, David Salinas, Valentin Flunkert, and Tim Januschowski. Probabilistic forecasting with spline quantile function rnns. In International Conference on Artificial Intelligence and Statistics (AISTAS), pages...
1901
-
[93]
Principe de moindre action, propagation de la chaleur et estimees sous elliptiques sur certains groupes nilpotents
Bernard Gaveau. Principe de moindre action, propagation de la chaleur et estimees sous elliptiques sur certains groupes nilpotents. Acta Mathematica, 139(none):95–153, January 1977
1977
-
[94]
Bayesian non-parametrics and the probabilistic approach to mod- elling
Zoubin Ghahramani. Bayesian non-parametrics and the probabilistic approach to mod- elling. Philosophical transactions. Series A, Mathematical, physical, and engineering sci- ences, 371:20110553, 02 2013
2013
-
[95]
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley , Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry . In International conference on machine learning, pages 1263–1272. PMLR, 2017
2017
-
[96]
Signatures, lipschitz-free spaces, and paths of persistence diagrams
Chad Giusti and Darrick Lee. Signatures, lipschitz-free spaces, and paths of persistence diagrams. SIAM Journal on Applied Algebra and Geometry, 7(4):828–866, 2023
2023
-
[97]
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery . Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007
2007
-
[98]
Webb, Rob Hyndman, and Pablo Montero-Manso
Rakshitha Wathsadini Godahewa, Christoph Bergmeir, Geoffrey I. Webb, Rob Hyndman, and Pablo Montero-Manso. Monash time series forecasting archive. In Thirty-fifth Con- ference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021
2021
-
[99]
Components of a new research resource for complex physiologic signals
AL Goldberger, LAN Amaral, L Glass, JM Hausdorff, P Ch Ivanov, RG Mark, JE Mietus, GB Moody , CK Peng, and HE Stanley . Components of a new research resource for complex physiologic signals. PhysioBank, PhysioToolkit, and Physionet, 2000. 183
2000
-
[100]
Concentration inequalities for poly- nomials inα-sub-exponential random variables
Friedrich Götze, Holger Sambale, and Arthur Sinulis. Concentration inequalities for poly- nomials inα-sub-exponential random variables. Electronic Journal of Probability, 26:1 – 22, 2021
2021
-
[101]
Sparse arrays of signatures for online character recognition
Benjamin Graham. Sparse arrays of signatures for online character recognition. arXiv preprint arXiv:1308.0371, 2013
2013 arXiv
-
[102]
An introduction to long-memory time series models and fractional differencing
Clive WJ Granger and Roselyne Joyeux. An introduction to long-memory time series models and fractional differencing. Journal of time series analysis, 1(1):15–29, 1980
1980
-
[103]
Graph neural networks in tensorflow and keras with spektral [application notes]
Daniele Grattarola and Cesare Alippi. Graph neural networks in tensorflow and keras with spektral [application notes]. IEEE Computational Intelligence Magazine, 16(1):99– 106, 2021
2021
-
[104]
Hybrid speech recognition with deep bidirectional lstm
Alex Graves, Navdeep Jaitly , and Abdel-rahman Mohamed. Hybrid speech recognition with deep bidirectional lstm. In 2013 IEEE workshop on automatic speech recognition and understanding, pages 273–278. IEEE, 2013
2013
-
[105]
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. Speech recognition with deep recurrent neural networks. In2013 IEEE international conference on acoustics, speech and signal processing, pages 6645–6649. IEEE, 2013
2013
-
[106]
Framewise phoneme classification with bidirec- tional lstm and other neural network architectures
Alex Graves and Jürgen Schmidhuber. Framewise phoneme classification with bidirec- tional lstm and other neural network architectures. Neural Networks, 18(5):602–610, 2005
2005
-
[107]
A kernel two-sample test
Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola. A kernel two-sample test. Journal of Machine Learning Research, 13(Mar):723– 773, 2012
2012
-
[108]
Heat kernel and analysis on manifolds, volume 47
Alexander Grigoryan. Heat kernel and analysis on manifolds, volume 47. American Math- ematical Soc., 2009
2009
-
[109]
node2vec: Scalable feature learning for networks
Aditya Grover and Jure Leskovec. node2vec: Scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining, pages 855–864, 2016
2016
-
[110]
Uniqueness for the Signature of a path of bounded variation and the reduced path group
Ben Hambly and Terry Lyons. Uniqueness for the Signature of a path of bounded variation and the reduced path group. Ann. of Math. (2), 171(1):109–167, 2010
2010
-
[111]
Hambly and Terry J
Ben M. Hambly and Terry J. Lyons. Uniqueness for the signature of a path of bounded variation and the reduced path group. Annals of Mathematics, 171(1):109–167, 2010. 184
2010
-
[112]
Inductive representation learning on large graphs
Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. Advances in neural information processing systems, 30, 2017
2017
-
[113]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016
2016
-
[114]
Heather and Benjamin Chain
James M. Heather and Benjamin Chain. The sequence of sequencers: The history of sequencing dna. Genomics, 107(1):1–8, 2016
2016
-
[115]
James Hensman, Alexander G. de G. Matthews, and Zoubin Ghahramani. Scalable varia- tional gaussian process classification. In Guy Lebanon and S. V . N. Vishwanathan, editors, AISTATS, volume 38 of JMLR Workshop and Conference Proceedings. JMLR.org, 2015
2015
-
[116]
Variational fourier features for Gaus- sian processes
James Hensman, Nicolas Durrande, and Arno Solin. Variational fourier features for Gaus- sian processes. Journal of Machine Learning Research, 18(151):1–52, 2018
2018
-
[117]
Lawrence
James Hensman, Nicoló Fusi, and Neil D. Lawrence. Gaussian processes for big data. CoRR, abs/1309.6835, 2013
2013 arXiv
-
[118]
Gaussian processes for big data
James Hensman, Nicolò Fusi, and Neil D Lawrence. Gaussian processes for big data. In Uncertainty in Artificial Intelligence (UAI), 2013
2013
-
[119]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems (NeurIPS), 33:6840–6851, 2020
2020
-
[120]
Long short-term memory .Neural computation, 9(8):1735–1780, 1997
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory .Neural computation, 9(8):1735–1780, 1997
1997
-
[121]
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory . Neural Computa- tion, 9(8):1735–1780, 1997
1997
-
[122]
Hypoelliptic second order differential equations
Lars Hörmander. Hypoelliptic second order differential equations. Acta Mathematica, 119:147–171, 1967
1967
-
[123]
Set functions for time series
Max Horn, Michael Moor, Christian Bock, Bastian Rieck, and Karsten Borgwardt. Set functions for time series. In ICML, 2020
2020
-
[124]
Approximation capabilities of multilayer feedforward networks
Kurt Hornik. Approximation capabilities of multilayer feedforward networks. Neural networks, 4(2):251–257, 1991. 185
1991
-
[125]
Deep multimodal multilinear fusion with high-order polynomial pooling
Ming Hou, Jiajia Tang, Jianhai Zhang, Wanzeng Kong, and Qibin Zhao. Deep multimodal multilinear fusion with high-order polynomial pooling. InAdvances in Neural Information Processing Systems, pages 12136–12145, 2019
2019
-
[126]
Forecasting: principles and practice
RJ Hyndman. Forecasting: principles and practice. OTexts, 2018
2018
-
[127]
Overcoming mean-field approximations in recurrent Gaussian process models
Alessandro Davide Ialongo, Mark Van Der Wilk, James Hensman, and Carl Edward Ras- mussen. Overcoming mean-field approximations in recurrent Gaussian process models. In International conference on machine learning (ICML), pages 2931–2940. PMLR, 2019
2019
-
[128]
Deep learning for time series classification: a review
Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhassane Idoumghar, and Pierre-Alain Muller. Deep learning for time series classification: a review. Data Min- ing and Knowledge Discovery, 33(4):917–963, July 2019
2019
-
[129]
Non-parametric online market regime detection and regime clustering for multidimensional and path-dependent data structures
Zacharia Issa and Blanka Horvath. Non-parametric online market regime detection and regime clustering for multidimensional and path-dependent data structures. arXiv preprint arXiv:2306.15835, 2023
2023 arXiv
-
[130]
Non-adversarial training of neural sdes with signature kernel scores
Zacharia Issa, Blanka Horvath, Maud Lemercier, and Cristopher Salvi. Non-adversarial training of neural sdes with signature kernel scores. In Advances in Neural Information Processing Systems, volume 36, pages 11102–11126, 2023
2023
-
[131]
On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables
Leon Isserlis. On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables. Biometrika, 12(1/2):134–139, 1918
1918
-
[132]
Neural tangent kernel: Convergence and generalization in neural networks.Advances in neural information processing systems, 31, 2018
Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks.Advances in neural information processing systems, 31, 2018
2018
-
[133]
Parametric Gaussian process regres- sors
Martin Jankowiak, Geoff Pleiss, and Jacob Gardner. Parametric Gaussian process regres- sors. In International conference on machine learning (ICML), 2020
2020
-
[134]
Gaussian Hilbert Spaces
Svante Janson. Gaussian Hilbert Spaces. Cambridge University Press, 1997
1997
-
[135]
Extensions of Lips- chitz maps into Banach spaces
William B Johnson, Joram Lindenstrauss, and Gideon Schechtman. Extensions of Lips- chitz maps into Banach spaces. Israel Journal of Mathematics, 54(2):129–138, 1986
1986
-
[136]
Multivariate lstm-fcns for time series classification
Fazle Karim, Somshubra Majumdar, Houshang Darabi, and Samuel Harford. Multivariate lstm-fcns for time series classification. Neural Networks, 116:237–245, 2019
2019
-
[137]
Generalized random shapelet forests
Isak Karlsson, Panagiotis Papapetrou, and Henrik Boström. Generalized random shapelet forests. Data Min. Knowl. Discov., 30(5):1053–1085, September 2016. 186
2016
-
[138]
Multiparameter processes: an introduction to random fields
Davar Khoshnevisan. Multiparameter processes: an introduction to random fields. Springer Science & Business Media, 2002
2002
-
[139]
Generalized tensor models for recurrent neural networks
Valentin Khrulkov, Oleksii Hrinchuk, and Ivan Oseledets. Generalized tensor models for recurrent neural networks. arXiv preprint arXiv:1901.10801, 2019
1901 arXiv
-
[140]
Deep signature transforms
Patrick Kidger, Patric Bonnier, Imanol Perez Arribas, Cristopher Salvi, and Terry Lyons. Deep signature transforms. In H. Wallach, H. Larochelle, A. Beygelzimer, F . d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32 , pages 3099...
2019
-
[141]
Neural SDEs as infinite-dimensional GANs
Patrick Kidger, James Foster, Xuechen Li, Harald Oberhauser, and Terry J Lyons. Neural SDEs as infinite-dimensional GANs. In International Conference on Machine Learning , pages 5453–5463, 2021
2021
-
[142]
Signatory: differentiable computations of the signa- ture and logsignature transforms, on both CPU and GPU
Patrick Kidger and Terry Lyons. Signatory: differentiable computations of the signa- ture and logsignature transforms, on both CPU and GPU. In International Conference on Learning Representations, 2021
2021
-
[143]
Adam: A method for stochastic optimization
Diederik P Kingma. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014
2014 arXiv
-
[144]
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling. Auto-encoding variational Bayes. In International Conference on Learning Representations (ICLR), 2014
2014
-
[145]
Semi-Supervised Classification with Graph Convolutional Networks
Thomas Kipf and Max Welling. Semi-Supervised Classification with Graph Convolutional Networks. In International Conference on Learning Representations, 2017
2017
-
[146]
Kernels for sequentially ordered data
Franz J Király and Harald Oberhauser. Kernels for sequentially ordered data. Journal of Machine Learning Research, 20(31):1–45, 2019
2019
-
[147]
Diffusion improves graph learning
Johannes Klicpera, Stefan Weissenberger, and Stephan Günnemann. Diffusion improves graph learning. ArXiv, abs/1911.05485, 2019
1911 arXiv
-
[148]
Tensor decompositions and applications
Tamara G Kolda and Brett W Bader. Tensor decompositions and applications. SIAM review, 51(3):455–500, 2009
2009
-
[149]
Predict, refine, synthesize: Self-guiding diffusion mod- els for probabilistic time series forecasting
Marcel Kollovieh, Abdul Fatir Ansari, Michael Bohlke-Schneider, Jasper Zschiegner, Hao Wang, and Yuyang Bernie Wang. Predict, refine, synthesize: Self-guiding diffusion mod- els for probabilistic time series forecasting. Advances in Neural Information Processing Systems (NeurI...
2024
-
[150]
Diffusion kernels on graphs and other discrete structures
Risi Kondor. Diffusion kernels on graphs and other discrete structures. In ICML 2002, 2002
2002
-
[151]
Tensor regression networks
Jean Kossaifi, Zachary C Lipton, Aran Khanna, Tommaso Furlanello, and Anima Anand- kumar. Tensor regression networks. arXiv preprint arXiv:1707.08308, 2017
2017 arXiv
-
[152]
Gaussian sketching yields a JL lemma in RKHS
Samory Kpotufe and Bharath Sriperumbudur. Gaussian sketching yields a JL lemma in RKHS. In International Conference on Artificial Intelligence and Statistics , pages 3928– 3937, 2020
2020
-
[153]
Moving beyond sub-Gaussianity in high-dimensional statistics: Applications in covariance estimation and linear regression
Arun Kumar Kuchibhotla and Abhishek Chakrabortty . Moving beyond sub-Gaussianity in high-dimensional statistics: Applications in covariance estimation and linear regression. Information and Inference: A Journal of the IMA, 11(4):1389–1456, 2022
2022
-
[154]
Modeling long-and short- term temporal patterns with deep neural networks
Guokun Lai, Wei-Cheng Chang, Yiming Yang, and Hanxiao Liu. Modeling long-and short- term temporal patterns with deep neural networks. In The 41st international ACM SIGIR conference on research & development in information retrieval, pages 95–104, 2018
2018
-
[155]
Serge Lang. Algebra. Springer, 2002
2002
-
[156]
Samuel Lanthaler and Nicholas H. Nelsen. Error bounds for learning with vector-valued random features. In Advances in Neural Information Processing Systems, 2023
2023
-
[157]
Inter-domain Gaussian processes for sparse inference using inducing features
Miguel Lázaro-Gredilla and Anibal Figueiras-Vidal. Inter-domain Gaussian processes for sparse inference using inducing features. In Advances in Neural Information Processing Systems, pages 1087–1095, 2009
2009
-
[158]
Sparse spectrum Gaussian process regression
Miguel Lázaro-Gredilla, Joaquin Quinonero-Candela, Carl Edward Rasmussen, and Aníbal R Figueiras-Vidal. Sparse spectrum Gaussian process regression. Journal of Ma- chine Learning Research (JMLR), 11:1865–1881, 2010
2010
-
[159]
Fastfood-approximating kernel expansions in loglinear time
Quoc Le, Tamás Sarlós, Alex Smola, et al. Fastfood-approximating kernel expansions in loglinear time. In International Conference on Machine Learning, page 8, 2013
2013
-
[160]
Convolutional networks for images, speech, and time series
Yann LeCun, Yoshua Bengio, et al. Convolutional networks for images, speech, and time series. The handbook of brain theory and neural networks, 3361(10):1995, 1995
1995
-
[161]
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998
1998
-
[162]
Path signatures on lie groups
Darrick Lee and Robert Ghrist. Path signatures on lie groups. arXiv preprint arXiv:2007.06633, 2020. 188
2007 arXiv
-
[163]
The signature kernel
Darrick Lee and Harald Oberhauser. The signature kernel. arXiv preprint arXiv:2305.04625, 2023
2023 arXiv
-
[164]
Self-attention graph pooling
Junhyun Lee, Inyeop Lee, and Jaewoo Kang. Self-attention graph pooling. In Interna- tional conference on machine learning, pages 3734–3743. PMLR, 2019
2019
-
[165]
Bonilla, Theodoros Damoulas, and Terry J Lyons
Maud Lemercier, Cristopher Salvi, Thomas Cass, Edwin V . Bonilla, Theodoros Damoulas, and Terry J Lyons. SigGPDE: Scaling sparse Gaussian processes on sequential data. In International Conference on Machine Learning, pages 6233–6242, 2021
2021
-
[166]
SigGPDE: Scaling sparse Gaussian processes on sequential data
Maud Lemercier, Cristopher Salvi, Thomas Cass, Edwin V Bonilla, Theodoros Damoulas, and Terry J Lyons. SigGPDE: Scaling sparse Gaussian processes on sequential data. In International conference on machine learning (ICML), pages 6233–6242. PMLR, 2021
2021
-
[167]
Trigonometric quadrature Fourier features for scalable Gaussian process regression
Kevin Li, Max Balakirsky , and Simon Mak. Trigonometric quadrature Fourier features for scalable Gaussian process regression. In International Conference on Artificial Intelligence and Statistics (AISTAS), pages 3484–3492. PMLR, 2024
2024
-
[168]
Disentangled sequential autoencoder
Yingzhen Li and Stephan Mandt. Disentangled sequential autoencoder. arXiv preprint arXiv:1803.02991, 2018
2018 arXiv
-
[169]
Gated graph sequence neural networks
Yujia Li, Richard Zemel, Marc Brockschmidt, and Daniel Tarlow. Gated graph sequence neural networks. In Proceedings of ICLR’16, 2016
2016
-
[170]
Towards a unified analysis of random Fourier features
Zhu Li, Jean-Francois Ton, Dino Oglic, and Dino Sejdinovic. Towards a unified analysis of random Fourier features. In International Conference on Machine Learning, pages 3905– 3914, 2019
2019
-
[171]
Learning representations from imperfect time series data via tensor rank regularization
Paul Pu Liang, Zhun Liu, Yao-Hung Hubert Tsai, Qibin Zhao, Ruslan Salakhutdinov, and Louis-Philippe Morency . Learning representations from imperfect time series data via tensor rank regularization. arXiv preprint arXiv:1907.01011, 2019
1907 arXiv
-
[172]
A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent
Zhenyu Liao, Romain Couillet, and Michael W Mahoney . A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent. In Advances in Neural Information Processing Systems, pages 13939–13950, 2020
2020
-
[173]
Random features for kernel approximation: A survey on algorithms, theory , and beyond.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):7128–7148, 2021
Fanghui Liu, Xiaolin Huang, Yudong Chen, and Johan AK Suykens. Random features for kernel approximation: A survey on algorithms, theory , and beyond.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):7128–7148, 2021. 189
2021
-
[174]
An intriguing failing of convolutional neural networks and the coord- conv solution
Rosanne Liu, Joel Lehman, Piero Molino, Felipe Petroski Such, Eric Frank, Alex Sergeev, and Jason Yosinski. An intriguing failing of convolutional neural networks and the coord- conv solution. In Proceedings of the 32nd International Conference on Neural Information Processing...
2018
-
[175]
Efficient low-rank multimodal fusion with modality-specific factors
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency . Efficient low-rank multimodal fusion with modality-specific factors. arXiv preprint arXiv:1806.00064, 2018
2018 arXiv
-
[176]
Text classification using string kernels
Huma Lodhi, Craig Saunders, John Shawe-Taylor, Nello Cristianini, and Chris Watkins. Text classification using string kernels. Journal of Machine Learning Research (JMLR) , 2(Feb):419–444, 2002
2002
-
[177]
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In ICLR, 2017
2017
-
[178]
Rough paths, signatures and the modelling of functions on streams
Terry Lyons. Rough paths, signatures and the modelling of functions on streams. In Proceedings of the International Congress of Mathematicians, 2014
2014
-
[179]
Terry Lyons, Michael Caruana, and Thierry Lévy .Differential Equations Driven by Rough Paths. Éc. Été Probab. St.-Flour. Springer-Verlag, Berlin Heidelberg, 2007
2007
-
[180]
Sketching the order of events
Terry Lyons and Harald Oberhauser. Sketching the order of events. arXiv preprint arXiv:1708.09708, 2017
2017 arXiv
-
[181]
Inversion of signature for paths of bounded variation
Terry Lyons and Weijun Xu. Inversion of signature for paths of bounded variation. arXiv preprint arXiv:1112.0452, 2011
2011 arXiv
-
[182]
Springer, 2007
Terry J Lyons, Michael Caruana, and Thierry Lévy .Differential equations driven by rough paths. Springer, 2007
2007
-
[183]
The M4 competi- tion: 100,000 time series and 61 forecasting methods
Spyros Makridakis, Evangelos Spiliotis, and Vassilios Assimakopoulos. The M4 competi- tion: 100,000 time series and 61 forecasting methods. International Journal of Forecast- ing, 36(1):54–74, 2020
2020
-
[184]
Provably powerful graph networks
Haggai Maron, Heli Ben-Hamu, Hadar Serviansky , and Yaron Lipman. Provably powerful graph networks. Advances in neural information processing systems, 32, 2019
2019
-
[185]
Scalable Gaussian process inference using variational meth- ods
Alexander G de G Matthews. Scalable Gaussian process inference using variational meth- ods. PhD thesis, Cambridge University , 2017. 190
2017
-
[186]
On sparse variational methods and the Kullback-Leibler divergence between stochastic processes
Alexander G de G Matthews, James Hensman, Richard Turner, and Zoubin Ghahramani. On sparse variational methods and the Kullback-Leibler divergence between stochastic processes. Journal of Machine Learning Research, 51:231–239, 2016
2016
-
[187]
Mattos, Zhenwen Dai, Andreas Damianou, Jeremy Forth, Guilherme A
César Lincoln C. Mattos, Zhenwen Dai, Andreas Damianou, Jeremy Forth, Guilherme A. Barreto, and Neil D. Lawrence. Recurrent Gaussian processes. InInternational Conference on Learning Representations (ICLR), volume 3, 2016
2016
-
[188]
Umap: Uniform manifold approximation and projection.The Journal of Open Source Software, 3(29):861, 2018
Leland McInnes, John Healy , Nathaniel Saul, and Lukas Grossberger. Umap: Uniform manifold approximation and projection.The Journal of Open Source Software, 3(29):861, 2018
2018
-
[189]
Universal kernels
Charles A Micchelli, Yuesheng Xu, and Haizhang Zhang. Universal kernels. Journal of Machine Learning Research, 7(Dec):2651–2667, 2006
2006
-
[190]
Shen-Orr, Shalev Itzkovitz, Nadav Kashtan, Dmitri B
Ron Milo, Shai S. Shen-Orr, Shalev Itzkovitz, Nadav Kashtan, Dmitri B. Chklovskii, and Uri Alon. Network motifs: simple building blocks of complex networks. Science, 298 5594:824–7, 2002
2002
-
[191]
Geometric deep learning on graphs and manifolds using mixture model cnns
Federico Monti, Davide Boscaini, Jonathan Masci, Emanuele Rodola, Jan Svoboda, and Michael M Bronstein. Geometric deep learning on graphs and manifolds using mixture model cnns. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 5115–5124, 2017
2017
-
[192]
A generalised signature method for multivariate time series feature extraction
James Morrill, Adeline Fermanian, Patrick Kidger, and Terry Lyons. A generalised signature method for multivariate time series feature extraction. arXiv preprint arXiv:2006.00873, 2020
2006 arXiv
-
[193]
Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe
Christopher Morris, Martin Ritzert, Matthias Fey , William L. Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and Leman Go Neural: Higher-Order Graph Neural Networks. Proceedings of the AAAI Conference on Artificial Intelligence , 33(01):4602–4609, July 2019
2019
-
[194]
On rings of operators
Francis J Murray and J v Neumann. On rings of operators. Annals of Mathematics, pages 116–229, 1936
1936
-
[195]
Reparameterization gradients through acceptance-rejection sampling algorithms
Christian Naesseth, Francisco Ruiz, Scott Linderman, and David Blei. Reparameterization gradients through acceptance-rejection sampling algorithms. In International Conference on Artificial Intelligence and Statistics (AISTAS), pages 489–498, 2017. 191
2017
-
[196]
Query-driven active surveying for collective classification
Galileo Namata, Ben London, Lise Getoor, Bert Huang, and U Edu. Query-driven active surveying for collective classification. In 10th International Workshop on Mining and Learning with Graphs, volume 8, page 1, 2012
2012
-
[197]
Handling in- complete heterogeneous data using vaes
Alfredo Nazabal, Pablo M Olmos, Zoubin Ghahramani, and Isabel Valera. Handling in- complete heterogeneous data using vaes. arXiv preprint arXiv:1807.03653, 2018
2018 arXiv
-
[198]
Nemenyi.Distribution-free Multiple Comparisons
P . Nemenyi.Distribution-free Multiple Comparisons. Princeton University , 1963
1963
-
[199]
Sig-Wasserstein GANs for time series generation
Hao Ni, Lukasz Szpruch, Marc Sabate-Vidales, Baoren Xiao, Magnus Wiese, and Shujian Liao. Sig-Wasserstein GANs for time series generation. In International Conference on AI in Finance, pages 1–8, 2022
2022
-
[200]
Random walk graph neural networks
Giannis Nikolentzos and Michalis Vazirgiannis. Random walk graph neural networks. Advances in Neural Information Processing Systems, 33:16211–16222, 2020
2020
-
[201]
Wavenet: A genera- tive model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. Wavenet: A genera- tive model for raw audio. arXiv preprint arXiv:1609.03499, 2016
2016 arXiv
-
[202]
Tensor-train decomposition
Ivan V Oseledets. Tensor-train decomposition. SIAM Journal on Scientific Computing , 33(5):2295–2317, 2011
2011
-
[203]
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury , Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. PyTorch: An imperative style, high-performance deep learning library . Advances in Neural Infor- mation Processing Systems ...
2019
-
[204]
Pedregosa, G
F . Pedregosa, G. Varoquaux, A. Gramfort, V . Michel, B. Thirion, O. Grisel, M. Blondel, P . Prettenhofer, R. Weiss, V . Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay . Scikit-learn: Machine learning in Python. Journal of Ma- chine L...
2011
-
[205]
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages 1532–1543, Doha, Qatar, October 2014. Association for Computati...
2014
-
[206]
Deepwalk: Online learning of so- cial representations
Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. Deepwalk: Online learning of so- cial representations. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 701–710, 2014. 192
2014
-
[207]
Satellite image time series analy- sis under time warping
Francois Petitjean, Jordi Inglada, and Pierre Gancarski. Satellite image time series analy- sis under time warping. IEEE transactions on geoscience and remote sensing, 50(8):3081– 3095, 2012
2012
-
[208]
A unifying view of sparse approximate Gaussian process regression
Joaquin Quiñonero-Candela and Carl Edward Rasmussen. A unifying view of sparse approximate Gaussian process regression. Journal of Machine Learning Research , 6(Dec):1939–1959, 2005
1939
-
[209]
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht. Random features for large-scale kernel machines. In Advances in Neural Information Processing Systems, pages 1177–1184, 2007
2007
-
[210]
Uniform approximation of functions with random bases
Ali Rahimi and Benjamin Recht. Uniform approximation of functions with random bases. In Allerton Conference on Communication, Control, and Computing, pages 555–561, 2008
2008
-
[211]
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht. Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning. Advances in Neural Information Processing Systems, pages 1313–1320, 2008
2008
-
[212]
Tensorized random projections
Beheshteh Rakhshan and Guillaume Rabusseau. Tensorized random projections. In In- ternational Conference on Artificial Intelligence and Statistics, pages 3306–3316, 2020
2020
-
[213]
Hierarchical graph neural nets can capture long-range interactions
Ladislav Rampášek and Guy Wolf. Hierarchical graph neural nets can capture long-range interactions. In 2021 IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP), pages 1–6. IEEE, 2021
2021
-
[214]
Deep state space models for time series forecasting
Syama Sundar Rangapuram, Matthias W Seeger, Jan Gasthaus, Lorenzo Stella, Yuyang Wang, and Tim Januschowski. Deep state space models for time series forecasting. Ad- vances in Neural Information Processing Systems (NeurIPS), 31, 2018
2018
-
[215]
Machine learning in Python: Main developments and technology trends in data science, machine learning, and artifi- cial intelligence
Sebastian Raschka, Joshua Patterson, and Corey Nolet. Machine learning in Python: Main developments and technology trends in data science, machine learning, and artifi- cial intelligence. Information, 11(4), 2020
2020
-
[216]
Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, 2006
2006
-
[217]
Free Lie algebras
Christophe Reutenauer. Free Lie algebras. Handbook of Algebra, 3:887–903, 2003
2003
-
[218]
Linda Preiss Rothschild and Elias M. Stein. Hypoelliptic differential operators and nilpo- tent groups. Acta Mathematica, 137:247–320, 1976
1976
-
[219]
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco. Generalization properties of learning with random features. In Advances in Neural Information Processing Systems, pages 3218–3228, 2017. 193
2017
-
[220]
Rudin.Fourier Analysis on Groups
W . Rudin.Fourier Analysis on Groups. Dover Publications, 2017
2017
-
[221]
Principles of mathematical analysis
Walter Rudin. Principles of mathematical analysis. McGraw-hill New York, 1976
1976
-
[222]
A survey on over- smoothing in graph neural networks
T Konstantin Rusch, Michael M Bronstein, and Siddhartha Mishra. A survey on over- smoothing in graph neural networks. arXiv preprint arXiv:2303.10993, 2023
2023 arXiv
-
[223]
R.A. Ryan. Introduction to Tensor Products of Banach Spaces. Springer London, 2013
2013
-
[224]
Spectral clustering of graphs with the bethe hessian
Alaa Saade, Florent Krzakala, and Lenka Zdeborová. Spectral clustering of graphs with the bethe hessian. In NIPS, 2014
2014
-
[225]
Sakoe and S
H. Sakoe and S. Chiba. Dynamic programming algorithm optimization for spoken word recognition. IEEE Transactions on Acoustics, Speech, and Signal Processing, 26(1):43–49, February 1978
1978
-
[226]
Deep Gaus- sian processes with importance-weighted variational inference
Hugh Salimbeni, Vincent Dutordoir, James Hensman, and Marc Deisenroth. Deep Gaus- sian processes with importance-weighted variational inference. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volum...
2019
-
[227]
DeepAR: Proba- bilistic forecasting with autoregressive recurrent networks
David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski. DeepAR: Proba- bilistic forecasting with autoregressive recurrent networks. International journal of fore- casting, 36(3):1181–1191, 2020
2020
-
[228]
The signature kernel is the solution of a Goursat PDE
Cristopher Salvi, Thomas Cass, James Foster, Terry Lyons, and Weixin Yang. The signature kernel is the solution of a Goursat PDE. SIAM Journal on Mathematics of Data Science , 3(3):873–899, 2021
2021
-
[229]
Some notes on concentration for α-subexponential random variables
Holger Sambale. Some notes on concentration for α-subexponential random variables. In High Dimensional Probability IX: The Ethereal Volume, pages 167–192. Springer, 2023
2023
-
[230]
Z. Sasvári. Multivariate Characteristic and Correlation Functions. De Gruyter, 2013
2013
-
[231]
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model. IEEE transactions on neural networks , 20(1):61–80, 2008
2008
-
[232]
Multivariate time series classification with weasel +muse
Patrick Schäfer and Ulf Leser. Multivariate time series classification with weasel +muse. ArXiv, abs/1711.11343, 2017. 194
2017 arXiv
-
[233]
Modeling relational data with graph convolutional networks
Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. Modeling relational data with graph convolutional networks. In European semantic web conference, pages 593–607. Springer, 2018
2018
-
[234]
Learning with Kernels: Support Vector Ma- chines, Regularization, Optimization, and Beyond
Bernhard Schölkopf and Alexander Smola. Learning with Kernels: Support Vector Ma- chines, Regularization, Optimization, and Beyond. MIT Press, 2002
2002
-
[235]
Bidirectional recurrent neural networks.IEEE trans- actions on Signal Processing, 45(11):2673–2681, 1997
Mike Schuster and Kuldip K Paliwal. Bidirectional recurrent neural networks.IEEE trans- actions on Signal Processing, 45(11):2673–2681, 1997
1997
-
[236]
Motifs for processes on networks
Alice C Schwarze and Mason A Porter. Motifs for processes on networks. SIAM Journal on Applied Dynamical Systems, 20(4):2516–2557, 2021
2021
-
[237]
Collective classification in network data
Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. Collective classification in network data. AI magazine, 29(3):93–93, 2008
2008
-
[238]
Financial time series forecasting with deep learning: A systematic literature review: 2005–2019.Applied soft computing, 90:106181, 2020
Omer Berat Sezer, Mehmet Ugur Gudelek, and Ahmet Murat Ozbayoglu. Financial time series forecasting with deep learning: A systematic literature review: 2005–2019.Applied soft computing, 90:106181, 2020
2005
-
[239]
Cambridge University Press, 2004
John Shawe-Taylor and Nello Cristianini.Kernel Methods for Pattern Analysis. Cambridge University Press, 2004
2004
-
[240]
Pitfalls of graph neural network evaluation
Oleksandr Shchur, Maximilian Mumme, Aleksandar Bojchevski, and Stephan Günne- mann. Pitfalls of graph neural network evaluation. arXiv preprint arXiv:1811.05868 , 2018
2018 arXiv
-
[241]
Sparse Gaussian processes using pseudo- inputs
Edward Snelson and Zoubin Ghahramani. Sparse Gaussian processes using pseudo- inputs. In Advances in neural information processing systems, pages 1257–1264, 2006
2006
-
[242]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. InInternational conference on machine learning (ICML), pages 2256–2265. PMLR, 2015
2015
-
[243]
On the relation between universality , characteristic kernels and RKHS embedding of measures
Bharath Sriperumbudur, Kenji Fukumizu, and Gert Lanckriet. On the relation between universality , characteristic kernels and RKHS embedding of measures. In International Conference on Artificial Intelligence and Statistics, pages 773–780, 2010
2010
-
[244]
Optimal rates for random Fourier features
Bharath Sriperumbudur and Zoltán Szabó. Optimal rates for random Fourier features. In Advances in Neural Information Processing Systems, pages 1144–1152, 2015. 195
2015
-
[245]
Universality , charac- teristic kernels and rkhs embedding of measures
Bharath K Sriperumbudur, Kenji Fukumizu, and Gert RG Lanckriet. Universality , charac- teristic kernels and rkhs embedding of measures. Journal of Machine Learning Research, 12(Jul):2389–2410, 2011
2011
-
[246]
Approximate kernel PCA: Computational versus statistical trade-off
Bharath K Sriperumbudur and Nicholas Sterge. Approximate kernel PCA: Computational versus statistical trade-off. The Annals of Statistics, 50(5):2713–2736, 2022
2022
-
[247]
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky , Ilya Sutskever, and Ruslan Salakhut- dinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014
1929
-
[248]
Support Vector Machines
Ingo Steinwart and Andreas Christmann. Support Vector Machines. Springer Science & Business Media, 2008
2008
-
[249]
Sub-riemannian geometry
Robert S Strichartz. Sub-riemannian geometry . Journal of Differential Geometry , 24(2):221–263, 1986
1986
-
[250]
Tensor random projection for low memory dimension reduction
Yiming Sun, Yang Guo, Joel A Tropp, and Madeleine Udell. Tensor random projection for low memory dimension reduction. arXiv preprint arXiv:2105.00105, 2021
2021 arXiv
-
[251]
But how does it work in theory? Linear SVM with random features
Yitong Sun, Anna Gilbert, and Ambuj Tewari. But how does it work in theory? Linear SVM with random features. In Advances in Neural Information Processing Systems, pages 3379–3388, 2018
2018
-
[252]
Translation modeling with bidirectional recurrent neural networks
Martin Sundermeyer, Tamer Alkhouli, Joern Wuebker, and Hermann Ney . Translation modeling with bidirectional recurrent neural networks. InProceedings of the 2014 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), pages 14–25, 2014
2014
-
[253]
On the error of random Fourier features
Dougal J Sutherland and Jeff Schneider. On the error of random Fourier features. In Conference on Uncertainty in Artificial Intelligence, pages 862–871, 2015
2015
-
[254]
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks. In Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 27, pages 3104–3112. Curran Associates, I...
2014
-
[255]
On kernel derivative approximation with random Fourier features
Zoltan Szabo and Bharath Sriperumbudur. On kernel derivative approximation with random Fourier features. InInternational Conference on Artificial Intelligence and Statistics (AISTAS), pages 827–836, 2019. 196
2019
-
[256]
Fourier features let networks learn high frequency functions in low dimensional domains
Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Ragha- van, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. In Advances in Neural Information...
2020
-
[257]
General tensor dis- criminant analysis and gabor features for gait recognition
Dacheng Tao, Xuelong Li, Xindong Wu, and Stephen J Maybank. General tensor dis- criminant analysis and gabor features for gait recognition. IEEE transactions on pattern analysis and machine intelligence, 29(10):1700–1715, 2007
2007
-
[258]
GRAND ++: Graph neural diffusion with a source term
Matthew Thorpe, Tan Minh Nguyen, Hedi Xia, Thomas Strohmer, Andrea Bertozzi, Stan- ley Osher, and Bao Wang. GRAND ++: Graph neural diffusion with a source term. In International Conference on Learning Representations, 2022
2022
-
[259]
Variational learning of inducing variables in sparse Gaussian processes
Michalis Titsias. Variational learning of inducing variables in sparse Gaussian processes. In David van Dyk and Max Welling, editors,Proceedings of the Twelth International Confer- ence on Artificial Intelligence and Statistics , volume 5 of Proceedings of Machine Learning Res...
2009
-
[261]
Bronstein
Jake Topping, Francesco Di Giovanni, Benjamin Paul Chamberlain, Xiaowen Dong, and Michael M. Bronstein. Understanding over-squashing and bottlenecks on graphs via cur- vature. In International Conference on Learning Representations, 2022
2022
-
[262]
Learning to for- get: Bayesian time series forecasting using recurrent sparse spectrum signature Gaussian processes
Csaba Tóth, Masaki Adachi, Michael A Osborne, and Harald Oberhauser. Learning to for- get: Bayesian time series forecasting using recurrent sparse spectrum signature Gaussian processes. arXiv preprint arXiv:2412.19727, 2024
2024 arXiv
-
[263]
Seq2Tens: An efficient representa- tion of sequences by low-rank tensor projections
Csaba Tóth, Patric Bonnier, and Harald Oberhauser. Seq2Tens: An efficient representa- tion of sequences by low-rank tensor projections. InInternational Conference on Learning Representations, 2021
2021
-
[264]
A user’s guide to KSig: GPU- accelerated computation of the signature kernel.arXiv preprint arXiv:2501.07145, 2025
Csaba Tóth, Danilo Jr Dela Cruz, and Harald Oberhauser. A user’s guide to KSig: GPU- accelerated computation of the signature kernel.arXiv preprint arXiv:2501.07145, 2025
2025 arXiv
-
[265]
Capturing graphs with hypo-elliptic diffusions
Csaba Tóth, Darrick Lee, Celia Hacker, and Harald Oberhauser. Capturing graphs with hypo-elliptic diffusions. In Advances in Neural Information Processing Systems , pages 38803–38817, 2022. 197
2022
-
[266]
Bayesian learning from sequential data using Gaus- sian processes with signature covariances
Csaba Tóth and Harald Oberhauser. Bayesian learning from sequential data using Gaus- sian processes with signature covariances. In International Conference on Machine Learn- ing, pages 9548–9560, 2020
2020
-
[267]
Random Fourier signature features
Csaba Tóth, Harald Oberhauser, and Zoltan Szabo. Random Fourier signature features. arXiv preprint arXiv:2311.12214, 2023
2023 arXiv
-
[268]
Some mathematical notes on three-mode factor analysis
Ledyard R Tucker. Some mathematical notes on three-mode factor analysis. Psychome- trika, 31(3):279–311, 1966
1966
-
[269]
Autoregressive forests for multivari- ate time series modeling
Kerem Sinan Tuncel and Mustafa Gokce Baydogan. Autoregressive forests for multivari- ate time series modeling. Pattern Recognition, 73:202–215, 2018
2018
-
[270]
Why are big data matrices approximately low rank? SIAM Journal on Mathematics of Data Science, 2019
Madeleine Udell and Alex Townsend. Why are big data matrices approximately low rank? SIAM Journal on Mathematics of Data Science, 2019
2019
-
[271]
Streaming kernel PCA with ˜O(pn) random features
Enayat Ullah, Poorya Mianjy , Teodor Vanislavov Marinov, and Raman Arora. Streaming kernel PCA with ˜O(pn) random features. In Advances in Neural Information Processing Systems, 2018
2018
-
[272]
Wellner.Weak Convergence and Empirical Processes: With Appli- cations to Statistics
AW van der Vaart and J. Wellner.Weak Convergence and Empirical Processes: With Appli- cations to Statistics. Springer, 1996
1996
-
[273]
Convolutional Gaus- sian processes
Mark van der Wilk, Carl Edward Rasmussen, and James Hensman. Convolutional Gaus- sian processes. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vish- wanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems , volume 30, 2017
2017
-
[274]
Analysis and geometry on groups
N Th Varopoulos. Analysis and geometry on groups. Cambridge Tracts in Math. , 100, 1992
1992
-
[275]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017
2017
-
[276]
Graph attention networks
Petar Veliˇckovi´c, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In International Conference on Learning Rep- resentations, 2018
2018
-
[277]
Vershynin
R. Vershynin. High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press, 2018. 198
2018
-
[278]
Order matters: Sequence to se- quence for sets
Oriol Vinyals, Samy Bengio, and Manjunath Kudlur. Order matters: Sequence to se- quence for sets. In ICLR, 2016
2016
-
[279]
Local random feature approximations of the Gaus- sian kernel
Jonas Wacker and Maurizio Filippone. Local random feature approximations of the Gaus- sian kernel. Procedia Computer Science, 207:987–996, 2022
2022
-
[280]
Improved random features for dot product kernels
Jonas Wacker, Motonobu Kanagawa, and Maurizio Filippone. Improved random features for dot product kernels. arXiv preprint arXiv:2201.08712, 2022
2022 arXiv
-
[281]
Complex-to-real sketches for ten- sor products with applications to the polynomial kernel
Jonas Wacker, Ruben Ohana, and Maurizio Filippone. Complex-to-real sketches for ten- sor products with applications to the polynomial kernel. In International Conference on Artificial Intelligence and Statistics, pages 5181–5212, 2023
2023
-
[282]
Watson, and George Karypis
Nikil Wale, Ian A. Watson, and George Karypis. Comparison of descriptor spaces for chemical compound retrieval and classification. Knowledge and Information Systems , 14(3):347–375, March 2008
2008
-
[283]
A review of deep learning for renewable energy forecasting
Huaizhi Wang, Zhenxing Lei, Xian Zhang, Bin Zhou, and Jianchun Peng. A review of deep learning for renewable energy forecasting. Energy Conversion and Management , 198:111799, 2019
2019
-
[284]
Z. Wang, W . Yan, and T . Oates. Time series classification from scratch with deep neural networks: A strong baseline. In 2017 International Joint Conference on Neural Networks (IJCNN), pages 1578–1585, 2017
2017
-
[285]
Transformers in time series: A survey .arXiv preprint arXiv:2202.07125, 2022
Qingsong Wen, Tian Zhou, Chaoli Zhang, Weiqi Chen, Ziqing Ma, Junchi Yan, and Liang Sun. Transformers in time series: A survey .arXiv preprint arXiv:2202.07125, 2022
2022 arXiv
-
[286]
A multi- horizon quantile recurrent forecaster
Ruofeng Wen, Kari Torkkola, Balakrishnan Narayanaswamy , and Dhruv Madeka. A multi- horizon quantile recurrent forecaster. arXiv preprint arXiv:1711.11053, 2017
2017 arXiv
-
[287]
Using the Nyström method to speed up kernel machines
Christopher Williams and Matthias Seeger. Using the Nyström method to speed up kernel machines. In Advances in Neural Information Processing Systems, pages 682–688, 2000
2000
-
[288]
Fast kernel learn- ing for multidimensional pattern extrapolation.Advances in Neural Information Processing Systems (NeurIPS), 27, 2014
Andrew G Wilson, Elad Gilboa, Arye Nehorai, and John P Cunningham. Fast kernel learn- ing for multidimensional pattern extrapolation.Advances in Neural Information Processing Systems (NeurIPS), 27, 2014
2014
-
[289]
Deep kernel learn- ing
Andrew G Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P Xing. Deep kernel learn- ing. In Artificial intelligence and statistics (AISTATS), pages 370–378. PMLR, 2016. 199
2016
-
[290]
Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P . Xing. Deep kernel learning. In Arthur Gretton and Christian C. Robert, editors, Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, volume 51 of Proceedings of Machine ...
2016
-
[291]
Random Walks on Infinite Graphs and Groups
Wolfgang Woess. Random Walks on Infinite Graphs and Groups . Cambridge Tracts in Mathematics. Cambridge University Press, Cambridge, 2000
2000
-
[292]
Random warping series: A random features method for time-series embedding
Lingfei Wu, Ian En-Hsu Yen, Jinfeng Yi, Fangli Xu, Qi Lei, and Michael Witbrock. Random warping series: A random features method for time-series embedding. In International Conference on Artificial Intelligence and Statistics, pages 793–802, 2018
2018
-
[293]
Group normalization
Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European confer- ence on computer vision (ECCV), pages 3–19, 2018
2018
-
[294]
Representing long-range context for graph neural networks with global at- tention
Zhanghao Wu, Paras Jain, Matthew Wright, Azalia Mirhoseini, Joseph E Gonzalez, and Ion Stoica. Representing long-range context for graph neural networks with global at- tention. Advances in Neural Information Processing Systems, 34:13266–13279, 2021
2021
-
[295]
Louis-Pascal A. C. Xhonneux, Meng Qu, and Jian Tang. Continuous graph neural net- works. In Proceedings of the 37th International Conference on Machine Learning, ICML’20, pages 10432–10441. JMLR.org, July 2020
2020
-
[296]
How powerful are graph neural networks? In International Conference on Learning Representations, 2019
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? In International Conference on Learning Representations, 2019
2019
-
[297]
Representation learning on graphs with jumping knowledge networks
Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka. Representation learning on graphs with jumping knowledge networks. In International Conference on Machine Learning, pages 5453–5462. PMLR, 2018
2018
-
[298]
Hamilton, and Jure Leskovec
Rex Ying, Jiaxuan You, Christopher Morris, Xiang Ren, William L. Hamilton, and Jure Leskovec. Hierarchical graph representation learning with differentiable pooling. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, pages 48...
2018
-
[299]
Tensor Spaces and Exterior Algebra
Takeo Yokonuma. Tensor Spaces and Exterior Algebra. American Mathematical Society , 1992
1992
-
[300]
Orthogonal random features
Felix Xinnan X Yu, Ananda Theertha Suresh, Krzysztof M Choromanski, Daniel N Holtmann-Rice, and Sanjiv Kumar. Orthogonal random features. In Advances in Neu- ral Information Processing Systems, pages 1975–1983, 2016. 200
1975
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.