Pith. sign in

REVIEW 2 major objections 5 minor 2 cited by

Gaussian processes for dynamics learning in model predictive control

T0 review · 2 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This review assembles a self-contained toolkit for Gaussian-process-based model predictive control, organizing scalable regression, uncertainty propagation, and closed-loop guarantees, and pointing to open problems in safe learning-based…

desk verdict A useful, wide-ranging review of GP-based MPC, but a missing-term error in the central propagation formula makes it unsafe as a formula reference. read the letter →

arxiv 2502.02310 v1 pith:S7TIW76N submitted 2025-02-04 eess.SY cs.SY

classification eess.SYcs.SY MSC 60G1562F1562G0562J0793B4593E20
keywords Gaussianprocessregressionmodelpredictivecontroluncertaintypropagationscalableprocesseschanceconstraintslearning-basedclosed-loopguarantees
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's aim is to turn Gaussian process-based model predictive control (GP-MPC) into a field a newcomer can enter from a map: it argues that every GP-MPC scheme is an assembly of three choices—how to make the Gaussian process regression scalable, how to propagate its uncertainty across the prediction horizon, and how to turn that uncertainty into guarantees that constraints hold with a chosen probability. It reviews each of these ingredients, including methods from outside control that have not yet been used there, and shows how they combine into one implementable MPC structure. If the review is right, a practitioner can pick a scalable GP approximation, a propagation rule, and a back-off method from surveyed options and assemble a working controller without reinventing the pipeline. It also identifies the open theoretical bottlenecks: online model updates that preserve guarantees, error bounds for scalable approximations, and uncertainty quantification tools that are still unexploited in control.

What carries the argument

The load-bearing object is the approximate stochastic optimal control problem recast as receding-horizon MPC: at each time k, minimize a cost over a control sequence while propagating the predicted state mean $\mu^x_{i+1|k} = \rho(u_{i|k}, \mu^x_{i|k}, \Sigma^x_{i|k})$ and covariance $\Sigma^x_{i+1|k} = \Phi(u_{i|k}, \mu^x_{i|k}, \Sigma^x_{i|k})$, and enforcing tightened constraints $h_j(\mu^x_{i|k}, u_{i|k}) + \nu_j(u_{i|k}, \mu^x_{i|k}, \Sigma^x_{i|k}) \leq 0$. The machinery is the decomposition itself: scalable GP approximations (subsets, inducing variables, spectral or truncated-kernel representations, numerical solvers) set the cost of evaluating $\rho$ and $\Phi$; uncertainty propagation schemes (linearization, moment matching, $\sigma$-points, Monte Carlo) define their functional form; and chance-constraint handling (bounded support, robust-in-probability, sampling) determines the back-off $\nu$. The worked example instantiates this with the Fully Independent Training Conditional (FITC) sparse GP, linearization-based propagation, and an inverse-Gaussian back-off, showing how the pieces fit.

What would settle it

Inspect the Section 4 statement that after Cholesky factorization, training computes $\log\det(K_{Z,Z})$ and $\partial\log\det(K_{Z,Z})/\partial\xi=\operatorname{tr}(K_{Z,Z}^{-1}\partial K_{Z,Z}/\partial\xi)$ in $O(N)$ operations. For a generic dense kernel matrix, computing that trace requires the diagonal of the inverse or N solves, costing $\Theta(N^2)$ after the factorization, so the $O(N)$ claim fails on a direct complexity check.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that the practical effectiveness of Gaussian-process MPC already outruns the theory, and that the theory can be advanced by explicitly separating three challenges: scalability of GP regression with growing and streaming data; propagation of uncertainty over a receding horizon despite the fact that exact distributions are non-Gaussian and intractable; and construction of closed-loop safety guarantees, which requires either robust-in-probability confidence regions or sampling-based back-offs. The review's organizing formulation is the moment-based MPC problem (45), where predicted state mean and covariance are propagated by user-chosen functions ρ and Φ and constraints are tightened by a back-off ν; every surveyed contribution is positioned by the choices it makes for these three ingredients. The paper also claims that several promising uncertainty-quantification tools—Bayesian uniform error bounds, conformal prediction, multi-step predictors—have not yet been exploited in GP-MPC and could yield less conservative safe controllers.

Load-bearing premise

The survey's value rests on its technical summaries being accurate, e.g., the claim that after Cholesky factorization the log-determinant gradient costs O(N), despite that gradient involving the trace of a matrix inverse, which is not O(N) in general.

Editorial extensions

If this is right

  • A reader can use the survey as a decision chart: choose a scalable GP method from Section 4, an uncertainty propagation rule from Section 5, and a constraint-tightening method from Section 6, and assemble the MPC problem (45) directly.
  • The moment-based formulation (45) is the common denominator of most existing GP-MPC applications, so new scalable methods will slot into the same pipeline rather than requiring bespoke controllers.
  • Rigorous closed-loop guarantees in the surveyed literature come only from robust-in-probability or sampling-based approaches; bounded-support heuristics should not be trusted beyond one-step-ahead predictions.
  • Online learning with GP-MPC improves the model but can destroy recursive feasibility and closed-loop guarantees, so the paper predicts that theoretically grounded online updates remain a key bottleneck.
  • Uncertainty-quantification tools from outside control (Bayesian uniform bounds, conformal prediction, multi-step predictors) are identified as the most promising source of less conservative safe MPC.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the complexity statements are corrected, the qualitative map of methods in Section 4 still stands; the taxonomy of scalability approaches does not depend on the specific O(N) claim.
  • The decomposition suggests a natural benchmarking protocol: fix a plant and compare each scalable GP paired with each propagation rule under the same back-off; such a systematic comparison is not provided in the review.
  • The moment-based MPC structure (45) could be extended to non-Markovian trajectory correlations in a receding-horizon setting, which the paper notes has so far only been done for shrinking horizons.
  • Conformal prediction, mentioned in the review as promising, could provide distribution-free back-offs that avoid hard-to-verify RKHS norm assumptions, making safety guarantees easier to certify in applications.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript is a review of Gaussian process-based model predictive control (GP-MPC). It introduces the GP regression background, formulates the stochastic optimal control problem approximated by MPC, and organizes the literature around three core challenges: scalability of GP regression (Section 4), uncertainty propagation over the prediction horizon (Section 5), and closed-loop safety guarantees (Section 6). The later sections survey how these ingredients are combined in practice, discuss alternative dynamics models and uncertainty quantification paradigms, and list open research problems. The authors aim to provide a self-contained toolkit for studying and advancing GP-MPC.

Significance. If its technical content is accurate, this review is a valuable and timely synthesis of a rapidly growing field. It systematically maps a large and fragmented literature, provides a useful application table, and identifies concrete open challenges. The paper's utility as a toolkit, however, depends on the correctness of the technical characterizations it presents. The errors identified below in the uncertainty-propagation formula and in the computational-complexity discussion directly affect two of the three core pillars of the review, so they must be corrected for the survey to serve its stated purpose.

major comments (2)
  1. [Section 5.1.2, Eq. (37)] The formula for the predictive variance in exact moment matching omits the prior variance term E[k(z_*, z_*)], which equals lambda for the squared-exponential kernel. Expanding E[Sigma(z_*)] + Var[mu(z_*)] gives E[k(z_*, z_*)] - Tr((K_{Z,Z} + sigma_w^2 I)^{-1} L) + beta^T L beta - (beta^T l)^2. As printed, the leading lambda is missing, so the expression systematically underestimates the propagated uncertainty; a sanity check with N=0 yields 0 instead of the prior variance lambda. Because Section 5 is one of the three core pillars of the review, this is a load-bearing error that should be corrected.
  2. [Section 4, first paragraph] The claim that, after Cholesky factorization, the training computations for log det(K_{Z,Z}) and its gradient d/dxi log det(K_{Z,Z}) = tr(K_{Z,Z}^{-1} dK_{Z,Z}/dxi) 'require O(N) operations' is inaccurate. The log-determinant is O(N) after the factorization, but the gradient involves an inner product with K_{Z,Z}^{-1}, which requires solving a linear system with N right-hand sides (O(N^3) naively, or O(N^2) using the Cholesky factors). This understates the cost of hyper-parameter optimization and should be corrected to avoid misleading readers about the scalability of full GP training.
minor comments (5)
  1. [Section 5.1.2] The matrix L is not explicitly defined; please define its entries as L_{ij} = E[k(z_i, z_*) k(z_*, z_j)] (or give the closed-form expression from the cited reference) so that Eq. (37) is reproducible.
  2. [Section 4.3.1, Eq. (30)] The normalization of the spectral density for the squared-exponential kernel is only correct for nz = 1; for nz-dimensional inputs the expression should read (2*pi*eta)^{nz/2} exp(-2*pi^2*eta ||s||^2), not sqrt(2*pi*eta^{nz}) exp(-2*pi^2*eta ||s||^2).
  3. [Table 2] The column headers 'Simulation only' and 'Online' are ambiguous; please clarify whether the former indicates that the results are simulation-based as opposed to experimental, and whether the latter indicates that online model updates are performed during operation.
  4. [Section 4.5] The reference numbering appears to jump from [383] to [343] in the text; please verify that the citations are consistent with the reference list.
  5. [Section 7.3.2] The statement that combining model (47) with the Gaussian process representation of [106] yields 'latent force models' would benefit from a direct citation to the latent force model literature, as that terminology is commonly attributed to a different line of work.

Circularity Check

0 steps flagged · score 1.0 of 10

Review is a literature synthesis; self-citations are illustrative, not load-bearing; no circular derivation found.

full rationale

The paper is a review, not a derivation: it makes no new technical predictions, fits no parameters, and presents no uniqueness theorem that would force a particular modeling choice. Its central deliverable, as stated in the abstract, is to provide a toolkit by surveying existing results on scalable Gaussian processes, uncertainty propagation, and closed-loop guarantees. The technical content in Sections 2, 4, 5, and 6 is taken from and attributed to the external literature; for example, the GP posterior equations in (9) are standard textbook results, and the uncertainty propagation methods in Section 5 are attributed to Girard, Deisenroth, and others. Self-citations such as [142], [196], and [65] appear as illustrative examples or as one among several references, but none is load-bearing: the review's conclusions do not reduce to those citations, and no alternative is excluded by an author-imported uniqueness theorem. The suspected omission in Eq. (37) of the E[k(z*,z*)] term and the Section 4 claim about O(N) training operations are technical-accuracy concerns, not circularity, because even if incorrect they do not make the paper's claims equivalent to their inputs by construction. Under the required quote-and-reduction standard, no circular step can be exhibited; the minor self-citation presence is not a circularity defect.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is a review paper, so its central claim is not a derivation but a synthesis. The ledger therefore contains no free parameters or invented entities. The main burden is on the accuracy of the survey: the comprehensiveness of the literature and the correctness of its technical summaries, which are unverified by any reproducible artifact.

assumptions (3)
  • domain assumption The surveyed literature is correctly and completely represented.
    The review does not provide a systematic search protocol or inclusion criteria; its value depends on the accuracy and representativeness of its reference list.
  • standard math Standard Gaussian process regression formulas (Section 2) are correct and apply to the reviewed MPC formulations.
    The review builds on these background results to frame the problem; they are standard and not in dispute.
  • domain assumption The complexity statements for scalable GP methods (Section 4) are accurate.
    One specific statement, that training computations for log-determinant and its gradient require O(N) operations, appears inaccurate, so this assumption is partially violated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Gaussian processes for dynamics learning in model predictive control." pith.science (2026). https://pith.science/paper/S7TIW76N

@misc{pith2026250202310,
  author       = {Pith},
  title        = {Pith review of: Gaussian processes for dynamics learning in model predictive control},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S7TIW76N}},
  note         = {Machine review of arXiv:2502.02310}
}
read the original abstract

Due to its state-of-the-art estimation performance complemented by rigorous and non-conservative uncertainty bounds, Gaussian process regression is a popular tool for enhancing dynamical system models and coping with their inaccuracies. This has enabled a plethora of successful implementations of Gaussian process-based model predictive control in a variety of applications over the last years. However, despite its evident practical effectiveness, there are still many open questions when attempting to analyze the associated optimal control problem theoretically and to exploit the full potential of Gaussian process regression in view of safe learning-based control. The contribution of this review is twofold. The first is to survey the available literature on the topic, highlighting the major theoretical challenges such as (i) addressing scalability issues of Gaussian process regression; (ii) taking into account the necessary approximations to obtain a tractable MPC formulation; (iii) including online model updates to refine the dynamics description, exploiting data collected during operation. The second is to provide an extensive discussion of future research directions, collecting results on uncertainty quantification that are related to (but yet unexploited in) optimal control, among others. Ultimately, this paper provides a toolkit to study and advance Gaussian process-based model predictive control.

Figures

Figures reproduced from arXiv: 2502.02310 by the authors.

Figure 1
Figure 1. Graphical paper roadmap Topics not covered. Finally, we would like to point out that in this review paper, we focus on Gaus￾sian process regression just for learning the dy￾namics, and do not dwell on their use for learn￾ing the constraints (e.g., for safe active learning) or the cost function (for problems of inverse optimal control or Bayesian optimization), nor for approx￾imating MPC laws – we refer to [5] for in… view at source ↗
Figure 2
Figure 2. Graphical overview of Section 2. with scaling factor λ > 0, and η > 0 denoting the kernel width. In this framework, λ plays the role of model order: because it belongs to the positive real numbers, it is more flexibly tunable than the discrete value denoting the number of features in parametric models. There are many other possible choices for the kernel, depending on the prior in￾formation to be encoded: some popul… view at source ↗
Figure 3
Figure 3. Graphical overview of Section 4. performed in O(N2 ) operations, while the required computations for training, i.e., log det (KZ,Z) and ∂ ∂ξ log det (KZ,Z) = tr  K −1 Z,Z ∂KZ,Z ∂ξ  , require O(N) operations. Note that for predictions with a fixed data-set and hyper-parameters, it suffices to pre-factorize the Gram matrix once, as it does not depend on the test points. For the posterior mean, the computational and … view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Qualitative visualization for the performance of [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Graphical overview of Section 5. independence assumption in Section 5.3. Note that in the following subsections we assume that the nominal dynamics gnom is equal to 0 and set the matrix Bd as the identity. The general case requires careful consideration of both propaga…
Figure 6
Figure 6. Figure 6: Independent one-step-ahead predictions: (Left p [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]
Figure 7
Figure 7. Figure 7: Visualization of [PITH_FULL_IMAGE:figures/full_fig_p026_7.png]
Figure 8
Figure 8. Figure 8: Graphical overview of Section 6. non-parametric estimation is assumed to be known before operation and is used to perform robust con￾trol design. In view of this, it is then key to assume that the model error g(·) lies in a compact set – a heuristic based on the practi…
Figure 9
Figure 9. Figure 9: Graphical overview of Section 7. a summary of the available papers on applications of Gaussian process-based MPC in [PITH_FULL_IMAGE:figures/full_fig_p030_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffusion-Residual Model Predictive Steering Control for Vehicle Stabilization at the Limit of Handling under Model Uncertainty

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Command-conditioned diffusion residual moments resize the MPC yaw reference and chance-tighten the handling envelope, cutting peak side-slip and recovering low-μ stability in simulation at 100 Hz.

  2. Learning, fast and slow: a two-fold algorithm for data-based model adaptation

    eess.SY 2025-07 conditional novelty 6.0 of 10

    Combining control-chart-triggered ensemble expansion with online Gaussian process residual correction lifts FIT from 69.5% to 94.2% on a district heating case.

Reference graph

Works this paper leans on

300 extracted references · 78 canonical work pages · cited by 2 Pith papers

  1. [1]

    D. Q. Mayne, Model predictive control: Recent de- velopments and future promise, Automatica 50 (2014) 2967–2986

  2. [2]

    Schwenzer, M

    M. Schwenzer, M. Ay, T. Bergs, D. Abel, Review on model predictive control: an engineering perspective, The International Journal of Advanced Manufacturing Technology 117 (2021) 1327–1349

  3. [3]

    J. B. Rawlings, D. Q. Mayne, M. Diehl, Model Predic- tive Control: Theory, Computation, and Design, Nob Hill Publishing, 2017

  4. [4]

    Hewing, K

    L. Hewing, K. P. Wabersich, M. Menner, M. N. Zeilinger, Learning-Based Model Predictive Control: Toward Safe Learning in Control, Annual Review of Control, Robotics, and Autonomous Systems 3 (2020) 269–296

  5. [5]

    Mesbah, K

    A. Mesbah, K. P. Wabersich, A. P. Schoellig, M. N. Zeilinger, S. Lucia, T. A. Badgwell, J. A. Paulson, Fusion of Machine Learning and MPC under Uncer- tainty: What Advances Are on the Horizon?, in: 2022 American Control Conference (ACC), 2022, pp. 342– 357

  6. [6]

    Bemporad, M

    A. Bemporad, M. Morari, Robust model predictive control: A survey, in: A. Garulli, A. Tesi (Eds.), Ro- bustness in identification and control, Springer, Lon- don, 1999, pp. 207–226

  7. [7]

    Kouvaritakis, M

    B. Kouvaritakis, M. Cannon, Model Predictive Con- trol, Advanced Textbooks in Control and Signal Pro- cessing, Springer International Publishing, Cham, 2016

  8. [8]

    Milanese, A

    M. Milanese, A. Vicino, Optimal estimation theory for dynamic systems with set membership uncertainty: An overview, Automatica 27 (1991) 997–1009. 39

Show all 300 references
  1. [9]

    Milanese, C

    M. Milanese, C. Novara, Set Membership identifica- tion of nonlinear systems, Automatica 40 (2004) 957– 975

  2. [10]

    Canale, L

    M. Canale, L. Fagiano, M. C. Signorile, Nonlinear model predictive control from data: a set membership approach, International Journal of Robust and Non- linear Control 24 (2014) 123–139

  3. [11]

    Mesbah, Stochastic Model Predictive Control: An Overview and Perspectives for Future Research, IEEE Control Systems Magazine 36 (2016) 30–44

    A. Mesbah, Stochastic Model Predictive Control: An Overview and Perspectives for Future Research, IEEE Control Systems Magazine 36 (2016) 30–44

  4. [12]

    Farina, L

    M. Farina, L. Giulioni, R. Scattolini, Stochastic line ar Model Predictive Control with chance constraints – A review, Journal of Process Control 44 (2016) 53–67

  5. [13]

    C. E. Rasmussen, C. K. I. Williams, Gaussian pro- cesses for machine learning, Adaptive computation and machine learning, MIT Press, Cambridge, Mass, 2006

  6. [14]

    Piche, B

    S. Piche, B. Sayyar-Rodsari, D. Johnson, M. Gerules, Nonlinear model predictive control using neural net- works, IEEE Control Systems Magazine 20 (2000) 53–62

  7. [15]

    Y. M. Ren, M. S. Alhajeri, J. Luo, S. Chen, F. Ab- dullah, Z. Wu, P. D. Christofides, A tutorial review of neural network modeling approaches for model predic- tive control, Computers & Chemical Engineering 165 (2022) 107956

  8. [16]

    Y. Bao, K. J. Chan, A. Mesbah, J. M. Velni, Learning- based adaptive-scenario-tree model predictive con- trol with improved probabilistic safety using robust Bayesian neural networks, International Journal of Robust and Nonlinear Control 33 (2023) 3312–3333

  9. [17]

    Y. Bao, H. S. Abbas, J. Mohammadpour Velni, A learning- and scenario-based MPC design for nonlinear systems in LPV framework with safety and stability guarantees, International Journal of Control 0 (2023) 1–20

  10. [18]

    Limon, J

    D. Limon, J. Calliess, J. M. Maciejowski, Learning- based Nonlinear Model Predictive Control**The au- thors would like to ackowledge to the Spanish MINECO Grant PRX15-00300 and projects DPI2013- 48243-C2-2-R and DPI2016-76493-C3-1-R as well as to the Engineering and Physical R...

  11. [19]

    J. M. Manzano, D. Limon, D. Mu˜ noz de la Pe˜ na, J. Calliess, Robust Data-Based Model Predictive Control for Nonlinear Constrained Systems, IFAC- PapersOnLine 51 (2018) 505–510

  12. [20]

    J. A. Paulson, A. Mesbah, Nonlinear Model Predictive Control with Explicit Backoffs for Stochastic Systems under Arbitrary Uncertainty, IFAC-PapersOnLine 51 (2018) 523–534

  13. [21]

    Murray-Smith, D

    R. Murray-Smith, D. Sbarbaro, C. E. Rasmussen, A. Girard, Adaptive, cautious, predictive control with gaussian process priors, IFAC Proceedings Volumes 36 (2003) 1155–1160

  14. [22]

    Kober, J

    J. Kober, J. A. Bagnell, J. Peters, Reinforcement learning in robotics: A survey, The International Jour- nal of Robotics Research 32 (2013) 1238–1274

  15. [23]

    A. S. Polydoros, L. Nalpantidis, Survey of Model- Based Reinforcement Learning: Applications on Robotics, Journal of Intelligent & Robotic Systems 86 (2017) 153–173

  16. [24]

    Garc´ ıa, F

    J. Garc´ ıa, F. Fern´ andez, A Comprehensive Survey on Safe Reinforcement Learning, Journal of Machine Learning Research 16 (2015) 1437–1480

  17. [25]

    Moerland, J

    T. Moerland, J. Broekens, A. Plaat, C. Jonker, Model- based Reinforcement Learning: A Survey, Now Foun- dations and Trentds, 2023

  18. [26]

    Plaat, W

    A. Plaat, W. Kosters, M. Preuss, High-accuracy model-based reinforcement learning, a survey, Arti- ficial Intelligence Review 56 (2023) 9541–9573

  19. [27]

    M. P. Deisenroth, C. E. Rasmussen, PILCO: A Model- Based and Data-Efficient Approach to Policy Search, in: Proceedings of the 28th International Confer- ence on International Conference on Machine Learn- ing, Omnipress, Bellevue, Washington, USA, 2011, pp. 465–472

  20. [28]

    Kamthe, M

    S. Kamthe, M. Deisenroth, Data-Efficient Reinforce- ment Learning with Probabilistic Model Predictive Control, in: Proceedings of the Twenty-First Interna- tional Conference on Artificial Intelligence and Statis- tics, PMLR, 2018, pp. 1701–1710

  21. [29]

    Berkenkamp, M

    F. Berkenkamp, M. Turchetta, A. Schoellig, A. Krause, Safe Model-based Reinforcement Learning with Stability Guarantees, in: Advances in Neural Information Processing Systems, volume 30, Curran Associates, Inc., 2017

  22. [30]

    Brunke, M

    L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, A. P. Schoellig, Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning, Annual Review of Control, Robotics, and Autonomous Systems 5 (2022) 411–444

  23. [31]

    G. S. Lima, S. Trimpe, W. M. Bessa, Sliding Mode Control with Gaussian Process Regression for Under- water Robots, Journal of Intelligent & Robotic Sys- tems 99 (2020) 487–498

  24. [32]

    T. Kim, W. Kim, S. Choi, H. Jin Kim, Path Track- ing for a Skid-steer Vehicle using Model Predictive Control with On-line Sparse Gaussian Process, IFAC- PapersOnLine 50 (2017) 5755–5760

  25. [33]

    Chowdhary, H

    G. Chowdhary, H. A. Kingravi, J. P. How, P. A. Vela, Bayesian Nonparametric Adaptive Control Us- ing Gaussian Processes, IEEE Transactions on Neural Networks and Learning Systems 26 (2015) 537–550

  26. [34]

    Gregorcic, G

    G. Gregorcic, G. Lightbody, Gaussian process ap- proaches to nonlinear modelling for control, in: In- telligent Control Systems using Computational Intel- ligence Techniques, volume 70 of IEE Control Series , London, 2005, pp. 177–217

  27. [35]

    Aˇ zman, J

    K. Aˇ zman, J. Kocijan, Fixed-structure Gaussian pro- cess model, International Journal of Systems Science 40 (2009) 1253–1262

  28. [36]

    Nguyen-Tuong, M

    D. Nguyen-Tuong, M. Seeger, J. Peters, Computed torque control with nonparametric regression models, in: 2008 American Control Conference, 2008, pp. 212– 217

  29. [37]

    Hastie, R

    T. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning, Springer Series in Statistics, Springer, New York, NY, 2009

  30. [38]

    Ljung, System identification: theory for the user, Prentice-Hall, Inc., USA, 1986

    L. Ljung, System identification: theory for the user, Prentice-Hall, Inc., USA, 1986

  31. [39]

    Matheron, Principles of geostatistics, Economic Geology 58 (1963) 1246–1266

    G. Matheron, Principles of geostatistics, Economic Geology 58 (1963) 1246–1266

  32. [40]

    Wahba, Spline Models for Observational Data, SIAM, 1990

    G. Wahba, Spline Models for Observational Data, SIAM, 1990

  33. [41]

    Poggio, F

    T. Poggio, F. Girosi, Networks for approximation and learning, Proceedings of the IEEE 78 (1990) 1481– 1497

  34. [42]

    Evgeniou, M

    T. Evgeniou, M. Pontil, T. Poggio, Regularization 40 Networks and Support Vector Machines, Advances in Computational Mathematics 13 (2000) 1–50

  35. [43]

    J. A. K. Suykens, T. V. Gestel, J. D. Brabanter, B. D. Moor, J. P. L. Vandewalle, Least Squares Support Vec- tor Machines, World Scientific, 2002

  36. [44]

    Sch¨ olkopf, A

    B. Sch¨ olkopf, A. J. Smola, Learning with Kernels: Support Vector Machines, Regularization, Optimiza- tion, and Beyond, The MIT Press, 2018

  37. [45]

    O’Hagan, Curve Fitting and Optimal Design for Prediction, Journal of the Royal Statistical Society

    A. O’Hagan, Curve Fitting and Optimal Design for Prediction, Journal of the Royal Statistical Society. Series B (Methodological) 40 (1978) 1–42

  38. [46]

    R. M. Neal, Bayesian Learning for Neural Networks, volume 118 of Lecture Notes in Statistics , Springer, New York, NY, 1996

  39. [47]

    Aronszajn, Theory of Reproducing Kernels, Trans- actions of the American Mathematical Society 68 (1950) 337–404

    N. Aronszajn, Theory of Reproducing Kernels, Trans- actions of the American Mathematical Society 68 (1950) 337–404

  40. [48]

    Berlinet, C

    A. Berlinet, C. Thomas-Agnan, Reproducing Kernel Hilbert Spaces in Probability and Statistics, Springer US, Boston, MA, 2004

  41. [49]

    Saitoh, Y

    S. Saitoh, Y. Sawano, Theory of Reproducing Ker- nels and Applications, volume 44 of Developments in Mathematics, Springer, Singapore, 2016

  42. [50]

    Gelman, Andrew, Carlin, John B., Stern, Hal S., Dunson, David B., Vehtari, Aki, Rubin, Donald B., Bayesian Data Analysis, 3 ed., Chapman and Hall/CRC, New York, 2015

  43. [51]

    D. T. Hristopulos, Gaussian Random Fields, in: D. T. Hristopulos (Ed.), Random Fields for Spatial Data Modeling: A Primer for Scientists and Engineers, Ad- vances in Geographic Information Science, Springer Netherlands, Dordrecht, 2020, pp. 245–307

  44. [52]

    C. A. Micchelli, Y. Xu, H. Zhang, Universal Kernels, Journal of Machine Learning Research 7 (2006) 2651– 2667

  45. [53]

    M¨ uller, S

    K.-R. M¨ uller, S. Mika, G. Ratsch, K. Tsuda, B. Sch¨ olkopf, An introduction to kernel-based learning algorithms, IEEE Transactions on Neural Networks 12 (2001) 181–201

  46. [54]

    Duvenaud, Automatic model construction with Gaussian processes (2014)

    D. Duvenaud, Automatic model construction with Gaussian processes (2014)

  47. [55]

    B. D. O. Anderson, J. B. Moore, Optimal Filtering, Courier Corporation, 2005

  48. [56]

    Cucker, D

    F. Cucker, D. X. Zhou, Learning Theory: An Approxi- mation Theory Viewpoint, Cambridge Monographs on Applied and Computational Mathematics, Cambridge University Press, 2007

  49. [57]

    Boucheron, G

    S. Boucheron, G. Lugosi, P. Massart, S. Boucheron, G. Lugosi, P. Massart, Concentration Inequalities: A Nonasymptotic Theory of Independence, Oxford Uni- versity Press, Oxford, New York, 2013

  50. [58]

    Smale, D.-X

    S. Smale, D.-X. Zhou, Learning Theory Estimates via Integral Operators and Their Approximations, Con- structive Approximation 26 (2007) 153–172

  51. [59]

    E. T. Maddalena, P. Scharnhorst, C. N. Jones, Deter- ministic error bounds for kernel-based learning tech- niques under bounded noise, Automatica 134 (2021) 109896

  52. [60]

    Wang, D.-X

    C. Wang, D.-X. Zhou, Optimal learning rates for least squares regularized regression with unbounded sam- pling, Journal of Complexity 27 (2011) 55–67

  53. [61]

    Guo, D.-X

    Z.-C. Guo, D.-X. Zhou, Concentration estimates for learning with unbounded sampling, Advances in Com- putational Mathematics 38 (2013) 207–223

  54. [62]

    Srinivas, A

    N. Srinivas, A. Krause, S. M. Kakade, M. W. Seeger, Information-Theoretic Regret Bounds for Gaussian Process Optimization in the Bandit Setting, IEEE Transactions on Information Theory 58 (2012) 3250– 3265

  55. [63]

    S. R. Chowdhury, A. Gopalan, On Kernelized Multi- armed Bandits, in: Proceedings of the 34th Interna- tional Conference on Machine Learning, PMLR, 2017, pp. 844–853

  56. [64]

    Fiedler, C

    C. Fiedler, C. W. Scherer, S. Trimpe, Practical and Rigorous Uncertainty Bounds for Gaussian Process Regression, Proceedings of the AAAI Conference on Artificial Intelligence 35 (2021) 7439–7447

  57. [65]

    Baggio, A

    G. Baggio, A. Car` e, A. Scampicchio, G. Pillonetto, Bayesian frequentist bounds for machine learning and system identification, Automatica 146 (2022) 110599

  58. [66]

    Shafer, V

    G. Shafer, V. Vovk, A Tutorial on Conformal Predic- tion, Journal of Machine Learning Research 9 (2008) 371–421

  59. [67]

    V. Vovk, A. Gammerman, G. Shafer, Algorithmic Learning in a Random World, Springer International Publishing, Cham, 2022

  60. [68]

    J. Lei, L. Wasserman, Distribution-free Prediction Bands for Non-parametric Regression, Journal of the Royal Statistical Society Series B: Statistical Method- ology 76 (2014) 71–96

  61. [69]

    R. J. Tibshirani, R. Foygel Barber, E. Candes, A. Ramdas, Conformal Prediction Under Covariate Shift, in: Advances in Neural Information Processing Systems, volume 32, Curran Associates, Inc., 2019

  62. [70]

    A. N. Angelopoulos, S. Bates, A Gentle Introduction to Conformal Prediction and Distribution-Free Uncer- tainty Quantification, 2022

  63. [71]

    Fontana, G

    M. Fontana, G. Zeni, S. Vantini, Conformal predic- tion: A unified review of theory and new challenges, Bernoulli 29 (2023) 1–23. Publisher: Bernoulli Society for Mathematical Statistics and Probability

  64. [72]

    Cauchois, S

    M. Cauchois, S. Gupta, A. Ali, J. C. Duchi, Robust Validation: Confident Predictions Even When Distributions Shift, Journal of the American Statistical Association 119 (2024) 3033–3044. Publisher: ASA Website eprint: https://doi.org/10.1080/01621459.2023.2298037

  65. [73]

    Gilks, Richardson, Sylvia, Spiegelhalter, Daniel (Eds.), Markov Chain Monte Carlo in Practice, Chap- man and Hall/CRC, New York, 1995

    W. Gilks, Richardson, Sylvia, Spiegelhalter, Daniel (Eds.), Markov Chain Monte Carlo in Practice, Chap- man and Hall/CRC, New York, 1995

  66. [74]

    Maritz, T

    J. Maritz, T. Lwin, Empirical Bayes Methods, Rout- ledge, London, 2018

  67. [75]

    A. N. Tikhonov, V. I. Arsenin, Solutions of Ill-posed Problems, Winston, 1977

  68. [76]

    G. S. Kimeldorf, G. Wahba, A Correspondence Be- tween Bayesian Estimation on Stochastic Processes and Smoothing by Splines, The Annals of Mathemat- ical Statistics 41 (1970) 495–502

  69. [77]

    A. Y. Aravkin, B. M. Bell, J. V. Burke, G. Pillonetto, The Connection Between Bayesian Estimation of a Gaussian Random Field and RKHS, IEEE Transac- tions on Neural Networks and Learning Systems 26 (2015) 1518–1524

  70. [78]

    Kanagawa, P

    M. Kanagawa, P. Hennig, D. Sejdinovic, B. K. Sripe- rumbudur, Gaussian Processes and Kernel Meth- ods: A Review on Connections and Equivalences, arXiv:1807.02582 [cs, stat] (2018)

  71. [79]

    M. F. Driscoll, The reproducing kernel Hilbert space structure of the sample paths of a Gaussian pro- cess, Zeitschrift f¨ ur Wahrscheinlichkeitstheorie und 41 Verwandte Gebiete 26 (1973) 309–316

  72. [80]

    Cucker, S

    F. Cucker, S. Smale, On the mathematical foundations of learning, Bulletin of the American Mathematical Society 39 (2002) 1–49

  73. [81]

    Hartikainen, Sequential Inference for Latent Tem- poral Gaussian Process Models, 2013

    J. Hartikainen, Sequential Inference for Latent Tem- poral Gaussian Process Models, 2013

  74. [82]

    McHutchon, Nonlinear Modelling and Control us- ing Gaussian Processes, PhD Thesis, University of Cambridge, 2014

    A. McHutchon, Nonlinear Modelling and Control us- ing Gaussian Processes, PhD Thesis, University of Cambridge, 2014

  75. [83]

    Frigola-Alcalde, Bayesian Time Series Learning with Gaussian Processes, PhD Thesis, University of Cambridge, 2015

    R. Frigola-Alcalde, Bayesian Time Series Learning with Gaussian Processes, PhD Thesis, University of Cambridge, 2015

  76. [84]

    Kocijan, Modelling and Control of Dynamic Sys- tems Using Gaussian Process Models, Advances in Industrial Control, Springer International Publishing, Cham, 2016

    J. Kocijan, Modelling and Control of Dynamic Sys- tems Using Gaussian Process Models, Advances in Industrial Control, Springer International Publishing, Cham, 2016

  77. [85]

    Svensson, Machine learning with state-space mod- els, Gaussian processes and Monte Carlo methods, PhD Thesis, Uppsala University, 2018

    A. Svensson, Machine learning with state-space mod- els, Gaussian processes and Monte Carlo methods, PhD Thesis, Uppsala University, 2018

  78. [86]

    S¨ arkk¨ a, Use of Gaussian Processes in System Iden- tification, in: J

    S. S¨ arkk¨ a, Use of Gaussian Processes in System Iden- tification, in: J. Baillieul, T. Samad (Eds.), Encyclo- pedia of Systems and Control, Springer International Publishing, Cham, 2021, pp. 2393–2402

  79. [87]

    Billings, Nonlinear System Identification, 1 ed., John Wiley & Sons, Ltd, 2013

    S. Billings, Nonlinear System Identification, 1 ed., John Wiley & Sons, Ltd, 2013

  80. [88]

    Akaike, Autoregressive model fitting for control, Annals of the Institute of Statistical Mathematics 23 (1971) 163–180

    H. Akaike, Autoregressive model fitting for control, Annals of the Institute of Statistical Mathematics 23 (1971) 163–180

  81. [89]

    Gregorcic, G

    G. Gregorcic, G. Lightbody, Gaussian processes for modelling of dynamic non-linear systems, Proceedings of the Irish Signals and Systems Conference (2002)

  82. [90]

    Girard, C

    A. Girard, C. E. Rasmussen, R. Murray-Smith, Gaus- sian Process priors with Uncertain Inputs: Multiple- Step-Ahead Prediction, 2002

  83. [91]

    Girard, R

    A. Girard, R. Murray-Smith, Learning a Gaussian Process Model with Uncertain Inputs, 2003

  84. [92]

    Kocijan, A

    J. Kocijan, A. Girard, B. Banko, R. Murray-Smith, Dynamic systems identification with Gaussian pro- cesses, Mathematical and Computer Modelling of Dy- namical Systems 11 (2005) 411–424

  85. [93]

    Groot, P

    P. Groot, P. Lucas, P. Bosch, Multiple-step Time Series Forecasting with Sparse Gaussian Processes, Ghent, 2011, pp. 105–112

  86. [94]

    Gutjahr, H

    T. Gutjahr, H. Ulmer, C. Ament, Sparse Gaussian Processes with Uncertain Inputs for Multi-Step Ahead Prediction, IFAC Proceedings Volumes 45 (2012) 107– 112

  87. [95]

    Krivec, G

    T. Krivec, G. Papa, J. Kocijan, Simulation of varia- tional Gaussian process NARX models with GPGPU, ISA Transactions 109 (2021) 141–151

  88. [96]

    Mattos, A

    C. Mattos, A. Damianou, G. Barreto, N. Lawrence, Latent Autoregressive Gaussian Processes Models for Robust System Identification, IFAC-PapersOnLine 49 (2016) 1121–1126

  89. [97]

    Damianou, N

    A. Damianou, N. Lawrence, Deep Gaussian Processes, in: Proceedings of the Sixteenth International Confer- ence on Artificial Intelligence and Statistics, PMLR, 2013, pp. 207–215

  90. [98]

    Mattos, Z

    C. Mattos, Z. Dai, A. Damianou, J. Forth, G. Barreto, N. D. Lawrence, Recurrent Gaussian Processes, CoRR (2015)

  91. [99]

    Lawrence, Gaussian Process Latent Variable Mod- els for Visualisation of High Dimensional Data, in: Advances in Neural Information Processing Systems, volume 16, MIT Press, 2003

    N. Lawrence, Gaussian Process Latent Variable Mod- els for Visualisation of High Dimensional Data, in: Advances in Neural Information Processing Systems, volume 16, MIT Press, 2003

  92. [100]

    Lawrence, Probabilistic Non-linear Principal Com - ponent Analysis with Gaussian Process Latent Vari- able Models, Journal of Machine Learning Research 6 (2005) 1783–1816

    N. Lawrence, Probabilistic Non-linear Principal Com - ponent Analysis with Gaussian Process Latent Vari- able Models, Journal of Machine Learning Research 6 (2005) 1783–1816

  93. [101]

    Jiang, J

    X. Jiang, J. Gao, X. Hong, Z. Cai, Gaussian Processes Autoencoder for Dimensionality Reduction, in: V. S. Tseng, T. B. Ho, Z.-H. Zhou, A. L. P. Chen, H.-Y. Kao (Eds.), Advances in Knowledge Discovery and Data Mining, Lecture Notes in Computer Science, Springer International Pu...

  94. [102]

    Takano, T

    J. Takano, T. Omori, Gaussian Process Dynamical Autoencoder Model, in: Proceedings of the 2019 3rd International Conference on Intelligent Systems, Metaheuristics & Swarm Intelligence, ISMSI ’19, As- sociation for Computing Machinery, New York, NY, USA, 2019, pp. 45–49

  95. [103]

    N. D. Lawrence, J. Qui˜ nonero-Candela, Local distance preservation in the GP-LVM through back constraints, in: Proceedings of the 23rd international conference on Machine learning - ICML ’06, ACM Press, Pittsburgh, Pennsylvania, 2006, pp. 513–520

  96. [104]

    Ferris, D

    B. Ferris, D. Fox, N. Lawrence, WiFi-SLAM using Gaussian process latent variable models, in: Proceed- ings of the 20th international joint conference on Ar- tifical intelligence, IJCAI’07, Morgan Kaufmann Pub- lishers Inc., San Francisco, CA, USA, 2007, pp. 2480– 2485

  97. [105]

    Lawrence, A

    N. Lawrence, A. Moore, Hierarchical Gaussian process latent variable models, in: Proceedings of the 24th international conference on Machine learning, ICML ’07, Association for Computing Machinery, New York, NY, USA, 2007, pp. 481–488

  98. [106]

    Hartikainen, S

    J. Hartikainen, S. S¨ arkk¨ a, Kalman filtering and smoothing solutions to temporal Gaussian process re- gression models, in: 2010 IEEE International Work- shop on Machine Learning for Signal Processing, 2010, pp. 379–384

  99. [107]

    J. M. Wang, D. J. Fleet, A. Hertzmann, Gaussian Process Dynamical Models for Human Motion, IEEE Transactions on Pattern Analysis and Machine Intel- ligence 30 (2008) 283–298

  100. [108]

    Titsias, N

    M. Titsias, N. D. Lawrence, Bayesian Gaussian Pro- cess Latent Variable Model, in: Proceedings of the Thirteenth International Conference on Artificial In- telligence and Statistics, JMLR Workshop and Con- ference Proceedings, 2010, pp. 844–851

  101. [109]

    Damianou, M

    A. Damianou, M. Titsias, N. Lawrence, Variational Gaussian Process Dynamical Systems, in: Advances in Neural Information Processing Systems, volume 24, Curran Associates, Inc., 2011

  102. [110]

    Damianou, M

    A. Damianou, M. Titsias, N. Lawrence, Variational Inference for Latent Variables and Uncertain Inputs in Gaussian Processes, Journal of Machine Learning Research 17 (2016) 1–62

  103. [111]

    D. d. Souza, D. Mesquita, J. P. Gomes, C. L. Mattos, Learning GPLVM with arbitrary kernels using the un- scented transformation, in: Proceedings of The 24th International Conference on Artificial Intelligence and Statistics, PMLR, 2021, pp. 451–459

  104. [112]

    J. Ko, D. Fox, Learning GP-BayesFilters via Gaussian process latent variable models, Autonomous Robots 30 (2011) 3–23

  105. [113]

    J. Ko, D. Fox, GP-BayesFilters: Bayesian filtering us- ing Gaussian process prediction and observation mod- els, in: 2008 IEEE/RSJ International Conference on 42 Intelligent Robots and Systems, 2008, pp. 3471–3476

  106. [114]

    R. D. Shachter, A Linear Approximation Method for Probabilistic Inference, in: R. D. Shachter, T. S. Levitt, L. N. Kanal, J. F. Lemmer (Eds.), Machine In- telligence and Pattern Recognition, volume 9 ofUncer- tainty in Artificial Intelligence , North-Holland, 1990, pp. 93–103

  107. [115]

    Simon, Optimal State Estimation: Kalman, H In- finity, and Nonlinear Approaches, Wiley-Interscience, USA, 2006

    D. Simon, Optimal State Estimation: Kalman, H In- finity, and Nonlinear Approaches, Wiley-Interscience, USA, 2006

  108. [116]

    E. Wan, R. Van Der Merwe, The unscented Kalman filter for nonlinear estimation, in: Proceedings of the IEEE 2000 Adaptive Systems for Signal Process- ing, Communications, and Control Symposium (Cat. No.00EX373), IEEE, Lake Louise, Alta., Canada, 2000, pp. 153–158

  109. [117]

    Julier, J

    S. Julier, J. Uhlmann, Unscented filtering and non- linear estimation, Proceedings of the IEEE 92 (2004) 401–422

  110. [118]

    P. Maybeck, Chapter 12 Nonlinear estimation, in: Stochastic models, estimation, and control, volume 141 of Series Mathematics in Science and Engineer- ing, academic press ed., Elsevier, 1979, pp. 212–271

  111. [119]

    Brigo, B

    D. Brigo, B. Hanzon, F. L. Gland, Approximate non- linear filtering by projection on exponential manifolds of densities, Bernoulli 5 (1999) 495–534

  112. [120]

    Gordon, D

    N. Gordon, D. Salmond, A. Smith, Novel approach to nonlinear/non-Gaussian Bayesian state estimation, IEE Proceedings F (Radar and Signal Processing) 140 (1993) 107–113

  113. [121]

    J. Ko, D. J. Klein, D. Fox, D. Haehnel, Gaussian Pro- cesses and Reinforcement Learning for Identification and Control of an Autonomous Blimp, in: Proceed- ings 2007 IEEE International Conference on Robotics and Automation, 2007, pp. 742–747

  114. [122]

    M. P. Deisenroth, M. F. Huber, U. D. Hanebeck, An- alytic moment-based Gaussian process filtering, in: Proceedings of the 26th Annual International Confer- ence on Machine Learning, ACM, Montreal Quebec Canada, 2009, pp. 225–232

  115. [123]

    M. P. Deisenroth, R. D. Turner, M. F. Huber, U. D. Hanebeck, C. E. Rasmussen, Robust Filtering and Smoothing with Gaussian Processes, IEEE Transac- tions on Automatic Control 57 (2012) 1865–1871

  116. [124]

    Deisenroth, S

    M. Deisenroth, S. Mohamed, Expectation Propagation in Gaussian Process Dynamical Systems, in: Advances in Neural Information Processing Systems, volume 25, Curran Associates, Inc., 2012

  117. [125]

    Turner, M

    R. Turner, M. Deisenroth, C. Rasmussen, State-Space Inference and Learning with Gaussian Processes, in: Proceedings of the Thirteenth International Confer- ence on Artificial Intelligence and Statistics, JMLR Workshop and Conference Proceedings, 2010, pp. 868– 875

  118. [126]

    Dempster, N

    A. Dempster, N. Laird, D. Rubin, Maximum Likeli- hood from Incomplete Data Via the EM Algorithm, Journal of the Royal Statistical Society: Series B (Methodological) 39 (1977) 1–22

  119. [127]

    Roweis, Z

    S. Roweis, Z. Ghahramani, Learning Nonlinear Dy- namical Systems Using the Expectation–Maximization Algorithm, in: Kalman Filtering and Neural Net- works, John Wiley & Sons, Ltd, 2001, pp. 175–220

  120. [128]

    T. B. Sch¨ on, A. Wills, B. Ninness, System identifica- tion of nonlinear state-space models, Automatica 47 (2011) 39–49

  121. [129]

    Frigola, F

    R. Frigola, F. Lindsten, T. B. Sch¨ on, C. E. Rasmussen, Identification of Gaussian Process State-Space Models with Particle Stochastic Approximation EM, IFAC Proceedings Volumes 47 (2014) 4097–4102

  122. [130]

    Frigola, F

    R. Frigola, F. Lindsten, T. B. Sch¨ on, C. E. Rasmussen, Bayesian Inference and Learning in Gaussian Process State-Space Models with Particle MCMC, in: Ad- vances in Neural Information Processing Systems, vol- ume 26, Curran Associates, Inc., 2013

  123. [131]

    T. B. Sch¨ on, F. Lindsten, J. Dahlin, J. W ˚ agberg, C. A. Naesseth, A. Svensson, L. Dai, Sequential Monte Carlo Methods for System Identification, IFAC- PapersOnLine 48 (2015) 775–786

  124. [132]

    Svensson, T

    A. Svensson, T. B. Sch¨ on, A. Solin, S. S¨ arkk¨ a, Non- linear state space model identification using a regular- ized basis function expansion, in: 2015 IEEE 6th In- ternational Workshop on Computational Advances in Multi-Sensor Adaptive Processing (CAMSAP), 2015, pp. 481–484

  125. [133]

    Svensson, A

    A. Svensson, A. Solin, S. S¨ arkk¨ a, T. Sch¨ on, Computa- tionally Efficient Bayesian Learning of Gaussian Pro- cess State Space Models, in: Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, PMLR, 2016, pp. 213–221

  126. [134]

    Svensson, T

    A. Svensson, T. B. Sch¨ on, A flexible state–space model for learning nonlinear dynamical systems, Automatica 80 (2017) 189–199

  127. [135]

    Frigola, Y

    R. Frigola, Y. Chen, C. E. Rasmussen, Variational Gaussian Process State-Space Models, in: Advances in Neural Information Processing Systems, volume 27, Curran Associates, Inc., 2014

  128. [136]

    Eleftheriadis, T

    S. Eleftheriadis, T. Nicholson, M. Deisenroth, J. Hen s- man, Identification of Gaussian Process State Space Models, in: Advances in Neural Information Process- ing Systems, volume 30, Curran Associates, Inc., 2017

  129. [137]

    A. D. Ialongo, M. van der Wilk, C. E. Rasmussen, Closed-form Inference and Prediction in Gaussian Pro- cess State-Space Models, in: Time Series Workshop at the 31st Conference on Neural Information Processing Systems, 2017

  130. [138]

    Doerr, C

    A. Doerr, C. Daniel, M. Schiegg, N.-T. Duy, S. Schaal, M. Toussaint, T. Sebastian, Probabilistic Recurrent State-Space Models, in: Proceedings of the 35th In- ternational Conference on Machine Learning, PMLR, 2018, pp. 1280–1289

  131. [139]

    A. D. Ialongo, M. V. D. Wilk, J. Hensman, C. E. Rasmussen, Overcoming Mean-Field Approximations in Recurrent Gaussian Process Models, in: Proceed- ings of the 36th International Conference on Machine Learning, PMLR, 2019, pp. 2931–2940

  132. [140]

    Lindinger, Variational inference for composite Gaus- sian process models, PhD Thesis, Universit¨ at Pots- dam, 2023

    J. Lindinger, Variational inference for composite Gaus- sian process models, PhD Thesis, Universit¨ at Pots- dam, 2023

  133. [141]

    Courts, A

    J. Courts, A. G. Wills, T. B. Sch¨ on, B. Ninness, Vari- ational system identification for nonlinear state-space models, Automatica 147 (2023) 110687

  134. [142]

    Arcari, M

    E. Arcari, M. V. Minniti, A. Scampicchio, A. Carron, F. Farshidian, M. Hutter, M. N. Zeilinger, Bayesian Multi-Task Learning MPC for Robotic Mobile Manip- ulation, IEEE Robotics and Automation Letters 8 (2023) 3222–3229

  135. [143]

    J. Hall, C. Rasmussen, J. Maciejowski, Modelling and control of nonlinear systems using Gaussian pro- cesses with partial model information, in: 2012 IEEE 51st IEEE Conference on Decision and Control (CDC), 43 2012, pp. 5266–5271

  136. [144]

    Liu, Y.-S

    H. Liu, Y.-S. Ong, X. Shen, J. Cai, When Gaussian Process Meets Big Data: A Review of Scalable GPs, IEEE Transactions on Neural Networks and Learning Systems 31 (2020) 4405–4423

  137. [145]

    T. F. Gonzalez, Clustering to minimize the maximum intercluster distance, Theoretical Computer Science 38 (1985) 293–306

  138. [146]

    Feder, D

    T. Feder, D. Greene, Optimal algorithms for approxi- mate clustering, in: Proceedings of the twentieth an- nual ACM symposium on Theory of computing, STOC ’88, Association for Computing Machinery, New York, NY, USA, 1988, pp. 434–444

  139. [147]

    Lawrence, M

    N. Lawrence, M. Seeger, R. Herbrich, Fast Sparse Gaussian Process Methods: The Informative Vector Machine, in: Advances in Neural Information Pro- cessing Systems, volume 15, MIT Press, 2002

  140. [148]

    Y. Lu, I. Cohen, X. S. Zhou, Q. Tian, Feature selec- tion using principal feature analysis, in: Proceedings of the 15th ACM international conference on Multime- dia, MM ’07, Association for Computing Machinery, New York, NY, USA, 2007, pp. 301–304

  141. [149]

    Keerthi, W

    S. Keerthi, W. Chu, A matching pursuit approach to sparse Gaussian process regression, in: Proceedings of the 18th International Conference on Neural Informa- tion Processing Systems, NIPS’05, MIT Press, Cam- bridge, MA, USA, 2005, pp. 643–650

  142. [150]

    Seeger, Bayesian Gaussian Process Models: PAC- Bayesian Generalisation Error Bounds and Sparse Ap- proximations (2003)

    M. Seeger, Bayesian Gaussian Process Models: PAC- Bayesian Generalisation Error Bounds and Sparse Ap- proximations (2003)

  143. [151]

    Hayashi, M

    K. Hayashi, M. Imaizumi, Y. Yoshida, On Ran- dom Subsampling of Gaussian Process Regression: A Graphon-Based Analysis, in: Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, PMLR, 2020, pp. 2055– 2065

  144. [152]

    S. Oh, Y. Xu, J. Choi, Explorative navigation of mo- bile sensor networks using sparse Gaussian processes, in: 49th IEEE Conference on Decision and Control (CDC), 2010, pp. 3851–3856

  145. [153]

    Y. Xu, J. Choi, S. Oh, Mobile Sensor Network Navi- gation Using Gaussian Processes With Truncated Ob- servations, IEEE Transactions on Robotics 27 (2011) 1118–1131

  146. [154]

    Krause, A

    A. Krause, A. Singh, C. Guestrin, Near-Optimal Sen- sor Placements in Gaussian Processes: Theory, Effi- cient Algorithms and Empirical Studies, The Journal of Machine Learning Research 9 (2008) 235–284

  147. [155]

    A. P. Bart´ ok, G. Cs´ anyi, Gaussian approximation po- tentials: A brief tutorial introduction, International Journal of Quantum Chemistry 115 (2015) 1051–1057

  148. [156]

    Wang, Y.-M

    H. Wang, Y.-M. Zhang, J.-X. Mao, Sparse Gaussian process regression for multi-step ahead forecasting of wind gusts combining numerical weather predictions and on-site measurements, Journal of Wind Engineer- ing and Industrial Aerodynamics 220 (2022) 104873

  149. [157]

    R. A. Jacobs, M. I. Jordan, S. J. Nowlan, G. E. Hinton, Adaptive Mixtures of Local Experts, Neural Compu- tation 3 (1991) 79–87

  150. [158]

    Tresp, Mixtures of Gaussian Processes, in: Ad- vances in Neural Information Processing Systems, vol- ume 13, MIT Press, 2000

    V. Tresp, Mixtures of Gaussian Processes, in: Ad- vances in Neural Information Processing Systems, vol- ume 13, MIT Press, 2000

  151. [159]

    S. E. Yuksel, J. N. Wilson, P. D. Gader, Twenty Years of Mixture of Experts, IEEE Transactions on Neural Networks and Learning Systems 23 (2012) 1177–1193

  152. [160]

    C. E. Rasmussen, Z. Ghahramani, Infinite Mixtures of Gaussian Process Experts, in: Advances in Neural In- formation Processing Systems, volume 14, MIT Press, 2001

  153. [161]

    Meeds, S

    E. Meeds, S. Osindero, An Alternative Infinite Mix- ture Of Gaussian Process Experts, in: Advances in Neural Information Processing Systems, volume 18, MIT Press, 2005

  154. [162]

    C. Yuan, C. Neubauer, Variational Mixture of Gaus- sian Process Experts, in: Advances in Neural Infor- mation Processing Systems, volume 21, Curran Asso- ciates, Inc., 2008

  155. [163]

    Nguyen, E

    T. Nguyen, E. Bonilla, Fast Allocation of Gaussian Process Experts, in: Proceedings of the 31st Interna- tional Conference on Machine Learning, PMLR, 2014, pp. 145–153

  156. [164]

    T. N. A. Nguyen, A. Bouzerdoum, S. L. Phung, Varia- tional inference for infinite mixtures of sparse Gaussian processes through KL-correction, in: 2016 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 2579–2583

  157. [165]

    Nguyen-Tuong, M

    D. Nguyen-Tuong, M. Seeger, J. Peters, Model Learn- ing with Local Gaussian Process Regression, Ad- vanced Robotics 23 (2009) 2015–2034

  158. [166]

    Nguyen-Tuong, J

    D. Nguyen-Tuong, J. Peters, M. Seeger, Local Gaus- sian Process Regression for Real Time Online Model Learning, in: Advances in Neural Information Pro- cessing Systems, volume 21, Curran Associates, Inc., 2008

  159. [167]

    Nguyen-Tuong, J

    D. Nguyen-Tuong, J. Peters, Incremental online spar- sification for model learning in real-time robot control, Neurocomputing 74 (2011) 1859–1867

  160. [168]

    T. Chen, J. Ren, Bagging for Gaussian process regres- sion, Neurocomputing 72 (2009) 1605–1610

  161. [169]

    S. Das, S. Roy, R. Sambasivan, Fast Gaussian Process Regression for Big Data, Big Data Research 14 (2018) 12–26

  162. [170]

    Schneider, W

    M. Schneider, W. Ertel, Robot Learning by Demon- stration with local Gaussian process regression, in: 2010 IEEE/RSJ International Conference on Intelli- gent Robots and Systems, 2010, pp. 255–260

  163. [171]

    C. D. McKinnon, A. P. Schoellig, Learning multi- modal models for robot dynamics online with a mix- ture of Gaussian process experts, in: 2017 IEEE In- ternational Conference on Robotics and Automation (ICRA), IEEE, Singapore, Singapore, 2017, pp. 322– 328

  164. [172]

    G. E. Hinton, Training Products of Experts by Min- imizing Contrastive Divergence, Neural Computation 14 (2002) 1771–1800

  165. [173]

    Tresp, A Bayesian Committee Machine, Neural Computation 12 (2000) 2719–2741

    V. Tresp, A Bayesian Committee Machine, Neural Computation 12 (2000) 2719–2741

  166. [174]

    Deisenroth, J

    M. Deisenroth, J. W. Ng, Distributed Gaussian Pro- cesses, in: Proceedings of the 32nd International Con- ference on Machine Learning, PMLR, 2015, pp. 1481– 1490

  167. [175]

    Lederer, A

    A. Lederer, A. J. O. Conejo, K. A. Maier, W. Xiao, J. Umlauft, S. Hirche, Gaussian Process-Based Real- Time Learning for Safety Critical Applications, in: Proceedings of the 38th International Conference on Machine Learning, PMLR, 2021, pp. 6055–6064

  168. [176]

    Urtasun, T

    R. Urtasun, T. Darrell, Sparse probabilistic regres- sion for activity-independent human pose inference, in: 2008 IEEE Conference on Computer Vision and Pat- tern Recognition, IEEE, Anchorage, AK, USA, 2008, 44 pp. 1–8

  169. [177]

    Z. Liu, L. Zhou, H. Leung, H. P. Shum, Kinect Posture Reconstruction Based on a Local Mixture of Gaussian Process Models, IEEE Transactions on Visualization and Computer Graphics 22 (2016) 2437–2450

  170. [178]

    Szab´ o, H

    B. Szab´ o, H. v. Zanten, An asymptotic analysis of dis- tributed nonparametric methods, Journal of Machine Learning Research 20 (2019) 1–30

  171. [179]

    Trapp, R

    M. Trapp, R. Peharz, F. Pernkopf, C. E. Rasmussen, Deep Structured Mixtures of Gaussian Processes, in: Proceedings of the Twenty Third International Con- ference on Artificial Intelligence and Statistics, PMLR, 2020, pp. 2251–2261

  172. [180]

    Qui˜ nonero-Candela, C

    J. Qui˜ nonero-Candela, C. E. Rasmussen, A Unifying View of Sparse Approximate Gaussian Process Regres- sion, J. Mach. Learn. Res. 6 (2005) 1939–1959

  173. [181]

    T. D. Bui, J. Yan, R. E. Turner, A unifying framework for Gaussian process pseudo-point approximations us- ing power expectation propagation, The Journal of Machine Learning Research 18 (2017) 3649–3720

  174. [182]

    Bauer, M

    M. Bauer, M. van der Wilk, C. E. Rasmussen, Under- standing Probabilistic Sparse Gaussian Process Ap- proximations, in: Advances in Neural Information Processing Systems, volume 29, Curran Associates, Inc., 2016

  175. [183]

    L´ azaro-Gredilla, A

    M. L´ azaro-Gredilla, A. Figueiras-Vidal, Inter-domain Gaussian Processes for Sparse Inference using Induc- ing Features, in: Advances in Neural Information Pro- cessing Systems, volume 22, Curran Associates, Inc., 2009

  176. [184]

    van der Wilk, C

    M. van der Wilk, C. E. Rasmussen, J. Hensman, Con- volutional Gaussian Processes, in: Advances in Neural Information Processing Systems, volume 30, Curran Associates, Inc., 2017

  177. [185]

    Dutordoir, N

    V. Dutordoir, N. Durrande, J. Hensman, Sparse Gaus- sian Processes with Spherical Harmonic Features, in: Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, pp. 2793–2802

  178. [186]

    Silverman, Some Aspects of the Spline Smoothing Approach to Non-Parametric Regression Curve Fit- ting, Journal of the Royal Statistical Society

    B. Silverman, Some Aspects of the Spline Smoothing Approach to Non-Parametric Regression Curve Fit- ting, Journal of the Royal Statistical Society. Series B (Methodological) 47 (1985) 1–52

  179. [187]

    Wahba, X

    G. Wahba, X. Lin, F. Gao, D. Xiang, R. Klein, B. Klein, The bias-variance tradeoff and the ran- domized GACV, in: Proceedings of the 11th Inter- national Conference on Neural Information Processing Systems, NIPS’98, MIT Press, Cambridge, MA, USA, 1998, pp. 620–626

  180. [188]

    A. J. Smola, B. Sch¨ okopf, Sparse Greedy Matrix Ap- proximation for Machine Learning, in: Proceedings of the Seventeenth International Conference on Ma- chine Learning, ICML ’00, Morgan Kaufmann Pub- lishers Inc., San Francisco, CA, USA, 2000, pp. 911– 918

  181. [189]

    Csat´ o, M

    L. Csat´ o, M. Opper, Sparse On-Line Gaussian Pro- cesses, Neural Computation 14 (2002) 641–668

  182. [190]

    Seeger, C

    M. Seeger, C. Williams, N. Lawrence, Fast Forward Selection to Speed Up Sparse Gaussian Process Re- gression, in: International Workshop on Artificial In- telligence and Statistics, PMLR, 2003, pp. 254–261

  183. [191]

    Billingsley, Probability and Measure, Wiley Serie s in Probability and Statistics, wiley ed., 2012

    P. Billingsley, Probability and Measure, Wiley Serie s in Probability and Statistics, wiley ed., 2012

  184. [192]

    Snelson, Z

    E. Snelson, Z. Ghahramani, Local and global sparse Gaussian process approximations, in: Proceedings of the Eleventh International Conference on Artificial In- telligence and Statistics, PMLR, 2007, pp. 524–531

  185. [193]

    Snelson, Z

    E. Snelson, Z. Ghahramani, Sparse Gaussian Processes using Pseudo-inputs, in: Advances in Neural Informa- tion Processing Systems, volume 18, MIT Press, 2005

  186. [194]

    Schwaighofer, V

    A. Schwaighofer, V. Tresp, Transductive and Inductive Methods for Approximate Gaussian Process Regres- sion, in: Advances in Neural Information Processing Systems, volume 15, MIT Press, 2002

  187. [195]

    Rossi, M

    S. Rossi, M. Heinonen, E. Bonilla, Z. Shen, M. Filip- pone, Sparse Gaussian Processes Revisited: Bayesian Approaches to Inducing-Variable Approximations, in: Proceedings of The 24th International Conference on Artificial Intelligence and Statistics, PMLR, 2021, pp. 1837–1845

  188. [196]

    Scampicchio, S

    A. Scampicchio, S. Chandrasekaran, M. N. Zeilinger, A Markov Chain Monte Carlo approach for Pseudo- Input selection in Sparse Gaussian Processes, IFAC- PapersOnLine 56 (2023) 10515–10520

  189. [197]

    Y. Cao, M. A. Brubaker, D. J. Fleet, A. Hertzmann, Efficient Optimization for Sparse Gaussian Process Regression, IEEE Transactions on Pattern Analysis and Machine Intelligence 37 (2015) 2415–2427

  190. [198]

    Yuan, Qi, A. H. Abdel-Gawad, T. P. Minka, Sparse- posterior Gaussian Processes for general likelihoods, 2012

  191. [199]

    X. Wei, J. Min, J. Chai, Physically valid statistical models for human motion generation, ACM Transac- tions on Graphics 30 (2011) 19:1–19:10

  192. [200]

    Ahmad, D

    T. Ahmad, D. Zhang, C. Huang, Methodological framework for short-and medium-term energy, solar and wind power forecasting with stochastic-based ma- chine learning approach to monetary and energy policy applications, Energy 231 (2021) 120911

  193. [201]

    K. Fan, Y. Wan, B. Jiang, State-of-charge dependent equivalent circuit model identification for batteries us- ing sparse Gaussian process regression, Journal of Pro- cess Control 112 (2022) 1–11

  194. [202]

    Y. Xue, G. Chen, Z. Li, G. Xue, W. Wang, Y. Liu, On- line identification of a ship maneuvering model using a fast noisy input Gaussian process, Ocean Engineering 250 (2022) 110704

  195. [203]

    T. N. Hoang, Q. M. Hoang, B. K. H. Low, A Unifying Framework of Anytime Sparse Gaussian Process Re- gression Models with Stochastic Variational Inference for Big Data, in: Proceedings of the 32nd Interna- tional Conference on Machine Learning, PMLR, 2015, pp. 569–578

  196. [204]

    Titsias, Variational Model Selection for Sparse Gaussian Process Regression, Technical Report, Uni- versity of Manchester, 2008

    M. Titsias, Variational Model Selection for Sparse Gaussian Process Regression, Technical Report, Uni- versity of Manchester, 2008

  197. [205]

    M. Titsias, Variational Learning of Inducing Variabl es in Sparse Gaussian Processes, in: Proceedings of the Twelth International Conference on Artificial Intelli- gence and Statistics, PMLR, 2009, pp. 567–574

  198. [206]

    D. M. Blei, A. Kucukelbir, J. D. McAuliffe, Variational Inference: A Review for Statisticians, Journal of the American Statistical Association 112 (2017) 859–877

  199. [207]

    Zhang, J

    C. Zhang, J. B¨ utepage, H. Kjellstr¨ om, S. Mandt, Ad- vances in Variational Inference, IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (2019) 2008–2026

  200. [208]

    Wilson, H

    A. Wilson, H. Nickisch, Kernel Interpolation for Scal - able Structured Gaussian Processes (KISS-GP), in: Proceedings of the 32nd International Conference on Machine Learning, PMLR, 2015, pp. 1775–1784. 45

  201. [209]

    A. G. Wilson, C. Dann, H. Nickisch, Thoughts on Mas- sively Scalable Gaussian Processes, 2015

  202. [210]

    Gardner, G

    J. Gardner, G. Pleiss, R. Wu, K. Weinberger, A. Wil- son, Product Kernel Interpolation for Scalable Gaus- sian Processes, in: Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics, PMLR, 2018, pp. 1407–1416

  203. [211]

    Evans, P

    T. Evans, P. Nair, Scalable Gaussian Processes with Grid-Structured Eigenfunctions (GP-GRIEF), in: Proceedings of the 35th International Conference on Machine Learning, PMLR, 2018, pp. 1417–1426

  204. [212]

    Izmailov, A

    P. Izmailov, A. Novikov, D. Kropotov, Scalable Gaus- sian Processes with Billions of Inducing Inputs via Tensor Train Decomposition, in: Proceedings of the Twenty-First International Conference on Artificial In- telligence and Statistics, PMLR, 2018, pp. 726–735

  205. [213]

    Wesel, K

    F. Wesel, K. Batselier, Tensor-based Kernel Machines with Structured Inducing Points for Large and High- Dimensional Data, in: Proceedings of The 26th In- ternational Conference on Artificial Intelligence and Statistics, PMLR, 2023, pp. 8308–8320

  206. [214]

    G.-L. Tran, D. Milios, P. Michiardi, M. Filippone, Sparse within Sparse Gaussian Processes using Neigh- bor Information, in: Proceedings of the 38th Interna- tional Conference on Machine Learning, PMLR, 2021, pp. 10369–10378

  207. [215]

    Lalchand, W

    V. Lalchand, W. Bruinsma, D. Burt, C. E. Rasmussen, Sparse Gaussian Process Hyperparameters: Optimize or Integrate?, Advances in Neural Information Pro- cessing Systems 35 (2022) 16612–16623

  208. [216]

    Hensman, A

    J. Hensman, A. G. Matthews, M. Filippone, Z. Ghahramani, MCMC for Variationally Sparse Gaussian Processes, in: Advances in Neural Infor- mation Processing Systems, volume 28, Curran Asso- ciates, Inc., 2015

  209. [217]

    I. Paun, D. Husmeier, C. J. Torney, Stochastic vari- ational inference for scalable non-stationary Gaussian process regression, Statistics and Computing 33 (2023) 44

  210. [218]

    A. G. d. G. Matthews, J. Hensman, R. Turner, Z. Ghahramani, On Sparse Variational Methods and the Kullback-Leibler Divergence between Stochastic Processes, in: Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, PMLR, 2016, pp. 231–239

  211. [219]

    Cheng, B

    C.-A. Cheng, B. Boots, Incremental variational spars e Gaussian process regression, in: Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, Curran Associates Inc., Red Hook, NY, USA, 2016, pp. 4410–4418

  212. [220]

    Jankowiak, G

    M. Jankowiak, G. Pleiss, J. Gardner, Parametric Gaussian Process Regressors, in: Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, pp. 4702–4712

  213. [221]

    Hensman, N

    J. Hensman, N. Fusi, N. D. Lawrence, Gaussian Pro- cesses for Big Data, 2013

  214. [222]

    M. D. Hoffman, D. M. Blei, C. Wang, J. Paisley, Stochastic Variational Inference, Journal of Machine Learning Research 14 (2013) 1303–1347

  215. [223]

    Cheng, B

    C.-A. Cheng, B. Boots, Variational Inference for Gaus- sian Process Models with Linear Complexity, in: Ad- vances in Neural Information Processing Systems, vol- ume 30, Curran Associates, Inc., 2017

  216. [224]

    Salimbeni, C.-A

    H. Salimbeni, C.-A. Cheng, B. Boots, M. Deisen- roth, Orthogonally Decoupled Variational Gaussian Processes, in: Advances in Neural Information Pro- cessing Systems, volume 31, Curran Associates, Inc., 2018

  217. [225]

    J. Shi, M. Titsias, A. Mnih, Sparse Orthogonal Vari- ational Inference for Gaussian Processes, in: Proceed- ings of the Twenty Third International Conference on Artificial Intelligence and Statistics, PMLR, 2020, pp. 1932–1942

  218. [226]

    Rogers, P

    T. Rogers, P. Gardner, N. Dervilis, K. Worden, A. E. Maguire, E. Papatheou, E. J. Cross, Probabilistic modelling of wind turbine power curves with appli- cation of heteroscedastic Gaussian Process regression, Renewable Energy 148 (2020) 1124–1136

  219. [227]

    McIntire, D

    M. McIntire, D. Ratner, S. Ermon, Sparse Gaussian processes for Bayesian optimization, in: Proceedings of the Thirty-Second Conference on Uncertainty in Ar- tificial Intelligence, UAI’16, AUAI Press, Arlington, Virginia, USA, 2016, pp. 517–526

  220. [228]

    Ament, M

    S. Ament, M. Amsler, D. R. Sutherland, M.-C. Chang, D. Guevarra, A. B. Connolly, J. M. Gregoire, M. O. Thompson, C. P. Gomes, R. B. van Dover, Au- tonomous materials synthesis via hierarchical active learning of nonequilibrium phase diagrams, Science Advances 7 (2021) eabg4930

  221. [229]

    Salimbeni, M

    H. Salimbeni, M. Deisenroth, Doubly Stochastic Vari- ational Inference for Deep Gaussian Processes, in: Ad- vances in Neural Information Processing Systems, vol- ume 30, Curran Associates, Inc., 2017

  222. [230]

    D. Burt, C. E. Rasmussen, M. V. D. Wilk, Rates of Convergence for Sparse Variational Gaussian Process Regression, in: Proceedings of the 36th International Conference on Machine Learning, PMLR, 2019, pp. 862–871

  223. [231]

    D. R. Burt, C. E. Rasmussen, M. Van Der Wilk, Con- vergence of sparse variational inference in Gaussian processes regression, The Journal of Machine Learn- ing Research 21 (2020) 131:5120–131:5182

  224. [232]

    Nieman, B

    D. Nieman, B. Szabo, H. v. Zanten, Contraction rates for sparse variational approximations in Gaussian pro- cess regression, Journal of Machine Learning Research 23 (2022) 1–26

  225. [233]

    J. H. Huggins, T. Campbell, M. Kasprzak, T. Broder- ick, Scalable Gaussian Process Inference with Finite- data Mean and Variance Guarantees, in: Proceedings of the Twenty-Second International Conference on Ar- tificial Intelligence and Statistics, PMLR, 2019, pp. 796–805

  226. [234]

    L´ azaro-Gredilla, J

    M. L´ azaro-Gredilla, J. Qui˜ nnero-Candela, C. E. Ras- mussen, An&#237, b. R. Figueiras-Vidal, Sparse Spectrum Gaussian Process Regression, Journal of Machine Learning Research 11 (2010) 1865–1881

  227. [235]

    Rahimi, B

    A. Rahimi, B. Recht, Random Features for Large- Scale Kernel Machines, in: Advances in Neural Infor- mation Processing Systems, volume 20, Curran Asso- ciates, Inc., 2007

  228. [236]

    Rahimi, B

    A. Rahimi, B. Recht, Uniform approximation of func- tions with random bases, in: 2008 46th Annual Aller- ton Conference on Communication, Control, and Com- puting, 2008, pp. 555–561

  229. [237]

    Q. Le, T. Sarlos, A. Smola, Fastfood - Computing Hilbert Space Expansions in loglinear time, in: Pro- ceedings of the 30th International Conference on Ma- chine Learning, PMLR, 2013, pp. 244–252

  230. [238]

    Z. Yang, A. Wilson, A. Smola, L. Song, A la Carte – Learning Fast Kernels, in: Proceedings of the Eigh- 46 teenth International Conference on Artificial Intelli- gence and Statistics, PMLR, 2015, pp. 1098–1106

  231. [239]

    Y. Gal, R. Turner, Improving the Gaussian Process Sparse Spectrum Approximation by Representing Un- certainty in Frequency Inputs, in: Proceedings of the 32nd International Conference on Machine Learning, PMLR, 2015, pp. 655–664

  232. [240]

    L. S. L. Tan, V. M. H. Ong, D. J. Nott, A. Jasra, Vari- ational inference for sparse spectrum Gaussian process regression, Statistics and Computing 26 (2016) 1243– 1261

  233. [241]

    Hensman, N

    J. Hensman, N. Durrande, A. Solin, Variational Fourier Features for Gaussian Processes, Journal of Machine Learning Research 18 (2018) 1–52

  234. [242]

    Sriperumbudur, Z

    B. Sriperumbudur, Z. Szabo, Optimal Rates for Ran- dom Fourier Features, in: Advances in Neural Infor- mation Processing Systems, volume 28, Curran Asso- ciates, Inc., 2015

  235. [243]

    A. Rudi, L. Rosasco, Generalization Properties of Learning with Random Features, in: Advances in Neu- ral Information Processing Systems, volume 30, Cur- ran Associates, Inc., 2017

  236. [244]

    Scampicchio, E

    A. Scampicchio, E. Arcari, M. N. Zeilinger, Error Analysis of Regularized Trigonometric Linear Regres- sion With Unbounded Sampling: A Statistical Learn- ing Viewpoint, IEEE Control Systems Letters 7 (2023) 3066–3071

  237. [245]

    Gijsberts, G

    A. Gijsberts, G. Metta, Real-time model learning using Incremental Sparse Spectrum Gaussian Process Regression, Neural Networks 41 (2013) 59–69

  238. [246]

    B. C. Levy, Karhunen Loeve Expansion of Gaussian Processes, in: B. C. Levy (Ed.), Principles of Sig- nal Detection and Parameter Estimation, Springer US, Boston, MA, 2008, pp. 1–47

  239. [247]

    G. Picci, On the optimality of the Karhunen-Lo` eve expansion, in: Ricordo di Antonio Lepschy, Istituto Veneto di Scienze, Lettere ed Arti, Adunanza Ac- cademica del 26 Novembre 2005, Istituto Veneto di Scienze, Lettere ed Arti, Venezia, 2006

  240. [248]

    Mercer, A

    J. Mercer, A. R. Forsyth, XVI. Functions of positive and negative type, and their connection the theory of integral equations, Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character 209 (1997) 415–446

  241. [249]

    H. Zhu, C. K. I. Williams, R. Rohwer, M. Morciniec, Gaussian regression and optimal finite dimensional lin- ear models, 1997

  242. [250]

    Ferrari-Trecate, C

    G. Ferrari-Trecate, C. Williams, M. Opper, Finite- Dimensional Approximation of Gaussian Processes, in: Advances in Neural Information Processing Systems, volume 11, MIT Press, 1998

  243. [251]

    Fasshauer, Positive definite kernels: past, presen t and future, Dolomites Research Notes on Approxima- tion (2011)

    G. Fasshauer, Positive definite kernels: past, presen t and future, Dolomites Research Notes on Approxima- tion (2011)

  244. [252]

    E. Wong, B. Hajek, Stochastic Processes in Engineer- ing Systems, Springer Texts in Electrical Engineering, Springer, New York, NY, 1985

  245. [253]

    Y. M. Marzouk, H. N. Najm, Dimensionality reduc- tion and polynomial chaos acceleration of Bayesian in- ference in inverse problems, Journal of Computational Physics 228 (2009) 1862–1902

  246. [254]

    H. Peng, Y. Qi, EigenGP: Gaussian process mod- els with adaptive eigenfunctions, in: Proceedings of the 24th International Conference on Artificial Intel- ligence, IJCAI’15, AAAI Press, Buenos Aires, Ar- gentina, 2015, pp. 3763–3769

  247. [255]

    Solin, S

    A. Solin, S. S¨ arkk¨ a, Hilbert space methods for reduced-rank Gaussian process regression, Statistics and Computing 30 (2020) 419–446

  248. [256]

    Riutort-Mayol, P.-C

    G. Riutort-Mayol, P.-C. B¨ urkner, M. R. Andersen, A. Solin, A. Vehtari, Practical Hilbert space approx- imate Bayesian Gaussian processes for probabilistic programming, Statistics and Computing 33 (2022) 17

  249. [257]

    Solin, M

    A. Solin, M. Kok, N. Wahlstr¨ om, T. B. Sch¨ on, S. S¨ arkk¨ a, Modeling and Interpolation of the Am- bient Magnetic Field by Gaussian Processes, IEEE Transactions on Robotics 34 (2018) 1112–1127

  250. [258]

    Williams, M

    C. Williams, M. Seeger, Using the Nystr¨ om Method to Speed Up Kernel Machines, in: Advances in Neu- ral Information Processing Systems, volume 13, MIT Press, 2000

  251. [259]

    S. Sun, J. Zhao, J. Zhu, A review of Nystr¨ om methods for large-scale machine learning, Information Fusion 26 (2015) 36–48

  252. [260]

    Bartels, W

    S. Bartels, W. Boomsma, J. Frellsen, D. Garreau, Kernel-Matrix Determinant Estimates from stopped Cholesky Decomposition, Journal of Machine Learn- ing Research 24 (2023) 1–57

  253. [261]

    Bartels, K

    S. Bartels, K. Stensbo-Smidt, P. Moreno-Munoz, W. Boomsma, J. Frellsen, S. Hauberg, Adaptive Cholesky Gaussian Processes, in: Proceedings of The 26th International Conference on Artificial Intelligence and Statistics, PMLR, 2023, pp. 408–452

  254. [262]

    S. Fine, K. Scheinberg, Efficient SVM training using low-rank kernel representations, Journal of Machine Learning Research 2 (2001) 243–264

  255. [263]

    F. Bach, M. Jordan, Kernel Independent Compo- nent Analysis, Journal of Machine Learning Research (2002)

  256. [264]

    C. K. I. Williams, C. E. Rasmussen, A. Schwaighofer, V. Tresp, Observations on the Nystr¨ om Method for Gaussian Processes, Technical Report, University of Edinburgh, 2002

  257. [265]

    F. Bach, M. Jordan, Predictive low-rank decomposi- tion for kernel methods, in: Proceedings of the 22nd international conference on Machine learning, ICML ’05, Association for Computing Machinery, New York, NY, USA, 2005, pp. 33–40

  258. [266]

    Harbrecht, M

    H. Harbrecht, M. Peters, R. Schneider, On the low- rank approximation by the pivoted Cholesky decom- position, Applied Numerical Mathematics 62 (2012) 428–440

  259. [267]

    Sch¨ afer, T

    F. Sch¨ afer, T. J. Sullivan, H. Owhadi, Compression, Inversion, and Approximate PCA of Dense Kernel Ma- trices at Near-Linear Computational Complexity, Mul- tiscale Modeling & Simulation 19 (2021) 688–730

  260. [268]

    Drineas, M

    P. Drineas, M. W. Mahoney, On the Nystrom Method for Approximating a Gram Matrix for Im- proved Kernel-Based Learning, Journal of Machine Learning Research 6 (2005) 2153–2175

  261. [269]

    Kumar, M

    S. Kumar, M. Mohri, A. Talwalkar, Sampling Methods for the Nystr¨ om Method, Journal of Machine Learning Research 13 (2012) 981–1006

  262. [270]

    Zhang, I

    K. Zhang, I. W. Tsang, J. T. Kwok, Improved Nystr¨ om low-rank approximation and error analysis, in: Proceedings of the 25th international conference on Machine learning, ICML ’08, Association for Comput- ing Machinery, New York, NY, USA, 2008, pp. 1232– 1239. 47

  263. [271]

    Drineas, M

    P. Drineas, M. Magdon-Ismail, M. W. Mahoney, D. P. Woodruff, Fast Approximation of Matrix Coherence and Statistical Leverage, Journal of Machine Learning Research 13 (2012)

  264. [272]

    Alaoui, M

    A. Alaoui, M. Mahoney, Fast Randomized Kernel Ridge Regression with Statistical Guarantees, in: Ad- vances in Neural Information Processing Systems, vol- ume 28, Curran Associates, Inc., 2015

  265. [273]

    Musco, C

    C. Musco, C. Musco, Recursive Sampling for the Nys- trom Method, in: Advances in Neural Information Processing Systems, volume 30, Curran Associates, Inc., 2017

  266. [274]

    Halko, P

    N. Halko, P. G. Martinsson, J. Tropp, Finding Struc- ture with Randomness: Probabilistic Algorithms for Constructing Approximate Matrix Decompositions, SIAM Review 53 (2011) 217–288

  267. [275]

    Martinsson, J

    P.-G. Martinsson, J. A. Tropp, Randomized numeri- cal linear algebra: Foundations and algorithms, Acta Numerica 29 (2020) 403–572

  268. [276]

    Banerjee, D

    A. Banerjee, D. B. Dunson, S. T. Tokdar, Effi- cient Gaussian process regression for large datasets, Biometrika 100 (2013) 75–89

  269. [277]

    Gittens, M

    A. Gittens, M. Mahoney, Revisiting the Nystrom method for improved large-scale machine learning, in: Proceedings of the 30th International Conference on Machine Learning, PMLR, 2013, pp. 567–575

  270. [278]

    R. Jin, T. Yang, M. Mahdavi, Y.-F. Li, Z.-H. Zhou, Improved Bound for the Nystrom’s Method and its Application to Kernel Classification, IEEE Transac- tions on Information Theory 59 (2013) 6939–6949

  271. [279]

    A. Rudi, R. Camoriano, L. Rosasco, Less is More: Nystr¨ om Computational Regularization, in: Advances in Neural Information Processing Systems, volume 28, Curran Associates, Inc., 2015

  272. [280]

    Cortes, M

    C. Cortes, M. Mohri, A. Talwalkar, On the Impact of Kernel Approximation on Learning Accuracy, in: Pro- ceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, JMLR Workshop and Conference Proceedings, 2010, pp. 113–120

  273. [281]

    Bach, Sharp analysis of low-rank kernel matrix approximations, in: Proceedings of the 26th An- nual Conference on Learning Theory, PMLR, 2013, pp

    F. Bach, Sharp analysis of low-rank kernel matrix approximations, in: Proceedings of the 26th An- nual Conference on Learning Theory, PMLR, 2013, pp. 185–209

  274. [282]

    Cutajar, M

    K. Cutajar, M. Osborne, J. Cunningham, M. Filip- pone, Preconditioning Kernel Matrices, in: Proceed- ings of The 33rd International Conference on Machine Learning, PMLR, 2016, pp. 2529–2538

  275. [283]

    A. Rudi, L. Carratino, L. Rosasco, FALKON: An Op- timal Large Scale Kernel Method, in: Advances in Neural Information Processing Systems, volume 30, Curran Associates, Inc., 2017

  276. [284]

    Gardner, G

    J. Gardner, G. Pleiss, K. Q. Weinberger, D. Bindel, A. G. Wilson, GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration, in: Advances in Neural Information Processing Sys- tems, volume 31, Curran Associates, Inc., 2018, pp. 7576–7586

  277. [285]

    Wenger, G

    J. Wenger, G. Pleiss, P. Hennig, J. Cunningham, J. Gardner, Preconditioning for Scalable Gaussian Process Hyperparameter Optimization, in: Proceed- ings of the 39th International Conference on Machine Learning, PMLR, 2022, pp. 23751–23780

  278. [286]

    Hestenes, E

    M. Hestenes, E. Stiefel, Methods of conjugate gradi- ents for solving linear systems, Journal of Research of the National Bureau of Standards 49 (1952) 409

  279. [287]

    W. E. Leithead, Y. Zhang, O(N 2)-Operation Ap- proximation of Covariance Matrix Inverse in Gaussian Process Regression Based on Quasi-Newton BFGS Method, Communications in Statistics - Simulation and Computation 36 (2007) 367–380

  280. [288]

    Skilling, Bayesian numerical analysis, Physics an d Probability (1993) 207–221

    J. Skilling, Bayesian numerical analysis, Physics an d Probability (1993) 207–221

  281. [289]

    Gibbs, D

    M. Gibbs, D. Mackay, Efficient implementation of gaussian processes, Technical Report, Cavendish Lab- oratory, Cambridge, UK, 1997

  282. [290]

    Gibbs, Bayesian Gaussian Processes for Regression and Classification, PhD Thesis, University of Cam- bridge, 1997

    M. Gibbs, Bayesian Gaussian Processes for Regression and Classification, PhD Thesis, University of Cam- bridge, 1997

  283. [291]

    Girard, Un algorithme simple et rapide pour la val- idation crois´ ee g´ en´ eralis´ ee sur des probl` emes de grande taille, Technical Report, 1987

    D. Girard, Un algorithme simple et rapide pour la val- idation crois´ ee g´ en´ eralis´ ee sur des probl` emes de grande taille, Technical Report, 1987

  284. [292]

    M. Hutchinson, A Stochastic Estimator of the Trace of the Influence Matrix for Laplacian Smoothing Splines, Communications in Statistics - Simulation and Com- putation 18 (1989) 1059–1076

  285. [293]

    Ubaru, J

    S. Ubaru, J. Chen, Y. Saad, Fast Estimation of \ tr(f(A))\ via Stochastic Lanczos Quadrature, SIAM Journal on Matrix Analysis and Applications 38 (2017) 1075–1099

  286. [294]

    K. Dong, D. Eriksson, H. Nickisch, D. Bindel, A. G. Wilson, Scalable log determinants for Gaussian pro- cess kernel learning, Advances in Neural Information Processing Systems 30 (2017)

  287. [295]

    G. H. Golub, C. F. Van Loan, Matrix Computa- tions, Johns Hopkins studies in the mathematical sci- ences, fourth edition ed., The Johns Hopkins Univer- sity Press, Baltimore, 2013

  288. [296]

    Zhang, W

    Y. Zhang, W. E. Leithead, Approximate implemen- tation of the logarithm of the matrix determinant in Gaussian process regression, Journal of Statistical Computation and Simulation 77 (2007) 329–348

  289. [297]

    Davies, Effective implementation of Gaussian pro- cess regression for machine learning, PhD Thesis, Uni- versity of Cambridge, 2015

    A. Davies, Effective implementation of Gaussian pro- cess regression for machine learning, PhD Thesis, Uni- versity of Cambridge, 2015

  290. [298]

    Artemev, D

    A. Artemev, D. R. Burt, M. v. d. Wilk, Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients, in: Proceedings of the 38th International Conference on Machine Learning, PMLR, 2021, pp. 362–372

  291. [299]

    Filippone, R

    M. Filippone, R. Engler, Enabling scalable stochasti c gradient-based inference for Gaussian processes by em- ploying the Unbiased LInear System SolvEr (ULISSE), in: Proceedings of the 32nd International Conference on Machine Learning, PMLR, 2015, pp. 1015–1024

  292. [300]

    Potapczynski, L

    A. Potapczynski, L. Wu, D. Biderman, G. Pleiss, J. P. Cunningham, Bias-Free Scalable Gaussian Processes via Randomized Truncations, in: Proceedings of the 38th International Conference on Machine Learning, PMLR, 2021, pp. 8609–8619

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.