Pith. sign in

REVIEW 3 major objections 4 minor 121 references

Scalable Gaussian Processes: Advances in Iterative Methods and Pathwise Conditioning

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A dissertation argues that rewriting Gaussian process computations as linear systems and pathwise sampling scales them to millions of data points with order-of-magnitude speed-ups.

desk verdict A well-written thesis that usefully unifies the author's published iterative-GP work, but the 'exact GP' framing is overstated because it rests on a fixed 2000-random-feature prior approximation. read the letter →

arxiv 2507.06839 v1 pith:U5YS7LNV submitted 2025-07-09 cs.LG stat.ML

classification cs.LGstat.ML
keywords GaussianprocessesscalableinferenceiterativelinearsystemsolverspathwiseconditioningstochasticdualdescentmarginallikelihoodoptimisationKroneckerstructurerandomFourierfeatures
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This dissertation tries to establish that Gaussian process inference can be made scalable to large modern datasets without moving to an approximate model. Its strategy is to write every expensive GP computation, posterior means, posterior samples, marginal-likelihood gradients, as the solution of a positive-definite linear system and to solve those systems with iterative solvers whose main operation is matrix-vector multiplication. Pathwise conditioning then turns the same solver output into posterior samples that can be evaluated at arbitrary locations, and the chapters build stochastic gradient and dual-descent solvers, warm-started and pathwise gradient estimators, and latent Kronecker structure on top of this core. If the claims are right, a practitioner gets GP inference in the exact model up to a numerical tolerance, with linear or sublinear scaling in time and memory, demonstrated on regression, Bayesian optimisation, and molecular binding affinity tasks, with speed-ups up to $72\times$ and datasets up to five million examples.

What carries the argument

The workhorse identity is $v=(K_{XX}+\sigma^2I)^{-1}b \iff v=\arg\min_u \tfrac12 u^\top (K_{XX}+\sigma^2I)u - u^\top b$, which turns GP computations into convex quadratic optimisation; for any iterative solver the gradient is the residual $Au-b$, so matrix-vector multiplication with the kernel matrix is the dominant cost. Pathwise conditioning, $f(\cdot)|y = f(\cdot)+K(\cdot)X(K_{XX}+\sigma^2I)^{-1}(y-(f_X+\varepsilon))$, converts a single linear solve into a posterior sample and decouples the expensive data-dependent solve from the number of prediction locations. Random Fourier features provide approximate prior samples $f_X$; the dual objective $\tfrac12\|\alpha\|^2_{K_{XX}+\sigma^2I}-\alpha^\top b$ with random-coordinate gradients gives stochastic dual descent better curvature than the primal objective; and latent Kronecker structure expresses the observed covariance as the projection of a Kronecker product so that fast matrix multiplication survives missing values and irregular inputs. These pieces are the machinery that carries every chapter's method.

What would settle it

Run the full pipeline (pathwise sampling, stochastic dual descent, warm-started pathwise marginal likelihood gradients) on a regression problem small enough for exact Cholesky inference (say $n=10^4$) using both the standard 2000 random Fourier features and exact prior samples or a much larger feature budget, while solving all linear systems to machine precision; if predictive log-likelihood, posterior sample moments, or the converged marginal likelihood differ systematically between the two, the random-feature approximation, not solver non-convergence, is the limiting assumption.

Watch

Extended reading notes

Core claim

The central claim is that the apparent cubic barrier of Gaussian processes is not inherent: the same kernel matrix inverse appears in prediction, sampling, and hyperparameter learning, and each appearance can be re-expressed as the solution of $(K_{XX}+\sigma^2I)v=b$, obtained by iterative optimisation of the equivalent convex quadratic. The dissertation argues that stochastic optimisation is a legitimate solver for these systems: stochastic gradient descent and its stochastic dual descent variant have favourable geometry and implicit-bias properties, so they produce accurate predictions even before full convergence. It further shows that a pathwise gradient estimator for the marginal likelihood both accelerates solver convergence and supplies posterior samples at no extra cost, that warm starting solvers across hyperparameter steps yields large speed-ups with negligible bias, and that projecting a latent Kronecker product keeps fast matrix multiplication available for non-grid data. The author would state the contribution as: iterative methods plus pathwise conditioning make Gaussian process inference practical at scales where exact Cholesky-based inference is impossible, while preserving the exact GP model up to a user-chosen numerical tolerance.

Load-bearing premise

The load-bearing premise is that the approximate prior samples used for pathwise conditioning and marginal-likelihood gradients, drawn from a random-feature model with a fixed number of features rather than from the exact GP, are faithful enough; if they are not, every posterior sample and hyperparameter gradient inherits the bias.

Editorial extensions

If this is right

  • Posterior means and posterior samples can be computed with asymptotically linear (or with inducing points, sublinear) time and memory in dataset size, replacing cubic Cholesky-based inference.
  • Iterative and pathwise inference targets the exact Gaussian process model up to a numerical tolerance rather than an approximate sparse or variational model, preserving calibrated uncertainty for decision-making.
  • Because pathwise conditioning requires one linear solve per posterior sample independent of evaluation locations, acquisition-function optimisation in Bayesian optimisation can reuse samples across many candidate points.
  • Hyperparameter learning can be accelerated by a pathwise marginal-likelihood gradient estimator and warm-started solvers, yielding up to $72\times$ speed-ups when solving to tolerance and up to $7\times$ lower residual norms on a fixed budget.
  • Latent Kronecker structure extends scalable GP inference to non-grid data, demonstrated on real-world datasets with up to five million examples in robotics, automated machine learning, and climate modelling.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: because the solver and sampler are decoupled, replacing random Fourier features with a higher-fidelity approximate prior sampler should remove the main approximation bias without changing the iterative or pathwise machinery; this is directly testable on the dissertation's own benchmarks.
  • Beyond the paper: the observation that low relative residual norms do not reliably predict predictive performance suggests that convergence criteria for GP linear systems should be tied to downstream quantities such as predictive log-likelihood or acquisition value rather than residual alone.
  • Beyond the paper: warm-started, pathwise marginal-likelihood gradients and latent Kronecker structure are complementary, so the large-scale hyperparameter learning of Chapter 5 should extend directly to the five-million-example non-grid setting of Chapter 6.
  • Beyond the paper: the dual-objective insight applies to kernel ridge regression generally, which may shift large-scale kernel methods toward first-order stochastic solvers on ill-conditioned problems where conjugate gradients struggle.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The dissertation develops scalable Gaussian process inference by combining iterative linear-system solvers with pathwise conditioning. Chapter 3 introduces stochastic gradient descent for posterior means and samples, including an inducing-point extension and a spectral analysis of its implicit bias. Chapter 4 proposes stochastic dual descent, based on a dual objective with more favorable conditioning, random-coordinate gradient estimators, momentum, and geometric averaging. Chapter 5 contributes a pathwise estimator of marginal-likelihood gradients, warm starting of inner-loop solvers, and an analysis of early stopping, reporting speed-ups of up to 72×. Chapter 6 introduces latent Kronecker structure for kernel matrices and demonstrates scalability to datasets with up to five million examples. The main theoretical derivations in Chapters 3–5 are internally consistent, while the empirical claims are scoped to large-scale or ill-conditioned regression, Bayesian optimization, and molecular binding-affinity prediction.

Significance. If the technical caveats discussed below are resolved, this is a valuable consolidation and extension of iterative methods for Gaussian processes. The dual-objective equivalence in Proposition 4.1, the initial-distance comparison in Eqs. (5.12)–(5.13), the variance bound in Proposition 5.1, and the spectral characterization of Chapter 3 are clean, well-explained contributions. The empirical program is broad and honestly compared against conjugate gradients, SVGP, and graph neural network baselines. The central limitation is that the claim of inference 'in the exact Gaussian process model' is weakened by the fixed random-feature approximation used for prior samples; this is a model approximation rather than a solver tolerance, and it is not subjected to sensitivity analysis.

major comments (3)
  1. [Section 2.2.4 / 3.2.2 / 5.2.4] The claim that iterative methods perform 'approximate inference in the exact Gaussian process model up to a specified numerical tolerance' conflates solver tolerance with model approximation. In Eq. (3.4) and Eqs. (5.7)–(5.8), every pathwise posterior sample and every probe vector is constructed from a prior sample f_X = Φ_X w with Φ_X Φ_X^T ≈ K_XX using m = 2000 random Fourier features. For finite m, the law of the posterior sample is not the exact GP posterior; this is an approximation of the model itself, not merely of the linear solve. Figure 5.5 confirms that the pathwise marginal-likelihood trajectory deviates from exact optimization and attributes the deviation to random features. Since m = 2000 is fixed for all datasets and kernels with no sensitivity analysis, the 'exact GP' guarantee is unsupported. Please provide either (i) a quantitative bound on ∥K - ΦΦ^T∥ or on the induced posterior perturbation in terms of m and the kernel spectrum, (ii) an m-ablation for representative datasets, including a slow-spectral-decay kernel and higher-dimensional inputs, showing that posterior samples and hyperparameter gradients are insensitive, or (iii) a revised claim that the method performs inference in the random-feature GP model rather than the exact GP model.
  2. [Section 3.2.3] The inducing-point sampling objective replaces f_X^[Z] ∼ N(0, K_XZ K_ZZ^{-1} K_ZX) with f_X ∼ N(0, K_XX), immediately after Eq. (3.24). This changes the sampling target, and the only justification is the unquantified statement that the approximation error is small 'when m is large and the inducing points are close enough to the data.' The sublinear-cost inducing-point results in Figure 3.2 rely on this replacement, but no bound or controlled experiment isolates its effect. Please provide a quantitative error analysis, or a comparison for moderate n between sampling with exact f_X^[Z] and with f_X, to show the replacement is benign in the regimes used.
  3. [Section 5.3.2] Proposition 5.2, which is used to justify that warm starting introduces only negligible bias, is stated informally as 'Under reasonable assumptions' and its formal proof is deferred to Lin et al. (2024c, Appendix A), which is not included in the dissertation. Warm starting fixes the probe vectors across outer-loop steps, so gradient estimates become coupled and the objective being optimized is no longer exactly the marginal likelihood; this is precisely the step that supports the 72× speed-up claim in Table 5.1. The dissertation should either state the formal theorem with explicit assumptions and a proof, or provide a stronger empirical verification, such as multiple seeds with different fixed probe vectors and a comparison of the final marginal likelihood values against exact optimization. The trajectory histograms in Figures 5.5 and 5.8 are suggestive but do not by themselves establish convergence to the exact optimum.
minor comments (4)
  1. [Algorithm 4.1] In the initialization line, 'α0 = 0' appears twice; presumably the third assignment should be to the geometric average \bar{α}_0. Please correct this.
  2. [Abstract / Section 1.1] The abstract claims 'state-of-the-art performance' without qualification, while Chapter 3 itself states the result holds 'on sufficiently large-scale or ill-conditioned regression tasks.' The abstract should carry the same qualification.
  3. [Section 5.2.3] The variance comparison in Proposition 5.1 assumes Gaussian probe vectors for both the standard and pathwise estimators. Since Rademacher or other probe distributions are common in Hutchinson estimators, the text should state explicitly that the comparison is restricted to Gaussian probes.
  4. [Table 3.1] In the SVGP negative log-likelihood row for ELEVATORS, the entry '0.43, ± 0.00' contains a stray comma between the value and the standard error.

Circularity Check

1 steps flagged · score 1.0 of 10

No material circularity: the central derivations are self-contained and the empirical claims rest on external benchmarks; only one minor, non-load-bearing self-citation (proof deferred to the authors' own prior paper) is present.

  1. self citation load bearing [Section 5.3.2 (Proposition 5.2, warm-start bias discussion)]
    "Fortunately, one can show that the marginal likelihood at the optimum implied by these gradients will converge in probability to the marginal likelihood of the true optimum. ... See Lin et al. (2024c, Appendix A) for a formal proof and further details."

    The theoretical guarantee that warm-started gradient trajectories are asymptotically unbiased is deferred to the authors' own prior NeurIPS paper (Lin, Padhy, Mlodozeniec, Antorán, and Hernández-Lobato, 2024c) rather than proved in the dissertation, so a contribution of Chapter 5 ('negligible bias' under warm starting) is partly justified by self-citation. This is not load-bearing to the chapter's central claims, however: the same claim is supported by the chapter's own empirical evidence (identical test log-likelihoods with and without warm start in Table 5.1), the cited proof is peer-reviewed with stated assumptions, and the main speed-up results (up to 72x) are measured, not derived from this proposition.

full rationale

I walked the claimed derivation chain equation by equation. (i) The SGD sampling objectives (3.5) and (3.6) are proved to share the same minimizer by an explicit gradient identity (3.7)-(3.12), not by fiat. (ii) The inducing-point pathwise posterior in Section 3.2.3 is checked against the Titsias posterior mean (2.49) and covariance (2.50) using the Woodbury identity, so Eqs. (3.14)-(3.21) are genuine verifications rather than definitions. (iii) The dual objective's shared minimizer and strong-duality relation are proved in Proposition 4.1, and the KXX-norm error bound in Proposition 4.2 is derived from the reproducing property. (iv) The pathwise gradient estimator's variance bound (Proposition 5.1) is a theorem comparing 2tr(A^2) with tr(A^2)+tr(AA^T); the initial-distance comparison (5.12)-(5.13) is an exact computation (tr(H^{-1}) versus n) with no fitted quantity. (v) The latent-Kronecker break-even point in Chapter 6 is an asymptotic complexity derivation. None of these steps equates a prediction to its fitted input, and none imports an 'uniqueness' or 'ansatz' from the authors' prior work: the random-feature and dual-objective ideas are explicitly attributed to Rahimi and Recht, Sutherland and Schneider, Wilson et al., Saunders et al., Schölkopf and Smola, and Shalev-Shwartz and Zhang. Empirically, the state-of-the-art claims are benchmarked against external datasets (UCI, DOCKSTRING) and external baselines (Attentive FP, MPNN, XGBoost, SVGP), with hyperparameters either tuned on exact GP marginal likelihood (Section 3.3.1) or shared across methods (Section 4.3.3), so the comparisons are not forced by construction. The two assumptions flagged by a careful reader, namely the fixed m=2000 random-Fourier-feature prior samples used for pathwise conditioning and the replacement of f_X^[Z] by f_X in Section 3.2.3, are acknowledged approximations (the text itself shows in Figure 5.5 that pathwise-estimator deviations disappear with exact prior samples, and Section 3.2.3 states the inducing-point error is small 'when m is large and inducing points are close'). These are real correctness/robustness risks, but they are not circular: no equation identifies the approximate prior GP(0, Phi Phi^T) with the exact prior GP(0, K) by definition, and no prediction is a renamed fit. The only circularity-adjacent item is the Proposition 5.2 proof deferred to the authors' own 2024 paper, which I record above as minor and non-load-bearing.

Assumptions & free parameters 6 free parameters · 6 assumptions · 1 invented entities

The central results rest on standard GP mathematics and on the random feature approximation as a finite surrogate for the GP prior. The free parameters are mostly optimizer hyperparameters and problem-specific counts (random features, inducing points, batches) chosen by hand. The latent Kronecker construction is an invented mathematical structure without independent evidence beyond the paper's own experiments.

free parameters (6)
  • Number of random features (m) = 2000 (all chapters)
    Controls fidelity of approximate prior samples used in pathwise conditioning and the pathwise gradient estimator; chosen by hand, no error bound given.
  • SGD/SDD learning rates = beta*n = 0.5 (mean) and 0.1 (samples) for SGD; beta*n = 3 and 0.003 for SDD; per-dataset grid selection
    Step sizes are tuned per dataset to avoid divergence; central to convergence results.
  • Momentum and averaging parameters = rho = 0.9, r = 100/t_max
    Chosen optimizer hyperparameters, not derived.
  • Mini-batch size and block size = 512 (Ch.3), 500/128 (Ch.4), block sizes 1000-2000 (Ch.5)
    Chosen values; affect stochastic gradient variance and solver behavior.
  • Inducing point count = 218k to 1099k for HOUSEELECTRIC; 1024-4096 for baselines
    Selected for inducing-point experiments; performance degrades less than 10% down to 218k.
  • Residual tolerance = 0.01
    Standard in GP literature (Maddox et al. 2021); chosen threshold for iterative solvers.
assumptions (6)
  • ad hoc to paper The random feature approximation Phi Phi^T is sufficiently accurate for prior sampling and marginal likelihood gradients, with 2000 features used throughout.
    All methods rely on this finite-dimensional approximation; no quantitative error control is provided.
  • ad hoc to paper Replacing f_X^[Z] with f_X in the inducing-point sampling objective does not materially change the posterior.
    Stated in Section 3.2.3 with a qualitative 'small when m is large and inducing points close' justification.
  • domain assumption Iterative solvers stopped at relative residual tolerance 0.01 (or early stopping budgets) produce solutions accurate enough for downstream predictions and hyperparameter learning.
    Empirically investigated in Chapter 5, but the relation between residual norm and predictive quality is shown to be weak, so this is an assumption.
  • domain assumption Zero-mean GP prior and homoscedastic Gaussian noise.
    Standard GP regression assumption used in all chapters (Section 2.1.1).
  • ad hoc to paper Fixing the random probe vectors and random feature parameters across outer-loop steps introduces only negligible bias into marginal likelihood optimization.
    Proposition 5.2 is informal and its proof is deferred to the author's separate paper; no proof is in this dissertation.
  • standard math Strong duality, Woodbury identity, Bochner's theorem, Courant-Fischer, and Hutchinson's estimator.
    Standard mathematical background used in derivations.
invented entities (1)
  • Latent Kronecker structure
    purpose: Express the covariance of observed values as a projection of a latent Kronecker product so that Kronecker-based speedups can be applied to non-gridded data.
    The construction is a modeling choice with empirical validation inside the paper; it does not yield an independently checkable prediction outside the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Scalable Gaussian Processes: Advances in Iterative Methods and Pathwise Conditioning." pith.science (2026). https://pith.science/paper/U5YS7LNV

@misc{pith2026250706839,
  author       = {Pith},
  title        = {Pith review of: Scalable Gaussian Processes: Advances in Iterative Methods and Pathwise Conditioning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/U5YS7LNV}},
  note         = {Machine review of arXiv:2507.06839}
}
read the original abstract

Gaussian processes are a powerful framework for uncertainty-aware function approximation and sequential decision-making. Unfortunately, their classical formulation does not scale gracefully to large amounts of data and modern hardware for massively-parallel computation, prompting many researchers to develop techniques which improve their scalability. This dissertation focuses on the powerful combination of iterative methods and pathwise conditioning to develop methodological contributions which facilitate the use of Gaussian processes in modern large-scale settings. By combining these two techniques synergistically, expensive computations are expressed as solutions to systems of linear equations and obtained by leveraging iterative linear system solvers. This drastically reduces memory requirements, facilitating application to significantly larger amounts of data, and introduces matrix multiplication as the main computational operation, which is ideal for modern hardware.

Figures

Figures reproduced from arXiv: 2507.06839 by the authors.

Figure 2.1
Figure 2.1. Samples from a Gaussian process prior using squared exponential covariance [PITH_FULL_IMAGE:figures/full_fig_p026_2_1.png] view at source ↗
Figure 2.2
Figure 2.2. Samples from a Gaussian process prior with Matérn covariance function using [PITH_FULL_IMAGE:figures/full_fig_p027_2_2.png] view at source ↗
Figure 2.3
Figure 2.3. Samples from a Gaussian process prior with periodic covariance function using [PITH_FULL_IMAGE:figures/full_fig_p028_2_3.png] view at source ↗
Figures from the paper (26 more)
Figure 3.1
Figure 3.1. Figure 3.1: Comparison of SGD, CG, and SVGP for GP inference with a squared exponential [PITH_FULL_IMAGE:figures/full_fig_p044_3_1.png]
Figure 3.2
Figure 3.2. Figure 3.2: Left: gradient variance throughout optimisation for a single-sample mini-batch [PITH_FULL_IMAGE:figures/full_fig_p046_3_2.png]
Figure 3.3
Figure 3.3. Figure 3.3: Convergence of GP posterior mean with SGD and CG as a function of time on [PITH_FULL_IMAGE:figures/full_fig_p049_3_3.png]
Figure 3.4
Figure 3.4. Figure 3.4: SGD error and spectral basis functions. Top-left: SGD (blue) and exact GP [PITH_FULL_IMAGE:figures/full_fig_p051_3_4.png]
Figure 3.5
Figure 3.5. Figure 3.5: Test RMSE and NLL as a function of compute time for CG and SGD. Step-like [PITH_FULL_IMAGE:figures/full_fig_p054_3_5.png]
Figure 3.6
Figure 3.6. Figure 3.6: Illustration of a single Thompson sampling acquisition step on a 1D problem. [PITH_FULL_IMAGE:figures/full_fig_p057_3_6.png]
Figure 3.7
Figure 3.7. Figure 3.7: Maximum function values (mean and standard error) obtained by Thompson [PITH_FULL_IMAGE:figures/full_fig_p058_3_7.png]
Figure 4.1
Figure 4.1. Figure 4.1: Comparison of full-batch primal and dual gradient descent on [PITH_FULL_IMAGE:figures/full_fig_p063_4_1.png]
Figure 4.2
Figure 4.2. Figure 4.2: Comparison of stochastic dual descent on the [PITH_FULL_IMAGE:figures/full_fig_p066_4_2.png]
Figure 4.3
Figure 4.3. Figure 4.3: Comparison of optimisation strategies for random coordinate estimator of the [PITH_FULL_IMAGE:figures/full_fig_p068_4_3.png]
Figure 4.4
Figure 4.4. Figure 4.4: Results for the parallel Thompson sampling task. Plots show mean and standard [PITH_FULL_IMAGE:figures/full_fig_p071_4_4.png]
Figure 5.1
Figure 5.1. Figure 5.1: Comparison of relative runtimes for different methods, linear system solvers, and [PITH_FULL_IMAGE:figures/full_fig_p076_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: Marginal likelihood optimisation for iterative GPs. Marginal likelihood optimisation for iterative GPs consists of bi-level optimisation, where the outer loop maximises the marginal likelihood (5.1) using stochastic estimates of its gradient (5.2). Computing these gr…
Figure 5.3
Figure 5.3. Figure 5.3: On the POL and ELEVATORS datasets, the pathwise estimator results in a lower RKHS distance (5.10) between solver initialisation and solution, as predicted by theory (left) (see Equations (5.12) and (5.13) for reference). This results in fewer AP iterations until reac…
Figure 5.4
Figure 5.4. Figure 5.4: On the POL dataset, increasing the number of posterior samples improves the performance of pathwise conditioning until diminishing returns start to manifest with more than 64 samples (left). Furthermore, with 4× as many probe vectors, the total cumulative runtime onl…
Figure 5.5
Figure 5.5. Figure 5.5: Across all datasets and marginal likelihood steps, most hyperparameter tra [PITH_FULL_IMAGE:figures/full_fig_p084_5_5.png]
Figure 5.6
Figure 5.6. Figure 5.6: Two-dimensional cross-sections of top eigendirections of the inner-loop quadratic [PITH_FULL_IMAGE:figures/full_fig_p085_5_6.png]
Figure 5.7
Figure 5.7. Figure 5.7: Required number of linear system solver iterations to reach the tolerance [PITH_FULL_IMAGE:figures/full_fig_p086_5_7.png]
Figure 5.8
Figure 5.8. Figure 5.8: Across marginal likelihood steps and datasets, warm starting results in hyperpa [PITH_FULL_IMAGE:figures/full_fig_p087_5_8.png]
Figure 5.9
Figure 5.9. Figure 5.9: Relative residual norms of the probe vector linear systems at each marginal [PITH_FULL_IMAGE:figures/full_fig_p089_5_9.png]
Figure 5.10
Figure 5.10. Figure 5.10: Relative residual norms and test log-likelihoods during marginal likelihood [PITH_FULL_IMAGE:figures/full_fig_p091_5_10.png]
Figure 6.1
Figure 6.1. Figure 6.1: The joint covariance matrix over {(s1, t1),(s1, t2),(s2, t1),(s2, t2),(s2, t3)}, namely two out of three time steps at spatial location s1 and three out of three time steps at spatial location s2, can be expressed as the projection of a latent Kronecker product. 6.2.…
Figure 6.2
Figure 6.2. Figure 6.2: Illustration of computational resources used during kernel evaluation and ma [PITH_FULL_IMAGE:figures/full_fig_p098_6_2.png]
Figure 6.3
Figure 6.3. Figure 6.3: Predicting the inverse dynamics of an anthropomorphic robot arm with seven [PITH_FULL_IMAGE:figures/full_fig_p101_6_3.png]
Figure 6.4
Figure 6.4. Figure 6.4: Learning curve prediction on the Fashion-MNIST dataset from the LCBench [PITH_FULL_IMAGE:figures/full_fig_p102_6_4.png]
Figure 6.5
Figure 6.5. Figure 6.5: Illustration of daily temperature and precipitation data from the Nordic Gridded [PITH_FULL_IMAGE:figures/full_fig_p104_6_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 79 canonical work pages

  1. [1]

    Antorán, S

    J. Antorán, S. Padhy, R. Barbano, E. T. Nalisnick, D. Janz, and J. M. Hernández- Lobato. Sampling-based inference for large linear models, with application to lin- earised Laplace. In International Conference on Learning Representations, 2023. Cited on pages 31, 34, 41, 57, 68, 70, 76

  2. [2]

    Artemev, D

    A. Artemev, D. R. Burt, and M. van der Wilk. Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients. In Interna- tional Conference on Machine Learning, 2021. Cited on page 66

  3. [3]

    Bauer, M

    M. Bauer, M. Van der Wilk, and C. E. Rasmussen. Understanding Probabilistic Sparse Gaussian Process Approximations. In Advances in Neural Information Processing Systems, 2016. Cited on page 83

  4. [4]

    Belkin, D

    M. Belkin, D. Hsu, S. Ma, and S. Mandal. Reconciling modern machine-learning prac- tice and the classical bias-variance trade-off. Proceedings of the National Academy of Sciences, 2019. Cited on page 30

  5. [5]

    Bernhardsson

    E. Bernhardsson. Approximate Nearest Neighbors Oh Yeah, 2012. Cited on page 43

  6. [6]

    C. M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006. Cited on page 17

  7. [7]

    Bo and C

    L. Bo and C. Sminchisescu. Greedy Block Coordinate Descent for Large Scale Gaussian Process Regression. In Uncertainty in Artificial Intelligence, 2008. Cited on page 57

  8. [8]

    E. V . Bonilla, K. M. A. Chai, and C. K. I. Williams. Multi-task Gaussian Process Prediction. In Advances in Neural Information Processing Systems, 2007. Cited on pages 1, 25, 83, 88

Show all 121 references
  1. [9]

    S. P. Boyd and L. Vandenberghe.Convex Optimization. Cambridge University Press,

  2. [10]

    H. Chen, L. Zheng, R. Al Kontar, and G. Raskutti. Gaussian Process Parameter Estimation Using Mini-batch Stochastic Gradient Descent: Convergence Guarantees and Empirical Benefits. Journal of Machine Learning Research, 23, 2022. Cited on page 30

  3. [11]

    H. Chen, L. Zheng, R. Al Kontar, and G. Raskutti. Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian Processes. In Advances in Neural Information Processing Systems, 2020. Cited on pages 30, 73

  4. [12]

    Csató and M

    L. Csató and M. Opper. Sparse On-Line Gaussian Processes. Neural Computation, 14, 3, 2002. Cited on page 20. 100 References

  5. [13]

    Cutajar, M

    K. Cutajar, M. Osborne, J. Cunningham, and M. Filippone. Preconditioning Kernel Matrices. In International Conference on Machine Learning, 2016. Cited on page 30

  6. [14]

    B. Dai, B. Xie, N. He, Y . Liang, A. Raj, M. -F. F. Balcan, and L. Song. Scalable Kernel Methods via Doubly Stochastic Gradients. InAdvances in Neural Information Processing Systems, 2014. Cited on pages 30, 48, 56, 57

  7. [15]

    Defazio and K

    A. Defazio and K. Mishchenko. Learning-Rate-Free Learning by D-Adaptation. In International Conference on Machine Learning, 2023. Cited on page 97

  8. [16]

    M. P. Deisenroth, D. Fox, and C. E. Rasmussen. Gaussian Processes for Data-Efficient Learning in Robotics and Control. IEEE Trans. Pattern Anal. Mach. Intell., 2015. Cited on page 64

  9. [17]

    M. P. Deisenroth and C. E. Rasmussen. PILCO: A Model-Based and Data-Efficient Approach to Policy Search. In International Conference on Machine Learning, 2011. Cited on page 82

  10. [18]

    Dieuleveut, N

    A. Dieuleveut, N. Flammarion, and F. Bach. Harder, Better, Faster, Stronger Conver- gence Rates for Least-Squares Regression. Journal of Machine Learning Research,

  11. [19]

    Dua and C

    D. Dua and C. Graff. UCI Machine Learning Repository, 2017. Cited on pages 41, 50, 58, 78

  12. [20]

    Elahi, F

    M. Elahi, F. Ricci, and N. Rubens. A survey of active learning in collaborative filtering recommender systems. Computer Science Review, 2016. Cited on page 44

  13. [21]

    Elsken, J

    T. Elsken, J. H. Metzen, and F. Hutter. Neural Architecture Search: A Survey.Journal of Machine Learning Research, 20, 55, 2019. Cited on page 89

  14. [22]

    E. N. Epperly, J. A. Tropp, and R. J. Webber. XTrace: Making the Most of Every Sample in Stochastic Trace Estimation. Matrix Analysis and Applications , 2024. Cited on page 72

  15. [23]

    Eschenhagen, A

    R. Eschenhagen, A. Immer, R. E. Turner, F. Schneider, and P. Hennig. Kronecker- Factored Approximate Curvature for Modern Neural Network Architectures. In Advances in Neural Information Processing Systems, 2023. Cited on page 98

  16. [24]

    T.-T. Frie, N. Cristianini, and C. Campbell. The Kernel-Adatron Algorithm: a Fast and Simple Learning Procedure for Support Vector Machines. In International Conference on Machine Learning, 1998. Cited on page 57

  17. [25]

    García-Ortegón, G

    M. García-Ortegón, G. N. C. Simm, A. J. Tripp, J. M. Hernández-Lobato, A. Bender, and S. Bacallado. DOCKSTRING: Easy Molecular Docking Yields Better Bench- marks for Ligand Design. Journal of Chemical Information and Modeling , 2022. Cited on pages 58, 60, 61

  18. [26]

    Gardner, G

    J. Gardner, G. Pleiss, K. Q. Weinberger, D. Bindel, and A. G. Wilson. GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration. In Advances in Neural Information Processing Systems, 2018. Cited on pages 1, 28, 30, 48, 64, 67, 70, 73, 76, 82, 88, 90, 92

  19. [27]

    Gardner, G

    J. Gardner, G. Pleiss, R. Wu, K. Weinberger, and A. Wilson. Product Kernel Inter- polation for Scalable Gaussian Processes. In International Conference on Artificial Intelligence and Statistics, 2018. Cited on pages 1, 26, 30. References 101

  20. [28]

    R. Garnett. Bayesian Optimization. Cambridge University Press, 2023. Cited on pages 2, 82

  21. [29]

    Ghasemi, M

    P. Ghasemi, M. Karbasi, A. Z. Nouri, M. S. Tabrizi, and H. M. Azamathulla. Ap- plication of Gaussian process regression to forecast multi-step ahead SPEI drought index. Alexandria Engineering Journal, 2021. Cited on page 64

  22. [30]

    Gilmer, S

    J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl. Neural Message Passing for Quantum Chemistry. In International Conference on Machine Learning,

  23. [31]

    Gómez-Bombarelli, J

    R. Gómez-Bombarelli, J. N. Wei, D. Duvenaud, J. M. Hernández-Lobato, B. Sánchez- Lengeling, D. Sheberla, J. Aguilera-Iparraguirre, T. D. Hirzel, R. P. Adams, and A. Aspuru-Guzik. Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules. American Che...

  24. [32]

    P. C. Hansen. Rank-Deficient and Discrete Ill-posed Problems: Numerical Aspects of Linear Inversion. Society for Industrial and Applied Mathematics, 1998. Cited on page 40

  25. [33]

    Hensman, N

    J. Hensman, N. Fusi, and N. D. Lawrence. Gaussian Processes for Big Data. In Uncertainty in Artificial Intelligence, 2013. Cited on pages 1, 22, 30, 34, 36, 40, 48, 58, 64, 82, 89

  26. [34]

    J. M. Hernández-Lobato, M. W. Hoffman, and Z. Ghahramani. Predictive Entropy Search for Efficient Global Optimization of Black-box Functions. In Advances in Neural Information Processing Systems, 2014. Cited on page 30

  27. [35]

    J. M. Hernández-Lobato, J. Requeima, E. O. Pyzer-Knapp, and A. Aspuru-Guzik. Parallel and Distributed Thompson Sampling for Large-scale Accelerated Exploration of Chemical Space. In International Conference on Machine Learning, 2017. Cited on pages 44, 59

  28. [36]

    M. R. Hestenes and E. Stiefel. Methods of Conjugate Gradients for Solving Linear Systems. Journal of Research of the National Bureau of Standards, 49, 6, 1952. Cited on page 27

  29. [37]

    R. A. Horn and C. R. Johnson. Matrix Analysis. Cambridge University Press, 2012. Cited on page 115

  30. [38]

    R. A. Horn and C. R. Johnson. Topics in Matrix Analysis. Cambridge University Press, 1991. Cited on page 72

  31. [39]

    Hutchinson

    M. Hutchinson. A Stochastic Estimator of the Trace of the Influence Matrix for Lapla- cian Smoothing Splines. Communications in Statistics - Simulation and Computation, 19(2), 1990. Cited on pages 28, 67

  32. [40]

    S. Ioffe. Improved Consistent Sampling, Weighted Minhash and L1 Sketching. In International Conference on Data Mining, 2010. Cited on page 61

  33. [41]

    Jankowiak, G

    M. Jankowiak, G. Pleiss, and J. Gardner. Parametric Gaussian Process Regressors. In International Conference on Machine Learning, 2020. Cited on page 83

  34. [42]

    Jiang, D

    S. Jiang, D. R. Jiang, M. Balandat, B. Karrer, J. R. Gardner, and R. Garnett. Efficient Nonmyopic Bayesian Optimization via One-Shot Multi-Step Trees. In Advances in Neural Information Processing Systems, 2020. Cited on page 13. 102 References

  35. [43]

    Jin and Ž

    B. Jin and Ž. Kereta. On the Convergence of Stochastic Gradient Descent for Linear Inverse Problems in Banach Spaces. SIAM Journal on Imaging Sciences, 2023. Cited on page 40

  36. [44]

    Kapoor, M

    S. Kapoor, M. Finzi, K. A. Wang, and A. G. G. Wilson. SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes. In International Conference on Machine Learning, 2021. Cited on pages 1, 26

  37. [45]

    Kivinen, A

    J. Kivinen, A. J. Smola, and R. C. Williamson. Online Learning with Kernels. IEEE Transactions on Signal Processing, 2004. Cited on pages 48, 57

  38. [46]

    Li, J.-F

    Z. Li, J.-F. Ton, D. Oglic, and D. Sejdinovic. Towards a Unified Analysis of Random Fourier Features. In International Conference on Machine Learning, 2019. Cited on pages 1, 55

  39. [47]

    J. A. Lin, S. Ament, M. Balandat, and E. Bakshy. Scaling Gaussian Processes for Learning Curve Prediction via Latent Kronecker Structure. In NeurIPS Bayesian Decision-making and Uncertainty Workshop, 2024. Cited on pages 1, 5

  40. [48]

    J. A. Lin, S. Ament, M. Balandat, D. Eriksson, J. M. Hernández-Lobato, and E. Bak- shy. Scalable Gaussian Processes with Latent Kronecker Structure. In International Conference on Machine Learning, 2025. Cited on pages 1, 5

  41. [49]

    J. A. Lin, J. Antorán, S. Padhy, D. Janz, J. M. Hernández-Lobato, and A. Terenin. Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent. In Advances in Neural Information Processing Systems, 2023. Cited on pages 1, 4, 48, 49, 57–60, 64, 65, 67–70, 73, 76–78, 88

  42. [50]

    J. A. Lin, S. Padhy, J. Antorán, A. Tripp, A. Terenin, C. Szepesvári, J. M. Hernández- Lobato, and D. Janz. Stochastic Gradient Descent for Gaussian Processes Done Right. In International Conference on Learning Representations, 2024. Cited on pages 1, 4, 65, 67, 70, 73, 76, 78

  43. [51]

    J. A. Lin, S. Padhy, B. Mlodozeniec, J. Antorán, and J. M. Hernández-Lobato. Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes. In Advances in Neural Information Processing Systems, 2024. Cited on pages 1, 5, 74

  44. [52]

    J. A. Lin, S. Padhy, B. Mlodozeniec, and J. M. Hernández-Lobato. Warm Start Marginal Likelihood Optimisation for Iterative Gaussian Processes. In Advances in Approximate Bayesian Inference, 2024. Cited on pages 1, 5, 73

  45. [53]

    W. Liu, P. P. Pokharel, and J. C. Principe. The Kernel Least-Mean-Square Algorithm. IEEE Transactions on Signal Processing, 2008. Cited on page 57

  46. [54]

    W. J. Maddox, M. Balandat, A. G. Wilson, and E. Bakshy. Bayesian Optimization with High-Dimensional Outputs. In Advances in Neural Information Processing Systems, 2021. Cited on page 85

  47. [55]

    W. J. Maddox, S. Kapoor, and A. G. Wilson. When are Iterative Gaussian Processes Reliably Accurate? In ICML OPTML Workshop, 2021. Cited on pages 67, 70

  48. [56]

    Mandt, M

    S. Mandt, M. D. Hoffman, and D. M. Blei. Stochastic Gradient Descent as Approxi- mate Bayesian Inference. Journal of Machine Learning Research, 2017. Cited on page 30. References 103

  49. [57]

    J. Mercer. Functions of Positive and Negative Type, and their Connection with the Theory of Integral Equations. Philosophical Transactions of the Royal Society of London, Series A, 209, 1909. Cited on page 23

  50. [58]

    Meyer, C

    R. Meyer, C. Musco, C. Musco, and D. Woodruff. Hutch++: Optimal Stochastic Trace Estimation. Symposium on Simplicity in Algorithms , 2021, 2021. Cited on page 72

  51. [59]

    Nesterov

    Y . Nesterov. A method for unconstrained convex minimization problem with the rate of convergence O(1/k2). In Doklady Akademii Nauk SSSR, 1983. Cited on page 56

  52. [60]

    K. B. Petersen and M. S. Pedersen. The Matrix Cookbook. Technical University of Denmark, 2006. Cited on page 71

  53. [61]

    Pinzi and G

    L. Pinzi and G. Rastelli. Molecular Docking: Shifting Paradigms in Drug Discovery. International Journal of Molecular Sciences, 2019. Cited on page 60

  54. [62]

    Pleiss, J

    G. Pleiss, J. R. Gardner, K. Q. Weinberger, and A. G. Wilson. Constant-Time Predic- tive Distributions for Gaussian Processes. In International Conference on Machine Learning, 2018. Cited on page 28

  55. [63]

    B. T. Polyak. New stochastic approximation type procedures. Avtomatika i Tele- mekhanika, 1990. Cited on page 56

  56. [64]

    B. T. Polyak. Some methods of speeding up the convergence of iteration meth- ods. USSR Computational Mathematics and Mathematical Physics, 1964. Cited on page 56

  57. [65]

    B. T. Polyak and A. B. Juditsky. Acceleration of Stochastic Approximation by Averaging. SIAM Journal on Control and Optimization, 1992. Cited on page 56

  58. [66]

    Quiñonero-Candela and C

    J. Quiñonero-Candela and C. E. Rasmussen. A Unifying View of Sparse Approximate Gaussian Process Regression. Journal of Machine Learning Research, 6, 1, 2005. Cited on pages 1, 19, 20, 64, 82

  59. [67]

    Rahimi and B

    A. Rahimi and B. Recht. Random Features for Large-scale Kernel Machines. In Advances in Neural Information Processing Systems, 2008. Cited on pages 1, 23, 32, 68

  60. [68]

    Ralaivola, S

    L. Ralaivola, S. J. Swamidass, H. Saigo, and P. Baldi. Graph kernels for chemical informatics. Neural Networks, 2005. Cited on page 60

  61. [69]

    C. E. Rasmussen and C. K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, 2006. Cited on pages 7, 64

  62. [70]

    C. Riis, F. Antunes, F. Hüttel, C. Lima Azevedo, and F. Pereira. Bayesian Active Learning with Fully Bayesian Gaussian Processes. InAdvances in Neural Information Processing Systems, 2022. Cited on page 82

  63. [71]

    Rogers and M

    D. Rogers and M. Hahn. Extended-Connectivity Fingerprints. Journal of Chemical Information and Modeling, 2010. Cited on page 60

  64. [72]

    Rubens, M

    N. Rubens, M. Elahi, M. Sugiyama, and D. Kaplan. Active Learning in Recommender Systems. Recommender Systems Handbook, 2015. Cited on page 44

  65. [73]

    A. Rudi, L. Carratino, and L. Rosasco. Falkon: An Optimal Large Scale Kernel Method. In Advances in Neural Information Processing Systems , 2017. Cited on page 48. 104 References

  66. [74]

    D. Ruppert. Efficient Estimations from a Slowly Convergent Robbins-Monro Process. Technical report, Cornell University, 1988. Cited on page 57

  67. [75]

    Saunders, A

    C. Saunders, A. Gammerman, and V . V ovk. Ridge Regression Learning Algorithm in Dual Variables. In International Conference on Machine Learning, 1998. Cited on page 57

  68. [76]

    Schölkopf, R

    B. Schölkopf, R. Herbrich, and A. J. Smola. A Generalized Representer Theorem. In Computational Learning Theory, 2001. Cited on page 31

  69. [77]

    Schölkopf and A

    B. Schölkopf and A. J. Smola. Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. MIT Press, 2002. Cited on pages 50, 57

  70. [78]

    Schwaighofer and V

    A. Schwaighofer and V . Tresp. Transductive and Inductive Methods for Approximate Gaussian Process Regression. InAdvances in Neural Information Processing Systems,

  71. [79]

    M. W. Seeger, C. K. I. Williams, and N. D. Lawrence. Fast Forward Selection to Speed Up Sparse Gaussian Process Regression. In International Workshop on Artificial Intelligence and Statistics, 2003. Cited on page 20

  72. [80]

    Shalev-Shwartz and T

    S. Shalev-Shwartz and T. Zhang. Stochastic Dual Coordinate Ascent Methods for Regularized Loss Minimization. Journal of Machine Learning Research, 2013. Cited on pages 48, 57, 64, 67

  73. [81]

    J. R. Shewchuk. An Introduction to the Conjugate Gradient Method Without the Ag- onizing Pain. Technical report, Carnegie Mellon University, 1994. Cited on page 28

  74. [82]

    Silverman

    B. Silverman. Some Aspects of the Spline Smoothing Approach to Non-Parametric Regression Curve Fitting. Journal of the Royal Statistical Society, Series B, 47, 1,

  75. [83]

    A. J. Smola and P. L. Bartlett. Sparse Greedy Gaussian Process Regression. In Advances in Neural Information Processing Systems, 2001. Cited on page 20

  76. [84]

    Snelson and Z

    E. Snelson and Z. Ghahramani. Sparse Gaussian Processes using Pseudo-inputs. In Advances in Neural Information Processing Systems, 2005. Cited on page 20

  77. [85]

    Snoek, H

    J. Snoek, H. Larochelle, and R. P. Adams. Practical Bayesian Optimization of Ma- chine Learning Algorithms. In Advances in Neural Information Processing Systems,

  78. [86]

    Stegle, C

    O. Stegle, C. Lippert, J. M. Mooij, N. Lawrence, and K. Borgwardt. Efficient infer- ence in matrix-variate Gaussian models with iid observation noise. In Advances in Neural Information Processing Systems, 2011. Cited on pages 1, 25, 83

  79. [87]

    D. J. Sutherland and J. Schneider. On the Error of Random Fourier Features. In Uncertainty in Artificial Intelligence, 2015. Cited on pages 1, 23, 68

  80. [88]

    Sutskever, J

    I. Sutskever, J. Martens, G. Dahl, and G. Hinton. On the importance of initialization and momentum in deep learning. In International Conference on Machine Learning,

  81. [89]

    Swersky, J

    K. Swersky, J. Snoek, and R. P. Adams. Freeze-Thaw Bayesian Optimization. arXiv:1406.3896, 2014. Cited on page 98

  82. [90]

    K. Tazi, J. A. Lin, R. Viljoen, A. Gardner, T. John, H. Ge, and R. E. Turner. Beyond Intuition, a Framework for Applying GPs to Real-World Data. InICML Structured Probabilistic Inference & Generative Modeling Workshop, 2023. Cited on page 64. References 105

  83. [91]

    Terenin, D

    A. Terenin, D. R. Burt, A. Artemev, S. Flaxman, M. van der Wilk, C. E. Rasmussen, and H. Ge. Numerically Stable Sparse Gaussian Processes via Minimum Separation using Cover Trees. Journal of Machine Learning Research, 2023. Cited on page 42

  84. [92]

    Y . Tian, Y . Zhang, and H. Zhang. Recent Advances in Stochastic Gradient Descent in Deep Learning. Mathematics, 11, 3, 2023. Cited on page 30

  85. [93]

    M. K. Titsias. Variational learning of inducing variables in sparse Gaussian processes. In Artificial Intelligence and Statistics, 2009. Cited on pages 1, 21, 22, 30, 34, 48, 64, 82, 83

  86. [94]

    M. K. Titsias. Variational model selection for sparse Gaussian process regression. Technical report, University of Manchester, 2009. Cited on pages 1, 21, 22, 64, 82

  87. [95]

    Tripp, S

    A. Tripp, S. Bacallado, S. Singh, and J. M. Hernández-Lobato. Tanimoto Random Features for Scalable Molecular Machine Learning. In Advances in Neural Informa- tion Processing Systems, 2023. Cited on pages 1, 60, 61

  88. [96]

    Trott and A

    O. Trott and A. J. Olson. AutoDock Vina: Improving the Speed and Accuracy of Docking with a New Scoring Function, Efficient Optimization, and Multithreading. Journal of Computational Chemistry, 2010. Cited on page 60

  89. [97]

    S. Tu, R. Roelofs, S. Venkataraman, and B. Recht. Large Scale Kernel Learning using Block Coordinate Descent. arXiv:1602.05310, 2016. Cited on pages 57, 64, 67

  90. [98]

    O. E. Tveito, I. Bjørdal, A. O. Skjelvåg, and B. Aune. A GIS-based agro-ecoglogical decision system based on gridded climatology. Metoeorl. Appl., 12, 2005. Cited on page 92

  91. [99]

    O. E. Tveito, E. J. Førland, R. Heino, I. Hanssen-Bauer, H. Alexandersson, B. Dahlström, A. Drebs, C. Kern-Hansen, T. Jónsson, E. Vaarby-Laursen, and E. West- man. Nordic Temperature Maps. DNMI Klima 9/00 KLIMA., 2000. Cited on page 92

  92. [100]

    V . Vapnik. The Nature of Statistical Learning. Springer, 1995. Cited on page 30

  93. [101]

    A. V . Varre, L. Pillaud-Vivien, and N. Flammarion. Last iterate convergence of SGD for Least-Squares in the Interpolation regime. In Advances in Neural Information Processing Systems, 2021. Cited on pages 48, 55, 57

  94. [102]

    Wahba, X

    G. Wahba, X. Lin, F. Gao, D. Xiang, R. Klein, and B. Klein. The Bias-Variance Tradeoff and the Randomized GACV. InAdvances in Neural Information Processing Systems, 1999. Cited on page 20

  95. [103]

    K. A. Wang, G. Pleiss, J. R. Gardner, S. Tyree, K. Q. Weinberger, and A. G. Wilson. Exact Gaussian Processes on a Million Data Points. In Advances in Neural Informa- tion Processing Systems, 2019. Cited on pages 1, 28, 30, 40, 41, 48, 58, 64, 66, 67, 69, 73, 78, 82

  96. [104]

    Wenger, G

    J. Wenger, G. Pleiss, M. Pförtner, Hennig, and J. P. Cunningham. Posterior and Computational Uncertainty in Gaussian processes. InAdvances in Neural Information Processing Systems, 2022. Cited on pages 1, 22

  97. [105]

    Wenger, K

    J. Wenger, K. Wu, P. Hennig, J. R. Gardner, G. Pleiss, and J. P. Cunningham. Computation-Aware Gaussian Processes: Model Selection And Linear-Time In- ference. In Advances in Neural Information Processing Systems , 2024. Cited on pages 1, 22, 89. 106 References

  98. [106]

    V . Wild, M. Kanagawa, and D. Sejdinovic. Connections and Equivalences between the Nyström Method and Sparse Variational Gaussian Processes. arXiv:2106.01121,

  99. [107]

    W. J. Wilkinson, P. E. Chang, M. R. Andersen, and A. Solin. State Space Expecta- tion Propagation: Efficient Inference Schemes for Temporal Gaussian Processes. In International Conference on Machine Learning, 2020. Cited on page 30

  100. [108]

    C. K. I. Williams and M. Seeger. Using the Nyström Method to Speed Up Kernel Machines. In Advances in Neural Information Processing Systems, 2000. Cited on pages 19, 21, 48

  101. [109]

    Wilson and H

    A. Wilson and H. Nickisch. Kernel Interpolation for Scalable Structured Gaussian Processes (KISS-GP). In International Conference on Machine Learning, 2015. Cited on pages 1, 26, 30

  102. [110]

    J. T. Wilson, V . Borovitskiy, A. Terenin, P. Mostowsky, and M. P. Deisenroth. Ef- ficiently Sampling Functions from Gaussian Process Posteriors. In International Conference on Machine Learning, 2020. Cited on pages 2, 11, 13, 28, 31, 45, 49, 67, 68, 85

  103. [111]

    J. T. Wilson, V . Borovitskiy, A. Terenin, P. Mostowsky, and M. P. Deisenroth. Path- wise Conditioning of Gaussian Processes. Journal of Machine Learning Research, 22, 1, 2021. Cited on pages 2, 11, 13, 28, 31, 49, 67, 85

  104. [112]

    K. Wu, J. Wenger, H. Jones, G. Pleiss, and J. R. Gardner. Large-Scale Gaussian Processes via Alternating Projection. In International Conference on Artificial Intel- ligence and Statistics, 2024. Cited on pages 57, 64, 66, 67, 69, 70, 73, 76, 78

  105. [113]

    L. Wu, G. Pleiss, and J. P. Cunningham. Variational Nearest Neighbor Gaussian Process. In International Conference on Machine Learning, 2022. Cited on pages 1, 22, 89

  106. [114]

    Xiong, D

    Z. Xiong, D. Wang, X. Liu, F. Zhong, X. Wan, X. Li, Z. Li, X. Luo, K. Chen, H. Jiang, et al. Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism. Journal of Medicinal Chemistry, 2019. Cited on page 61

  107. [115]

    Y . Yang, K. Yao, M. P. Repasky, K. Leswing, R. Abel, B. K. Shoichet, and S. V . Jerome. Efficient Exploration of Chemical Space with Docking and Deep Learning. Journal of Chemical Theory and Computation, 2021. Cited on page 60

  108. [116]

    S. Zhe, W. Xing, and R. M. Kirby. Scalable High-Order Gaussian Process Regression. In International Conference on Artificial Intelligence and Statistics, 2019. Cited on page 83

  109. [117]

    Zimmer, M

    L. Zimmer, M. Lindauer, and F. Hutter. Auto-PyTorch Tabular: Multi-Fidelity Met- aLearning for Efficient and Robust AutoDL. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43, 9, 2021. Cited on pages 89, 90

  110. [118]

    D. Zou, J. Wu, V . Braverman, Q. Gu, and S. M. Kakade. Benign Overfitting of Constant-Stepsize SGD for Linear Regression. In Conference on Learning Theory,

  111. [2012]

    Cited on pages 2, 30, 64

  112. [2017]

    Cited on pages 48, 55, 57

  113. [2021]

    Appendix A Mathematical Derivations This appendix contains mathematical derivations which are too verbose to be included in the main body of this dissertation

    Cited on pages 30, 40. Appendix A Mathematical Derivations This appendix contains mathematical derivations which are too verbose to be included in the main body of this dissertation. The following derivations are based on and adapted from: • J. A. Lin, J. Antorán, S. Padhy, D....

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.