REVIEW 3 cited by
Connections and Equivalences between the Nystr\"om Method and Sparse Variational Gaussian Processes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We investigate the connections between sparse approximation methods for making kernel methods and Gaussian processes (GPs) scalable to large-scale data, focusing on the Nystr\"om method and the Sparse Variational Gaussian Processes (SVGP). While sparse approximation methods for GPs and kernel methods share some algebraic similarities, the literature lacks a deep understanding of how and why they are related. This may pose an obstacle to the communications between the GP and kernel communities, making it difficult to transfer results from one side to the other. Our motivation is to remove this obstacle, by clarifying the connections between the sparse approximations for GPs and kernel methods. In this work, we study the two popular approaches, the Nystr\"om and SVGP approximations, in the context of a regression problem, and establish various connections and equivalences between them. In particular, we provide an RKHS interpretation of the SVGP approximation, and show that the Evidence Lower Bound of the SVGP contains the objective function of the Nystr\"om approximation, revealing the origin of the algebraic equivalence between the two approaches. We also study recently established convergence results for the SVGP and how they are related to the approximation quality of the Nystr\"om method.
Forward citations
Cited by 3 Pith papers
-
The Price of Linear Time: Error Analysis of Structured Kernel Interpolation
For cubic SKI the inducing-point count should grow as n^{d/3}; the advertised linear-time regime d≤3 is incorrect because at d=3 the paper's own inequality forces error to grow with n.
-
Adaptive Nystr\"om for Gaussian Process Regression
Interleaving greedy trace-residual landmark selection with hyperparameter updates yields Nyström GPs that match exact GP accuracy far more stably than random landmarks on standard test functions.
-
Scalable Gaussian Processes: Advances in Iterative Methods and Pathwise Conditioning
The thesis shows that iterative linear solvers plus pathwise conditioning scale Gaussian processes to millions of data points, introducing SGD-based, dual-descent, warm-started, and latent-Kronecker methods for infere...
Discussion (0). Continue with ORCID to comment.