REVIEW 2 cited by
Large Scale Kernel Learning using Block Coordinate Descent
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We demonstrate that distributed block coordinate descent can quickly solve kernel regression and classification problems with millions of data points. Armed with this capability, we conduct a thorough comparison between the full kernel, the Nystr\"om method, and random features on three large classification tasks from various domains. Our results suggest that the Nystr\"om method generally achieves better statistical accuracy than random features, but can require significantly more iterations of optimization. Lastly, we derive new rates for block coordinate descent which support our experimental findings when specialized to kernel methods.
Forward citations
Cited by 2 Pith papers
-
Joker: Joint Optimization Framework for Lightweight Kernel Machines
A dual block-coordinate trust-region solver with random Fourier features trains KRR, KLR, and SVM on millions of samples with 1 to 5 GB GPU memory and accuracy matching or beating Falkon, EigenPro3, and ThunderSVM.
-
Scalable Gaussian Processes: Advances in Iterative Methods and Pathwise Conditioning
The thesis shows that iterative linear solvers plus pathwise conditioning scale Gaussian processes to millions of data points, introducing SGD-based, dual-descent, warm-started, and latent-Kronecker methods for infere...
Discussion (0). Continue with ORCID to comment.