REVIEW 2 cited by
Kokkos Kernels: Performance Portable Sparse/Dense Linear Algebra and Graph Kernels
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As hardware architectures are evolving in the push towards exascale, developing Computational Science and Engineering (CSE) applications depend on performance portable approaches for sustainable software development. This paper describes one aspect of performance portability with respect to developing a portable library of kernels that serve the needs of several CSE applications and software frameworks. We describe Kokkos Kernels, a library of kernels for sparse linear algebra, dense linear algebra and graph kernels. We describe the design principles of such a library and demonstrate portable performance of the library using some selected kernels. Specifically, we demonstrate the performance of four sparse kernels, three dense batched kernels, two graph kernels and one team level algorithm.
Forward citations
Cited by 2 Pith papers
-
Fast Entropy Decoding for Sparse MVM on GPUs
Encoding CSR sparse matrices with dtANS, a decoupled segment-parallel tANS variant, compresses them up to 11.77x over cuSPARSE formats and accelerates GPU SpMVM up to 3.48x on large matrices.
-
ShyLU node: On-node Scalable Solvers and Preconditioners Recent Progresses and Current Performance
ShyLU-node's Basker, Tacho, and FastILU solvers show competitive performance against Pardiso, SuperLU, and Kokkos-Kernels ILU in benchmarks for circuit, ice-sheet, and 3D elasticity problems.
Discussion (0). Continue with ORCID to comment.