A Multi-Stage CUDA Kernel for Floyd-Warshall
classification
💻 cs.DC
cs.PF
keywords
algorithmcudafloyd-warshallachieveall-pairsallowappliedapproximately
read the original abstract
We present a new implementation of the Floyd-Warshall All-Pairs Shortest Paths algorithm on CUDA. Our algorithm runs approximately 5 times faster than the previously best reported algorithm. In order to achieve this speedup, we applied a new technique to reduce usage of on-chip shared memory and allow the CUDA scheduler to more effectively hide instruction latency.
This paper has not been read by Pith yet.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.