Bucketed approximate top-k, retrieving a few elements per chunk, is 2-4x faster than exact top-k on GPUs with negligible downstream loss, and different bucket settings are optimal for small versus large k.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Approximate Top-$k$ for Increased Parallelism
Bucketed approximate top-k, retrieving a few elements per chunk, is 2-4x faster than exact top-k on GPUs with negligible downstream loss, and different bucket settings are optimal for small versus large k.