On an irregular hash-blocked TSDF kernel, CUDA C++ and Rust are close (1.0-3.3x) on the insertion stage, Triton is 11-32x slower, and Triton's fixed-bound probe silently drops surface data at higher load factors.
TartanAir: A dataset to push the limits of visual SLAM
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload
On an irregular hash-blocked TSDF kernel, CUDA C++ and Rust are close (1.0-3.3x) on the insertion stage, Triton is 11-32x slower, and Triton's fixed-bound probe silently drops surface data at higher load factors.