CuTile achieves up to 2.5x FlashAttention-2 throughput on B200 with 60 lines of Python but shows significant cross-architecture portability gaps, reaching only 53% of FlashAttention-2 on RTX PRO 6000.
Roofline: An Insightful Visual Performance Model for Multicore Architectures
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
CuTile achieves up to 2.5x FlashAttention-2 throughput on B200 with 60 lines of Python but shows significant cross-architecture portability gaps, reaching only 53% of FlashAttention-2 on RTX PRO 6000.