A two-stage prompting workflow with a TL language lets LLMs generate FlashAttention-class GPU kernels that match or beat hand-optimized libraries on multiple GPU generations.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
A two-stage prompting workflow with a TL language lets LLMs generate FlashAttention-class GPU kernels that match or beat hand-optimized libraries on multiple GPU generations.