A TVM-based pipeline generates 20,000 CUDA-CPU code pairs, and fine-tuning code LLMs on this data improves transpilation success and CPU performance, with an average speedup improvement of 43.8%.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
A TVM-based pipeline generates 20,000 CUDA-CPU code pairs, and fine-tuning code LLMs on this data improves transpilation success and CPU performance, with an average speedup improvement of 43.8%.