NPUEval evaluates LLM-generated NPU kernels on real AMD hardware, finding that frontier models rarely produce vectorized code, with average vectorization scores near 10%.
Evocodebench: An evolving code generation benchmark aligned with real-world code repositories, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.PL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
NPUEval: Optimizing NPU Kernels with LLMs and Open Source Compilers
NPUEval evaluates LLM-generated NPU kernels on real AMD hardware, finding that frontier models rarely produce vectorized code, with average vectorization scores near 10%.