In a six-model comparison, DeepSeek-Coder-33B generated the most passable directive-based compiler tests (Pass@1=0.434) and Qwen2.5-Coder-32B judged test validity best (F1=0.735).
Using a larg e language model as a building block to generate usablevalidation and v erification suite for openmp,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
LLM4VV: Evaluating Cutting-Edge LLMs for Generation and Evaluation of Directive-Based Parallel Programming Model Compiler Tests
In a six-model comparison, DeepSeek-Coder-33B generated the most passable directive-based compiler tests (Pass@1=0.434) and Qwen2.5-Coder-32B judged test validity best (F1=0.735).