COMPASS benchmark shows that evaluating code generation with only correctness misses large differences in runtime efficiency and code quality among frontier LLMs.
Managing technical debt with the sqale method,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
other 1
citation-polarity summary
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models
COMPASS benchmark shows that evaluating code generation with only correctness misses large differences in runtime efficiency and code quality among frontier LLMs.