Semantically similar prompt mutations cause substantial score shifts and can overturn model rankings in code benchmarks, especially within the same model family.
Code llama: Open foundation models for code, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
Semantically similar prompt mutations cause substantial score shifts and can overturn model rankings in code benchmarks, especially within the same model family.