Code LLMs drop by more than 10% in accuracy when problem details are subtly changed, and fine-tuning on such counterfactual variants boosts performance on standard benchmarks.
If there are errors in the Sample Input/Output or in the Test Cases , correct them
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Success is in the Details: Evaluate and Enhance Details Sensitivity of Code LLMs through Counterfactuals
Code LLMs drop by more than 10% in accuracy when problem details are subtly changed, and fine-tuning on such counterfactual variants boosts performance on standard benchmarks.