Some base LLMs, particularly codegen-multi, show strong memorization of the Defects4J bug benchmark, while newer models like LLaMa 3.1 show weaker leakage signals.
A survey on software fault localization,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Are Large Language Models Memorizing Bug Benchmarks?
Some base LLMs, particularly codegen-multi, show strong memorization of the Defects4J bug benchmark, while newer models like LLaMa 3.1 show weaker leakage signals.