Some base LLMs, particularly codegen-multi, show strong memorization of the Defects4J bug benchmark, while newer models like LLaMa 3.1 show weaker leakage signals.
Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Are Large Language Models Memorizing Bug Benchmarks?
Some base LLMs, particularly codegen-multi, show strong memorization of the Defects4J bug benchmark, while newer models like LLaMa 3.1 show weaker leakage signals.