Pith. sign in

Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies,

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.SE 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Are Large Language Models Memorizing Bug Benchmarks?

cs.SE · 2024-11-20 · conditional · novelty 5.0

Some base LLMs, particularly codegen-multi, show strong memorization of the Defects4J bug benchmark, while newer models like LLaMa 3.1 show weaker leakage signals.

citing papers explorer

Showing 1 of 1 citing paper.

  • Are Large Language Models Memorizing Bug Benchmarks? cs.SE · 2024-11-20 · conditional · none · ref 5

    Some base LLMs, particularly codegen-multi, show strong memorization of the Defects4J bug benchmark, while newer models like LLaMa 3.1 show weaker leakage signals.