A toolkit of 11 code refactoring operators reduces n-gram overlap with training corpora by up to 65 percentage points, though this drop is partly by construction and is not tied to downstream task performance.
NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CODECLEANER: Elevating Standards with A Robust Data Contamination Mitigation Toolkit
A toolkit of 11 code refactoring operators reduces n-gram overlap with training corpora by up to 65 percentage points, though this drop is partly by construction and is not tied to downstream task performance.