A 328-problem execution-based benchmark shows the best AI code generators pass version-specific hidden tests only about half the time.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
A 328-problem execution-based benchmark shows the best AI code generators pass version-specific hidden tests only about half the time.