GitGoodBench introduces a benchmark with 900 evaluation, 120 lite, and 17,469 training samples for three Git scenarios, with a 21.11% GPT-4o baseline solve rate on lite.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
GitGoodBench introduces a benchmark with 900 evaluation, 120 lite, and 17,469 training samples for three Git scenarios, with a 21.11% GPT-4o baseline solve rate on lite.