FAUN-Eval is a curated benchmark of 300 real GitHub issue-PR pairs that scores LLMs separately on question answering, fault localization, and code editing.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models
FAUN-Eval is a curated benchmark of 300 real GitHub issue-PR pairs that scores LLMs separately on question answering, fault localization, and code editing.