BouncerBench evaluates whether LLM coding agents can abstain from acting on underspecified tickets and incorrect patches, and shows current models rarely abstain correctly.
Improving the effectiveness of peer code review in identifying security defects,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Is Your Automated Software Engineer Trustworthy?
BouncerBench evaluates whether LLM coding agents can abstain from acting on underspecified tickets and incorrect patches, and shows current models rarely abstain correctly.