SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation

Kaicheng Sun; Keke Han; Ruikai Shi; Rui Xie; Shikun Zhang; Wei Ye; Yidong Wang; Yixin Li; Zhengran Zeng; Zhuohao Yu

arxiv: 2509.01494 · v2 · pith:N75JPUYRnew · submitted 2025-09-01 · 💻 cs.SE

SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment Generation

Zhengran Zeng , Ruikai Shi , Keke Han , Yixin Li , Kaicheng Sun , Yidong Wang , Zhuohao Yu , Rui Xie

show 2 more authors

Wei Ye Shikun Zhang

This is my paper

classification 💻 cs.SE

keywords evaluationswrbenchcodecurrentreviewbenchmarkbenchmarkscontext

0 comments

read the original abstract

Automated Code Review (ACR) is crucial for software quality, yet existing benchmarks often fail to reflect real-world complexities, hindering the evaluation of modern Large Language Models (LLMs). Current benchmarks frequently focus on fine-grained code units, lack complete project context, and use inadequate evaluation metrics. To address these limitations, we introduce SWRBench , a new benchmark comprising 1000 manually verified Pull Requests (PRs) from GitHub, offering PR-centric review with full project context. SWRBench employs an objective LLM-based evaluation method that aligns strongly with human judgment (~90 agreement) by verifying if issues from a structured ground truth are covered in generated reviews. Our systematic evaluation of mainstream ACR tools and LLMs on SWRBench reveals that current systems underperform, and ACR tools are more adept at detecting functional errors. Subsequently, we propose and validate a simple multi-review aggregation strategy that significantly boosts ACR performance, increasing F1 scores by up to 43.67%. Our contributions include the SWRBench benchmark, its objective evaluation method, a comprehensive study of current ACR capabilities, and an effective enhancement approach, offering valuable insights for advancing ACR research.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Code Review Agent Benchmark
cs.SE 2026-03 unverdicted novelty 7.0

c-CRAB benchmark shows state-of-the-art code review agents solve only around 40% of tasks derived from human reviews, suggesting potential for human-AI collaboration.
On the Footprints of Reviewer Bots Feedback on Agentic Pull Requests in OSS GitHub Repositories
cs.SE 2026-04 unverdicted novelty 6.0

Reviewer bots' higher comment volume on AI agent PRs is associated with slower resolutions and poorer average feedback quality, while feedback quality itself has no association with PR outcomes.
Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review
cs.SE 2026-05 unverdicted novelty 5.0

The paper presents a vision for an agentic code review framework spanning PR Creation, Augmentation, Reviewer Selection, AI-Assisted Review, and Retrospective, with humans retained at quality gates.