{"work":{"id":"a0e2ddd0-ff8a-46d5-a2d8-234dae74cc22","openalex_id":null,"doi":null,"arxiv_id":"2410.06992","raw_key":null,"title":"SWE-Bench+: Enhanced Coding Benchmark for LLMs","authors":null,"authors_text":"Reem Aleithan, Haoran Xue, Mohammad Mahdi Mohajer, Elijah Nnorom, Gias Uddin, and Song Wang","year":2024,"venue":"cs.SE","abstract":"Large Language Models (LLMs) in Software Engineering (SE) can offer assistance for coding. To facilitate a rigorous evaluation of LLMs in practical coding contexts, Carlos et al. introduced the SWE-bench dataset, which comprises 2,294 real-world GitHub issues and their corresponding pull requests, collected from 12 widely used Python repositories. Several impressive LLM-based toolkits recently are developed and evaluated on this dataset. However, a systematic evaluation of the quality of SWE-bench remains missing. In this paper, we addressed this gap by presenting an empirical analysis of the SWE-bench dataset. We conducted a manual screening of instances where SWEAgent + GPT-4 successfully resolved issues by comparing the model-generated patches with the actual pull requests. SWE-Agent+GPT-4 was at the top of SWE-bench leaderboard during the time of our study. Our analysis reveals some critical issues with the SWE-bench dataset: 1) 32.67% of the successful patches involve cheating as the solutions were directly provided in the issue report or the comments. We refer to as solution leakage problem. 2) 31.08% of the passed patches are suspicious patches due to weak test cases, i.e., the tests were not adequate to verify the correctness of a patch. When we filtered out these problematic issues, the resolution rate of SWE-Agent+GPT-4 dropped from 12.47% to 3.97%. We also observed that the same data quality issues also exist in the two variants of SWE-bench, i.e., SWE-bench Lite and SWE-Bench Verified. In addition, over 94% of the issues were created before LLM's knowledge cutoff dates, posing potential data leakage issues.","external_url":"https://arxiv.org/abs/2410.06992","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-09T06:06:01.699658+00:00","pith_arxiv_id":"2410.06992","created_at":"2026-05-09T06:25:39.784205+00:00","updated_at":"2026-07-09T06:06:01.699658+00:00","title_quality_ok":true,"display_title":"ArXiv, abs/2410.06992","render_title":"ArXiv, abs/2410.06992"},"hub":{"state":{"work_id":"a0e2ddd0-ff8a-46d5-a2d8-234dae74cc22","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":23,"external_cited_by_count":null,"distinct_field_count":5,"first_pith_cited_at":"2025-03-20T17:59:23+00:00","last_pith_cited_at":"2026-07-08T16:13:15+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-13T11:59:42.431574+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":4}],"polarity_counts":[{"context_polarity":"background","n":4}],"runs":{},"summary":{},"graph":{},"authors":[]}}