{"work":{"id":"a561c78a-4b02-4053-a92a-bc5c7c5f6b9b","openalex_id":"https://openalex.org/W4415252687","doi":"10.48550/arxiv.2509.16941","arxiv_id":"2509.16941","raw_key":null,"title":"SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?","authors":null,"authors_text":"Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg","year":2025,"venue":"cs.SE","abstract":"We introduce SWE-Bench Pro, a substantially more challenging benchmark that builds upon the best practices of SWE-BENCH [25], but is explicitly designed to capture realistic, complex, enterprise-level problems beyond the scope of SWE-BENCH. SWE-BENCH PRO contains 1,865 problems sourced from a diverse set of 41 actively maintained repositories spanning business applications, B2B services, and developer tools. The benchmark is partitioned into a public set with open access to problems sourced from 11 repositories, a held-out set of 12 repositories and a commercial set of 18 proprietary repositories where we have formal partnership agreements with early-stage startups. Problems in the held-out and the commercial set are not publicly accessible, but we release results on the commercial set. Our benchmark features long-horizon tasks that may require hours to days for a professional software engineer to complete, often involving patches across multiple files and substantial code modifications. All tasks are human-verified and augmented with sufficient context to ensure resolvability. To better understand these limitations, we cluster the failure modes observed in the collected agent trajectories for a clearer characterization of the error patterns exhibited by current models. Overall, SWE-BENCH PRO provides a contamination-resistant testbed that more faithfully captures the complexity and diversity of real-world software development, advancing the pursuit of truly autonomous software engineering agents at a professional level.","external_url":"https://arxiv.org/abs/2509.16941","cited_by_count":0,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2509.16941","created_at":"2026-05-09T06:25:39.762456+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?","render_title":"SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?"},"hub":{"state":{"work_id":"a561c78a-4b02-4053-a92a-bc5c7c5f6b9b","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":95,"external_cited_by_count":0,"distinct_field_count":9,"first_pith_cited_at":"2025-12-20T19:08:15+00:00","last_pith_cited_at":"2026-07-08T21:45:34+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T23:09:22.927220+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":10},{"context_role":"dataset","n":6},{"context_role":"baseline","n":1}],"polarity_counts":[{"context_polarity":"background","n":13},{"context_polarity":"use_dataset","n":3},{"context_polarity":"baseline","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}