Pith. sign in

super hub Canonical reference

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Canonical reference. 71% of citing Pith papers cite this work as background.

139 Pith papers citing it
2 external citations · Pith
Background 71% of classified citations
abstract

We introduce SWE-Bench Pro, a substantially more challenging benchmark that builds upon the best practices of SWE-BENCH [25], but is explicitly designed to capture realistic, complex, enterprise-level problems beyond the scope of SWE-BENCH. SWE-BENCH PRO contains 1,865 problems sourced from a diverse set of 41 actively maintained repositories spanning business applications, B2B services, and developer tools. The benchmark is partitioned into a public set with open access to problems sourced from 11 repositories, a held-out set of 12 repositories and a commercial set of 18 proprietary repositories where we have formal partnership agreements with early-stage startups. Problems in the held-out and the commercial set are not publicly accessible, but we release results on the commercial set. Our benchmark features long-horizon tasks that may require hours to days for a professional software engineer to complete, often involving patches across multiple files and substantial code modifications. All tasks are human-verified and augmented with sufficient context to ensure resolvability. To better understand these limitations, we cluster the failure modes observed in the collected agent trajectories for a clearer characterization of the error patterns exhibited by current models. Overall, SWE-BENCH PRO provides a contamination-resistant testbed that more faithfully captures the complexity and diversity of real-world software development, advancing the pursuit of truly autonomous software engineering agents at a professional level.

hub tools

citation-role summary

background 16 dataset 7 baseline 1

citation-polarity summary

representative citing papers

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

cs.CR · 2026-08-26 · conditional · novelty 7.0

A self-evolving coding agent can be poisoned through its own tool-authoring step: reading a planted skill makes the agent author, store, and later run a malicious copy that persists even after the original skill is removed.

DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

cs.AI · 2026-08-21 · accept · novelty 7.0

A controlled multi-session benchmark shows that later coding tasks can depend on earlier-session memory, and that a simple verbatim event-memory baseline is surprisingly strong, while the benchmark reliably discriminates memory-bearing conditions from no memory.

Dockerless: Environment-Free Program Verifier for Coding Agents

cs.SE · 2026-06-26 · unverdicted · novelty 7.0

Dockerless uses agentic repository exploration to verify patches without execution, enabling SFT and RL training of coding agents that reach 62.0/50.0/35.2% resolve rates on SWE-bench Verified/Multilingual/Pro while matching environment-based results.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns

cs.SE · 2026-06-17 · unverdicted · novelty 7.0

StaminaBench evaluates coding agents over 100 procedurally generated change requests to a REST API, finding that tested models fail within 5-6 turns without feedback but improve up to 12x with test feedback and good harnesses.

citing papers explorer

Showing 50 of 139 citing papers.