Pith. sign in

ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next. Yet existing software engineering (SWE) benchmarks evaluate models one bug at a time: the repository is reset, the codebase is re-read, and a single self-contained issue is graded in isolation. This setting collapses a continuous maintenance workflow into a series of independent sessions, ignoring the cumulative dependencies that make real-world bug fixing challenging. To bridge this gap, we introduce ChainSWE, the first benchmark for evaluating agents on sequential, dependent bug fixes within a shared codebase. We collect chronological chains of 304 issues across 54 Python projects, mined from six SWE-bench-family datasets. Our evaluation across a range of agents and models reveals a consistent performance drop by up to 70% as the chain length increases.

citation-role summary

dataset 1

citation-polarity summary

fields

cs.SE 1

years

2026 1

verdicts

CONDITIONAL 1

roles

dataset 1

polarities

use dataset 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.