SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Jialiang Liang; Kai Cai; Maoquan Wang; Ningyuan Xu; Shaoqiu Zhang; Shilin He; Siyu Ye; Wenhao Zeng; Xiaodong Gu; Yuhang Wang

arxiv: 2606.07297 · v1 · pith:27HZ6HPCnew · submitted 2026-06-05 · 💻 cs.SE · cs.CL

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Shaoqiu Zhang , Yuhang Wang , Jialiang Liang , Yuling Shi , Wenhao Zeng , Maoquan Wang , Shilin He , Ningyuan Xu

show 3 more authors

Siyu Ye Kai Cai Xiaodong Gu

This is my paper

classification 💻 cs.SE cs.CL

keywords codingagentsswe-explorecoderepositoryretrievalacrossagent

0 comments

read the original abstract

Repository-level coding benchmarks such as SWE-bench have driven a rapid surge in the capabilities of coding agents. Yet they usually treat coding tasks as a holistic, binary prediction problem (e.g., resolved or unresolved), neglecting fine-grained agent capabilities such as repository understanding, context retrieval, code localization, and bug diagnosis. In this paper, we introduce SWE-Explore, a benchmark that isolates the evaluation of repository exploration, a critical capability of coding agents. Given a repository and an issue, SWE-Explore asks an explorer to return a ranked list of relevant code regions under a fixed line budget. SWE-Explore covers 848 issues across 10 programming languages and 203 open-source repositories. For each instance, we derive line-level ground truth from independent agent trajectories that successfully solved the same issue, distilling the specific code regions their solution paths actually consulted. We evaluate exploration along coverage, ranking, and context-efficiency dimensions, showing that these metrics strongly track downstream repair behavior. Across a broad set of retrieval methods, general coding agents, and specialized localizers, we find that agentic explorers form a clear tier above classical retrieval. While file-level localization is already strong for modern methods, line-level coverage and efficient ranking remain the key axes differentiating state-of-the-art explorers.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

LLM Agents Can See Code Repositories
cs.SE 2026-06 unverdicted novelty 7.0

Visual graphs of repository structure added to text inputs for multimodal LLM agents reduce token consumption by up to 26% while maintaining or improving issue-resolution accuracy.
Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent
cs.AI 2026-06 unverdicted novelty 4.0

Ablation study finds that a structural codebase index improves localization and resolve rates in coding agents on two SWE benchmarks without raising per-cell cost.