Pith. sign in

Paper Citation Record · LEDGER

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

As of 7 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2507.09063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09063 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:08:50.185042Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:25:09.639122Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:59:37.262688Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50d5f51c-86b7-461d-bbf0-380d477824ac · outbound

This paper cites CodeMirage: Hallucinations in Code Generated by Large Language Models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments CodeMirage: Hallucinations in Code Generated by Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:42.937431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:42.937431Z digest=sha256:47094453565b97cda9ac30ed43bfd58319909cf8ae956a1b3353242f55d1e8a2

Observation 9297bddf-9034-42d8-a918-4095eea14f66 · outbound

This paper cites Aider code editing.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Aider code editing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:53.176750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:43.101216Z digest=sha256:cc5c68320e055c528fea8d276d54fd3c956c52f828bb4af20c562c861e8e1919

Observation 6da50e90-6837-4278-b2a5-d5a6b69f218d · outbound

This paper cites Introducing claude 4.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Introducing claude 4

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:53.007049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:43.534078Z digest=sha256:8fb275fbbbc5abe179bb32ecf89173213d59a2cd4c0fb5116bb5624e7bccb6c6

Observation 3890ee47-51a6-443c-b208-cb69d2cba843 · outbound

This paper cites Meet devin, the first ai software engineer.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Meet devin, the first ai software engineer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.841914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:44.557141Z digest=sha256:de00aa6496018dd1eb6436fadb2e5c10128687da4469ebd44448c46ee3d4e252

Observation d879eca9-43f8-45dd-88cc-d681e778c0de · outbound

This paper cites Envbench: A benchmark for automated environment setup.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Envbench: A benchmark for automated environment setup

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.672129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:45.312916Z digest=sha256:53db22e936ab96dd6d571c3f20a9ec0995ee338ffdd6ac83f579b7f90e5b81ba

Observation e48daeb8-4b10-4042-91ba-34afbfcdc666 · outbound

This paper cites Code completions with github copilot.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Code completions with github copilot

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.561243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:46.040321Z digest=sha256:7ab13d364dfe13186370fbc2cb496bcfd114c2c243a7145bc737b4dc46885114

Observation 033ba43f-9853-492f-a666-53e636a432df · outbound

This paper cites Meet the new github copilot coding agent.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Meet the new github copilot coding agent

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.434166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:46.588514Z digest=sha256:c7ecf3d1708129ec9534b2d9cb0e159223bf04a966d63acb6b71e2fd9e19ff60

Observation e125de41-40e0-4d9b-ac1e-2b951395e8a3 · outbound

This paper cites S table T ool B ench: Towards stable large-scale benchmarking on tool learning of large language models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments S table T ool B ench: Towards stable large-scale benchmarking on tool learning of large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:46.856121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:46.856121Z digest=sha256:8a59c2b95932349f3b207966a4f64c5605088e54340f821a48d101ce1b543e07

Observation 37340459-c733-4faf-be3e-941b902b7aa6 · outbound

This paper cites Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:46.990851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:46.990851Z digest=sha256:9f8c3d83b2d500b748e611e870f0e8f61a6bedc0afce6e5023147f4dcbbd3700

Observation 64aadbcc-7bd0-4e5e-b9fd-a6a96feeaaf6 · outbound

This paper cites SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.207136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.207136Z digest=sha256:1c54d6b0d07ce487d7f5a993d50a3c2b3ea57dbc27b2fea6e91bb0ffaab45d0d

Observation 189efb8e-726d-4d8c-a012-748ab8cbcddd · outbound

This paper cites LADs: Leveraging LLMs for AI-Driven DevOps.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments LADs: Leveraging LLMs for AI-Driven DevOps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.452716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.452716Z digest=sha256:7ae2217cc13bd83ffb5f77bf1b809fd6674575ad96118ad39fced3351e513990

Observation 2e45cd78-68a0-416c-b7a8-3838f9ccc14d · outbound

This paper cites Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.747783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.747783Z digest=sha256:47cef4f0f70d44bf44a36424dfe4778013401f53439a17583bbbaef078eaeaed

Observation ef28d23c-67e6-42d2-ae3b-65b2ce6c1acb · outbound

This paper cites ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.901841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.901841Z digest=sha256:6fcc376f22b941d08f47083eab0c6cbb545f712a99a8e2c75e06ba34fba55f68

Observation a1b86a78-b459-40bb-ae34-1b76186d72e9 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Long-context LLMs Struggle with Long In-context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.218778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.218778Z digest=sha256:ba6f21e0b8d6f8d17d0c6b89a9b60c165b3b11ca087cc1fd1eb909ce2785eb93

Observation 172e390e-a508-4fed-81b8-56cdc56880a2 · outbound

This paper cites NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.471016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.471016Z digest=sha256:65592f778a8a31d7285d6f209e381589e291d3c1ff82d9a9743b7cf8f8552658

Observation b22a4f2a-f6e4-41ac-ba8b-904495d780e1 · outbound

This paper cites GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.894069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.894069Z digest=sha256:af6b8fd00ede79a6ac63fb763607fc442ec7ff10a9b4a60565c953beb372623b

Observation 1f1c29eb-498d-45fc-8586-90fdbe2c509a · outbound

This paper cites Agent B ench: Evaluating LLM s as agents.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Agent B ench: Evaluating LLM s as agents

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.185296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.409404Z digest=sha256:3e86fbb1def298903aa4d33640699f35fc7fa8cce3177a9601b4b9b5e53f54b0

Observation 465e73db-8c85-484f-bde7-400bdb029460 · outbound

This paper cites OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:49.556866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:49.556866Z digest=sha256:f86fa86bdc94631599b590ff1745079834d227cdb56c5931d2164942cd93b2c6

Observation f4d4ff97-6a2b-4a42-9c02-08faba3b1908 · outbound

This paper cites Beyond pip install : Evaluating llm agents for the automated installation of python projects.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Beyond pip install : Evaluating llm agents for the automated installation of python projects

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.886387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.657176Z digest=sha256:04fe80ead65c4d0e16b123ef851894fc3eaf549df783ccf1d75f5881793b5a07

Observation 873b81fd-f42a-4b8f-b0ec-65c6c81668f8 · outbound

This paper cites Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.730388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.760261Z digest=sha256:61f51e30dc11e4d850edb073011898e96721a591884a25050ddca34040be1f74

Observation d1f09481-964e-487a-a003-74c3ccfc8592 · outbound

This paper cites Mitigating Configuration Differences Between Development and Production Environments: A Catalog of Strategies.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Mitigating Configuration Differences Between Development and Production Environments: A Catalog of Strategies

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.560978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.838672Z digest=sha256:55822bdebbea5bb23d02be9a5ebfcc4efee814e768fe5e848895ef301fc0894d

Observation a745daca-bc41-4f02-b73e-483b5074a561 · outbound

This paper cites Identifying Factors Contributing to Bad Days for Software Developers: A Mixed Methods Study.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Identifying Factors Contributing to Bad Days for Software Developers: A Mixed Methods Study

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.393493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.949818Z digest=sha256:7a8d626f7277fde948d79f24882609e696c81d574bb22c96bd81f5ac4f116ac5

Observation 1f59e4bb-8f10-4883-9797-937ad07ba375 · outbound

This paper cites Introducing codex.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Introducing codex

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.532796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:50.015354Z digest=sha256:b5b7e85ba072320ff93b5c408efd45a9566a478e266c51a237a1c3e040bb4e06

Observation 6c10f197-0d49-4891-b04d-e5f550e034ff · outbound

This paper cites Tool LLM : Facilitating large language models to master 16000+ real-world API s.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Tool LLM : Facilitating large language models to master 16000+ real-world API s

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:50.078862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:50.078862Z digest=sha256:a32ed4bf08b43717ff94bab3e4d3418ad55280c4b3ea4fd73e1326f894fdf32c

Observation 14cdc954-7a61-4ff9-ae2b-9710a708ce93 · outbound

This paper cites Retrieval models aren't tool-savvy: Benchmarking tool retrieval for large language models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Retrieval models aren't tool-savvy: Benchmarking tool retrieval for large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.106233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:08:50.185042Z digest=sha256:9070c1b34fa62315434dd8ad3040bc1c7efb006fbabe57815bbfcc0e48ef15b3

Pith citing papers

Observation ffbff7db-862b-4a39-a151-2dd1eed36654 · inbound

BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge cites this paper.

BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:08:41.972644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T17:04:07.449733Z digest=sha256:f7552a9ebb6da8fd56e449f7009ac633f03cf5773cab0aa986d745252393f70b

Observation a46cff9f-fb2b-450e-a320-c89f9805d8d8 · inbound

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment cites this paper.

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:16:49.263539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T05:35:12.027697Z digest=sha256:c270f7f1354544e648d39de7d80382337ee8233e2fe47ee5cf90a0f827c5590c

Observation 019662f6-356d-4c7d-9f4c-df963549c716 · inbound

Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer cites this paper.

Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.264275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T14:02:37.404821Z digest=sha256:f3837ed48da21afe2b6dc00e2d14470fd55f4a534cfd1e6ec1770446501e7474

Observation d0300806-5c46-477b-aae5-08da90ed168a · inbound

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents cites this paper.

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:48.608177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T01:19:54.081755Z digest=sha256:26d6fcb2e2f0f9fe01ec047fd503b5164d4696607c6430f62196011fcc732439

Observation e5abfd60-86ad-4964-a30f-6ce9e6e87fca · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:09.639122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:09.639122Z digest=sha256:8d8c4942571c604855c2c8798bd16c9261481c88e7935f7d87cbf0b825bd05d8