Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:08:50.185042Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2507.09063.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:08:50.185042Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T00:25:09.639122Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:59:37.262688Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 50d5f51c-86b7-461d-bbf0-380d477824ac · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments CodeMirage: Hallucinations in Code Generated by Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9297bddf-9034-42d8-a918-4095eea14f66 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Aider code editing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6da50e90-6837-4278-b2a5-d5a6b69f218d · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Introducing claude 4
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3890ee47-51a6-443c-b208-cb69d2cba843 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Meet devin, the first ai software engineer
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d879eca9-43f8-45dd-88cc-d681e778c0de · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Envbench: A benchmark for automated environment setup
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e48daeb8-4b10-4042-91ba-34afbfcdc666 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Code completions with github copilot
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 033ba43f-9853-492f-a666-53e636a432df · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Meet the new github copilot coding agent
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e125de41-40e0-4d9b-ac1e-2b951395e8a3 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments S table T ool B ench: Towards stable large-scale benchmarking on tool learning of large language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37340459-c733-4faf-be3e-941b902b7aa6 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64aadbcc-7bd0-4e5e-b9fd-a6a96feeaaf6 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 189efb8e-726d-4d8c-a012-748ab8cbcddd · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments LADs: Leveraging LLMs for AI-Driven DevOps
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e45cd78-68a0-416c-b7a8-3838f9ccc14d · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef28d23c-67e6-42d2-ae3b-65b2ce6c1acb · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b86a78-b459-40bb-ae34-1b76186d72e9 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Long-context LLMs Struggle with Long In-context Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172e390e-a508-4fed-81b8-56cdc56880a2 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b22a4f2a-f6e4-41ac-ba8b-904495d780e1 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f1c29eb-498d-45fc-8586-90fdbe2c509a · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Agent B ench: Evaluating LLM s as agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 465e73db-8c85-484f-bde7-400bdb029460 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d4ff97-6a2b-4a42-9c02-08faba3b1908 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Beyond pip install : Evaluating llm agents for the automated installation of python projects
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 873b81fd-f42a-4b8f-b0ec-65c6c81668f8 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1f09481-964e-487a-a003-74c3ccfc8592 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Mitigating Configuration Differences Between Development and Production Environments: A Catalog of Strategies
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a745daca-bc41-4f02-b73e-483b5074a561 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Identifying Factors Contributing to Bad Days for Software Developers: A Mixed Methods Study
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f59e4bb-8f10-4883-9797-937ad07ba375 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Introducing codex
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c10f197-0d49-4891-b04d-e5f550e034ff · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Tool LLM : Facilitating large language models to master 16000+ real-world API s
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14cdc954-7a61-4ff9-ae2b-9710a708ce93 · outbound
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Retrieval models aren't tool-savvy: Benchmarking tool retrieval for large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ffbff7db-862b-4a39-a151-2dd1eed36654 · inbound
BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a46cff9f-fb2b-450e-a320-c89f9805d8d8 · inbound
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 019662f6-356d-4c7d-9f4c-df963549c716 · inbound
Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d0300806-5c46-477b-aae5-08da90ed168a · inbound
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5abfd60-86ad-4964-a30f-6ce9e6e87fca · inbound
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.