Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-07T16:02:16.598404Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 26 inbound Pith citation observations for arXiv:2605.03546.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-07T16:02:16.598404Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:12:42.031152Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
21 of 21 outbound references displayed
External citation measurements
73
pith, observed 2026-08-05T02:28:24.338817Z
Observation 23491487-6b18-45e6-a9a2-a55a076fc98e · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2732c2d3-dbcb-43b6-8b21-67d4a240221f · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-chat: Coding Agent Interactions From Real Users in the Wild
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c658db2a-4b2d-408f-ad8c-9d960b5f59cb · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? Agentless: Demystifying LLM-based Software Engineering Agents
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3ebf38f3-0b04-491a-b912-460c64803631 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da61c64a-fd80-4b54-b3b4-14ed41c5919d · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? Effective Strategies for Asynchronous Software Engineering Agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5d2872e8-76fe-4b6e-ab67-5d80d71f2491 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? Measuring Coding Challenge Competence With APPS
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3f2c7c33-e23f-4fa7-b197-dbbfec9f3567 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? EffiBench: Benchmarking the Efficiency of Automatically Generated Code
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b5dc6ee8-1492-41db-b8e1-c2609ac3cc8b · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74e1376c-a74e-4f76-b723-c28bd2414375 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? doi: 10.1109/TSE.2015.2454513.https://doi.org/10
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32612799-a066-49b0-b250-2d8d92b0940e · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3ee170c3-ed9c-4fd4-be20-f3831f7bf6f3 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? KernelBench: Can LLMs Write Efficient GPU Kernels?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3ebec05b-a78d-4660-b4c0-8907d37d49a4 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? K., Krupke, D., Kidger, P., Sajed, T., Stellato, B., Park, J., et al
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06e4e3f1-2dae-412d-960e-490812c3f472 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? 14 Manish Shetty, Naman Jain, Jinjian Liu, Vijay Kethanaboyina, Koushik Sen, and Ion Stoica
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b74c08e-5aca-4489-9732-84c2bc29875e · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 96b3cbc3-d577-4412-87c4-f28dba077f05 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 521c5a30-4917-4235-8906-c77e34d7aa29 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bf80bdb8-7412-4e37-9977-190650dea79a · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa33546a-0a4e-435e-888e-ddbcc12cc5e6 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-smith: Scaling Data for Software Engineering Agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c3fc952-9b6d-4c25-a771-e3cc3a85fb49 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e3bf5a8-11fb-41de-816a-6d8e5049f672 · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? blocklist
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a1e91875-76e0-439e-9b01-84a3f117e77d · outbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? Dependency count.Of the 200 repositories, 171 (85.5%) contain a recognized package manifest file; among these, the median repository declares 17 total dependencies (12 runtime)
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cb4518c9-e379-45d2-b783-58f0f704ad54 · inbound
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8877253a-ba49-450b-be39-ee7688061346 · inbound
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aea32eb-cab9-4419-8c73-3f45ced95e0a · inbound
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 01073e1b-ede1-4856-ac7f-9a4a6c3f8804 · inbound
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 444b9378-2a23-4aac-8961-0bc25fd993cf · inbound
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 562e2e17-d6bf-42d4-9af3-7f3aa429ca7d · inbound
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2385a71b-7635-43c1-8e9c-3787a865c67f · inbound
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End? ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f9b5b92d-69a0-4255-a079-67aa72182e16 · inbound
PhoneBuddy: Training Open Models for Agentic Phone Use ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7cb494e5-ef3d-49d3-82d3-6e13fd75130f · inbound
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b24001df-3fae-46eb-8a3e-42837d9dabad · inbound
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ab37f4-185c-4fcb-ade7-06407dd454bf · inbound
MirrorCode: AI can rebuild entire programs from behavior alone ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b417fda7-9a20-4bd3-8b91-112047e66d5a · inbound
MirrorCode: AI can rebuild entire programs from behavior alone ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80ecb84a-28a8-4e1e-b014-8024f51ce7ee · inbound
Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3e2e693c-94df-46e4-a847-871633a1880f · inbound
Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd1e738-f856-47c2-9f7a-0a8fbf2aeef7 · inbound
ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d33732-dc73-463e-8f34-8923d7b0426f · inbound
GameEngineBench: Evaluating Coding Agents on Real C++ Runtime Environments ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3b249d-c456-4308-b02c-24b056bc14a8 · inbound
GameEngineBench: Evaluating Coding Agents on Real C++ Runtime Environments ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8663cfd-75fb-42d0-8bda-d2b42f5110ea · inbound
ArchEval: Measuring AI Agents as Computer Architects ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83dd5da8-a8d8-400f-84f0-19ca19df90eb · inbound
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0e8a81-bdd9-415e-8be6-07d311a8254b · inbound
Large Language Models Have Unreliable Understanding of Software Engineering Terminology ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b83b7f6f-68a8-4d7d-afd2-bda795aafdd3 · inbound
PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f92f42de-b44c-4560-8ff4-913c39d620e5 · inbound
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57c807f-6848-47f9-9cee-72b0b565b2f6 · inbound
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793a67f3-9468-431b-9c3c-d506b09ad2d2 · inbound
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fbf6837-852b-425b-9cb2-ebdd3652d86a · inbound
SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd783a3-029e-4a27-9fc8-711f071249e1 · inbound
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution ProgramBench: Can Language Models Rebuild Programs From Scratch?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.