Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T13:48:53.691192Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 100 inbound Pith citation observations for arXiv:2509.16941.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T13:48:53.691192Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:17:17.509512Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
24 of 24 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation b58572c5-3108-493a-9244-811664d4d9a7 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-Bench+: Enhanced Coding Benchmark for LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4fa22eb7-74dd-4d6f-896e-412f97d37a18 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2730ce7e-26c4-4526-ad71-7659dfc36c60 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Brown, B
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d00b35bf-a19e-4727-80c4-519562a10538 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86b4d980-55e6-415f-8f05-d037e9fda9a3 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? A Survey on Data Contamination for Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 781e070c-9bf1-4920-951e-e58502d4508d · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 547ba755-0716-4334-9fd2-073fe81acbb7 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f4ac406-f073-49d6-b2f8-711ac0c9b261 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd7f8f2a-2500-45cf-bf98-c977351137b0 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Hendrycks, S
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61dfd1a7-6b83-4089-92fd-84cf07c1a199 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 41215c8a-1442-4c82-8315-342ee5a72f8f · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? URLhttps://openai.com/index/introducing-swe-bench-verified/
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a3f1750e-fb59-499c-b510-21b55d3c962b · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Steidl, B
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3a1428d-2403-40a5-b69c-8fc4fd377703 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 947d9844-7be4-42d4-9912-6e649fc11685 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Agentless: Demystifying LLM-based Software Engineering Agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50e60ca1-3688-455e-a3a9-d82c29367e6e · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Benchmark Data Contamination of Large Language Models: A Survey
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb4a0324-6647-4cc5-a03b-c60c208bf59d · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c816a82b-f825-466f-9d70-8f6afb83e9bc · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f0b4741-ac2d-4463-a347-5a80fe12f348 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? A Gauss-Seidel method for solving multi-leader-multi-follower games
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a9f94058-bd67-404e-a148-0ea32a88f3b1 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-bench Goes Live!
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69e784c3-d604-4142-8de4-1b074e2caeaa · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8e8c8fb-7d89-4c65-993b-71d55550c1a1 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Book 978
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f712ad7-fe0a-4caa-a996-1993c4edd292 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? iOS Contacts)
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e24d14d4-c7ef-4ac7-97fb-0d7918ec8ef3 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47a54694-2997-4dba-94c4-cf6eaad0c878 · outbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? No actions recorded
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 13f16d5c-f481-4a6f-a3d1-39ca7c7ba0f7 · inbound
Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7e9f8cb-82d3-4683-8574-95f7359de454 · inbound
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f06ad605-bd3d-4005-a720-f5c4fbdb8310 · inbound
Toward Training Superintelligent Software Agents through Self-Play SWE-RL SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3173a31c-9f42-4acb-a605-f8e95113e822 · inbound
Toward Training Superintelligent Software Agents through Self-Play SWE-RL SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b595c520-873a-4502-89d0-04b38fa983e0 · inbound
Token-Level LLM Collaboration via FusionRoute SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ce49000-131b-4fba-af4c-00b792727133 · inbound
Kimi K2.5: Visual Agentic Intelligence SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4d4d8e91-4238-43a2-bc70-f90979d13cbc · inbound
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efab4aec-579f-4fb9-8c2b-1a25fe507601 · inbound
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72356ece-1253-4191-8d6d-7e65b1af3de5 · inbound
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 857d370e-323e-4cbd-9d14-fc454972b0e1 · inbound
SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c877ba6d-c61a-4fee-9e42-e86dba31844d · inbound
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11f9be90-3195-48cf-9e38-bc5bf5c83ee4 · inbound
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac9ac8a3-bcd2-4802-82da-6601d83b922a · inbound
SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962a08fc-1c02-42ef-9fb9-073405967f2e · inbound
Effective Strategies for Asynchronous Software Engineering Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca76b21-4b78-47cd-b3b0-9599ad8d659b · inbound
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a319e9c7-0805-4589-ac66-9ef7b6cdb150 · inbound
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f2fbd2-f229-4841-b32c-c2e956e3fbb9 · inbound
Beyond Resolution Rates: Behavioral Drivers of Coding Agent Success and Failure SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5fe1d313-cf97-414c-9b64-77cc0c915ceb · inbound
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63be1b4f-e3d6-4aba-8384-0395a19e2e57 · inbound
Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 58ff9215-614e-4895-9c28-56911a4cc6cf · inbound
Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03ad9450-e42e-466d-8228-29516e394f06 · inbound
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b677eb3d-a0bf-4cce-a88f-80d918952e7e · inbound
REAgent: Requirement-Driven LLM Agents for Software Issue Resolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6077b97-22f5-4326-a9c8-a3fd6521f465 · inbound
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eab2639-01f2-41c0-9071-b72e9123a2fd · inbound
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a5df470-1c62-4ded-8643-a3592547b717 · inbound
Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f0d715d-d53d-4fd3-8e5d-d106104f7ea0 · inbound
Evaluating Plan Compliance in Autonomous Programming Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd6c239f-0d6f-4699-bbdc-5a86a72e6a27 · inbound
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4af8321-575e-4797-8840-69e3b113f30e · inbound
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8bd41b9-d737-4592-9b4c-5503b1b9a1b9 · inbound
What Should Frontier AI Developers Disclose About Internal Deployments? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69b70e61-fb90-4039-96e5-803f5cd20d3f · inbound
KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d410cc8c-814f-4668-841f-cd6609322f23 · inbound
KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9cd76b4c-6bd2-4e90-bdea-b5b53fe87986 · inbound
Risk Reporting for Developers' Internal AI Model Use SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d381dd1-bc04-424a-b780-bf6687025065 · inbound
From Threads to Trajectories: A Multi-LLM Pipeline for Community Knowledge Extraction from GitHub Issue Discussions SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5dcb735e-0ec5-4e3c-b9af-e2542d42a03d · inbound
An Empirical Study of Speculative Decoding on Software Engineering Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a413685-774a-4426-aa15-6a61bee37eb5 · inbound
The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5226126a-18da-4d8e-827e-b9eaa1bf7171 · inbound
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 95b77cf5-352a-4122-95bd-7aadd0001c70 · inbound
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ebf38f3-0b04-491a-b912-460c64803631 · inbound
ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2252a6bc-f4b8-4c1d-8e49-8054c3554788 · inbound
Reproduction Test Generation for Java SWE Issues SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f5e6106c-2853-4b7c-8a7c-7c119cba2f7c · inbound
Reproduction Test Generation for Java SWE Issues SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5097a72-fd71-451d-b850-25b82c93f511 · inbound
Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f7cafb1d-2016-4127-9ca4-ea4caad9df08 · inbound
Constraint Decay: The Fragility of LLM Agents in Backend Code Generation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54fcd544-adc8-4f04-942c-471ea125fde7 · inbound
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 77150a33-e2b1-46f7-813a-1d98e59558ff · inbound
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a504efaa-d9c0-4462-a00c-eda78f883fed · inbound
LLM Agents Already Know When to Call Tools -- Even Without Reasoning SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8ef8e8c-9e34-4861-b2a7-a19be5da0d1f · inbound
LLM Agents Already Know When to Call Tools -- Even Without Reasoning SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a84ee52e-36b6-4a4d-be4e-32d46abc0950 · inbound
PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d83f093e-718b-40f0-bcea-3fed8c8fd736 · inbound
The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb042689-a284-479a-8f12-f6a3a55d4fd3 · inbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c8ade44b-5959-4279-a90b-bffac6b7c730 · inbound
Revisiting DAgger in the Era of LLM-Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06d7c0a3-9c97-4644-b93b-5a6fc3091519 · inbound
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b59972cb-0de7-4587-8e8a-2fac2b9f984d · inbound
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ab82e84-cc86-4718-b255-fdc308fd9099 · inbound
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d198de4-fee1-4182-9719-31488dd2f51d · inbound
SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9a80e3eb-b345-4428-ab23-02f0038ead57 · inbound
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7bc1a6d-a13d-4e66-8202-8bc9bd28b69f · inbound
AI for Auto-Research: Roadmap & User Guide SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 311134f0-7ca3-4a19-90ff-6b762e1931c3 · inbound
AI for Auto-Research: Roadmap & User Guide SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e16258-0172-4c27-9d02-66cab342ed7b · inbound
Open-World Evaluations for Measuring Frontier AI Capabilities SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f26fd98-371a-42f8-9caf-d8ac484d49e4 · inbound
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bee8a10-af87-431e-842f-9cf15f3f0aa1 · inbound
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b89c8d84-bfd6-4425-9fde-090fade0186f · inbound
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a92ef363-b45f-4990-b02c-2b60d60f0cff · inbound
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f0114c10-1b66-4e8e-8580-10434f5800a5 · inbound
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90a487ad-326a-4460-8376-7407687d6d3c · inbound
Mem-$\pi$: Adaptive Memory through Learning When and What to Generate SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c10373ab-a14e-46ad-9db8-3d85178d6f65 · inbound
From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05de791f-d537-42cb-b540-d2f9f442d50c · inbound
Design and Report Benchmarks for Knowledge Work SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5985207-77b9-45da-80a5-780d2bacb82c · inbound
Stop Comparing LLM Agents Without Disclosing the Harness SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc23a8af-ad45-4bc1-8394-9afb0e7626ed · inbound
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3d95fc5b-f528-4013-b57c-824d67e1275a · inbound
Agentic AI Workload Characteristics SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 74362d7c-e805-4960-93e9-4499ec8a3d65 · inbound
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6cd2488a-f694-4ced-898c-f6f4d4c8cbf8 · inbound
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6315f07-7902-4aa9-8442-d029cbd258d9 · inbound
TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46ca2011-88b9-40a1-9473-b1e2431d08b9 · inbound
TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d2dc92-3a74-47c3-b03e-2a31eefbb616 · inbound
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51a6c18d-b57e-404d-ab47-1d6f4aac611c · inbound
PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 037da2a1-4226-4704-b11d-4c336c4f9fc1 · inbound
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 46bad182-e56a-4676-ad27-cfd6422c3c78 · inbound
BlueFin: Benchmarking LLM Agents on Financial Spreadsheets SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bceb751-5100-452e-a60a-4ac3614bccdb · inbound
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c6ac21d-2210-41c6-a001-7b6bae2bc7d9 · inbound
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc911c64-2378-4cb3-8db6-81ffbc66ac0f · inbound
TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe56401d-1c7d-4c4b-a87f-6c00174200d9 · inbound
Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ebdbb5be-f447-4acd-ab0e-212b6ac70a40 · inbound
Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7f14fcfe-164d-46dc-a39a-0f60100073ef · inbound
SWE-Explore: Benchmarking How Coding Agents Explore Repositories SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0656c1f3-66af-4007-a018-7be7e408ce6f · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 151
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54d18632-e300-4c7e-a44d-55bdc2d6c79f · inbound
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eeb84929-671d-412e-88bc-453e6662a8e7 · inbound
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b428d440-dbb6-4f23-94ef-820c10d84236 · inbound
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4de31687-758a-479c-8de5-17dd8a57e262 · inbound
Dissecting model behavior through agent trajectories SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 08baac0d-7a43-4972-a0ba-a83fc0b01753 · inbound
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a91eb18-0a35-4798-8dd1-c3ad097ae358 · inbound
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a4a2b6b-4888-46f3-913f-78879420a5a9 · inbound
A Framework for Evaluating Agentic Skills at Scale SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc0d4a20-cd2a-42be-8c3a-c3dd72d3102b · inbound
Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90e7165b-426d-408f-9739-52eba309fccd · inbound
Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68aedaf1-b5a4-4b14-a4ef-e54b0d753528 · inbound
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97151aee-e408-496f-acda-bd857ebc1ccf · inbound
Sakana Fugu Technical Report SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bd2557e-bdf9-4d9c-935d-b65da3acb0d0 · inbound
SwarmX: Agentic Scheduling for Low-Latency Agentic Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 750246b3-5d42-40d2-b85e-48eda54ea020 · inbound
SwarmX: Agentic Scheduling for Low-Latency Agentic Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1f62e3e-8627-456f-bd3e-a213102da0cd · inbound
CFAgentBench: A Reproducible Environment and Benchmark for Autonomous Construction-Finance Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 91914b4e-2ab9-4738-a16c-95041df088df · inbound
Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5fef82c6-8b56-48dc-8757-99ded1855f77 · inbound
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.