Pith. sign in

Paper Citation Record · LEDGER

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

As of 6 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 100 inbound Pith citation observations for arXiv:2509.16941.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.16941 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T13:48:53.691192Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 100 of 153 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:17:17.509512Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact12
  • verified fuzzy7
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b58572c5-3108-493a-9244-811664d4d9a7 · outbound

This paper cites SWE-Bench+: Enhanced Coding Benchmark for LLMs.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-Bench+: Enhanced Coding Benchmark for LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.720917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:8d5ed3a8d99896cc328241d6f517502446569a4f61f1818aec86c44d918048c1

Observation 4fa22eb7-74dd-4d6f-896e-412f97d37a18 · outbound

This paper cites Program Synthesis with Large Language Models.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Program Synthesis with Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T13:48:53.728075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:4855bb541c39451f05c25b755054263dcf330fc1043c5480f928f0a9d28a031a

Observation 2730ce7e-26c4-4526-ad71-7659dfc36c60 · outbound

This paper cites Brown, B.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Brown, B

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:48:53.806530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:0dc49f0839d8577e385645f96b8a55b6ab95dccd1e7ecbdaf7b5aa5a4f09ad90

Observation d00b35bf-a19e-4727-80c4-519562a10538 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T13:48:53.734613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:09cfeb71199a307b8a19f2fbd182af4e09e858b80500d7fe0e4133fae658c182

Observation 86b4d980-55e6-415f-8f05-d037e9fda9a3 · outbound

This paper cites A Survey on Data Contamination for Large Language Models.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? A Survey on Data Contamination for Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.741016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:e39345f3aada29a8776077242521297895e13f224ed6d58cd4d7a7a0c37d8857

Observation 781e070c-9bf1-4920-951e-e58502d4508d · outbound

This paper cites Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.748215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:09b600a66c97b320fda298dd1a2352ab4f2bae25665777effdbacb2b80b5294a

Observation 547ba755-0716-4334-9fd2-073fe81acbb7 · outbound

This paper cites an unresolved cited work.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:48:53.821287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:1730a83c021bbebab6aeb3fdfe7b9b9695cc991e23d84305b3c1e530086a9b0f

Observation 0f4ac406-f073-49d6-b2f8-711ac0c9b261 · outbound

This paper cites an unresolved cited work.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:48:53.824830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:914d4a4e59612fae07110c71def02daa376308d9846599a02fc35e8fa1d711f5

Observation bd7f8f2a-2500-45cf-bf98-c977351137b0 · outbound

This paper cites Hendrycks, S.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Hendrycks, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:48:53.829641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:a139288c8cfbc51c6b768990d8130e40735fffc3550b7d3dbc75fd4bb126e67f

Observation 61dfd1a7-6b83-4089-92fd-84cf07c1a199 · outbound

This paper cites an unresolved cited work.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:48:53.837044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:4db86f9814641496328e1cb5e708333b8e7aeafceea37f25139e8d109cb645a4

Observation 41215c8a-1442-4c82-8315-342ee5a72f8f · outbound

This paper cites URLhttps://openai.com/index/introducing-swe-bench-verified/.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? URLhttps://openai.com/index/introducing-swe-bench-verified/

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:48:53.840782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:b2200da711f14d385e3d9adee38ad4f0c1ef3d59066b828ab786c84bfbd3c29f

Observation a3f1750e-fb59-499c-b510-21b55d3c962b · outbound

This paper cites Steidl, B.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Steidl, B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:48:53.847696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:e7426383d6f6aff744ee2e45cf10dec80bb602d8827775e5fbf7806882c488c6

Observation c3a1428d-2403-40a5-b69c-8fc4fd377703 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.603034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:43898bbb689dd7d30d25749827d8cc315acfecb1648911e51e78a3a75ee5518f

Observation 947d9844-7be4-42d4-9912-6e649fc11685 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Agentless: Demystifying LLM-based Software Engineering Agents

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T13:48:53.760722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:5123bb033014e450bf7c89f81f6fc40f06844b4a5487e5d14777db99637c3116

Observation 50e60ca1-3688-455e-a3a9-d82c29367e6e · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Benchmark Data Contamination of Large Language Models: A Survey

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.376209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:c788ce6c027cbf9936dd26d46e01bdd1b0a22bb1ab69611c2dd760bb18cd6138

Observation eb4a0324-6647-4cc5-a03b-c60c208bf59d · outbound

This paper cites an unresolved cited work.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:48:53.814455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:d68f0dc9c74dcff011d9d26b3d8a2629b5157381af9d0df62be3a4da09e248a4

Observation c816a82b-f825-466f-9d70-8f6afb83e9bc · outbound

This paper cites SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.774713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:d4bf812533cf00e59e195f8e053f108481a12a4e45ee0f188045ef4bda94ec67

Observation 3f0b4741-ac2d-4463-a347-5a80fe12f348 · outbound

This paper cites A Gauss-Seidel method for solving multi-leader-multi-follower games.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? A Gauss-Seidel method for solving multi-leader-multi-follower games

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.781577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:e68b52cc24f2ae7d248a4f648008a1517e03dbb05a30b12fab368865893cf082

Observation a9f94058-bd67-404e-a148-0ea32a88f3b1 · outbound

This paper cites SWE-bench Goes Live!.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-bench Goes Live!

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.788715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:6325b8ddad4266c14a0744e565d24131e01a31b02368005518e574cd5d214e65

Observation 69e784c3-d604-4142-8de4-1b074e2caeaa · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.795243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:841525a265a0e7c6710a308d21881252856741a240564ccee23f31057d7f6386

Observation a8e8c8fb-7d89-4c65-993b-71d55550c1a1 · outbound

This paper cites Book 978.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Book 978

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:48:53.818015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:7d46dc0480b10ac993f0bb5d7ed8a4bfb539076aa7de85db7190233808206290

Observation 0f712ad7-fe0a-4caa-a996-1993c4edd292 · outbound

This paper cites iOS Contacts).

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? iOS Contacts)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:48:53.799271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:e9a3fa843c93b50d06da6a81c5b6061da52a92c33385420e0554fefc7743f8d1

Observation e24d14d4-c7ef-4ac7-97fb-0d7918ec8ef3 · outbound

This paper cites an unresolved cited work.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:48:53.802767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:8a8e60fb2cb66734c5e373c41265c8d94620218653bbfde764147c7632894359

Observation 47a54694-2997-4dba-94c4-cf6eaad0c878 · outbound

This paper cites No actions recorded.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? No actions recorded

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:48:53.810753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:168709d74732ea1c01e965305476ea96d97e73edaa70bdb2108d2fe9e9ee73ca

Pith citing papers

Observation 13f16d5c-f481-4a6f-a3d1-39ca7c7ba0f7 · inbound

Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning cites this paper.

Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:17:17.509512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:17:17.509512Z digest=sha256:a12f2c1ab90b2b0ff91fd82cb1e09c426f06888ed7cdcd9ae9465376246b4586

Observation a7e9f8cb-82d3-4683-8574-95f7359de454 · inbound

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios cites this paper.

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:28:24.562895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T20:24:40.939455Z digest=sha256:9790564b1fe6bce0275c08c886c6b0365a2764cd4115423cf3049c20dd458d4f

Observation f06ad605-bd3d-4005-a720-f5c4fbdb8310 · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T16:10:20.250611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T16:07:48.570995Z digest=sha256:dfd3e463b6d1cab276aa2cd723f1399513393dca8d58968e2b7fb150da953530

Observation 3173a31c-9f42-4acb-a605-f8e95113e822 · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T15:02:10.639537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:02:10.639537Z digest=sha256:202fe8c2e44b107f8eceb6ec0ecc23b5263327abd1e485faf6974ab50e5a5ec9

Observation b595c520-873a-4502-89d0-04b38fa983e0 · inbound

Token-Level LLM Collaboration via FusionRoute cites this paper.

Token-Level LLM Collaboration via FusionRoute SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:26:31.643260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T12:25:59.747665Z digest=sha256:f352c08642c7176f97e024f1b6a3375ebd836aa02ec6f6033d8b43e8016a5a95

Observation 9ce49000-131b-4fba-af4c-00b792727133 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:b6f15e59ee787a1b0124c675160569cf56005d26b0a3aee0062a1540d2a112c9

Observation 4d4d8e91-4238-43a2-bc70-f90979d13cbc · inbound

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation cites this paper.

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:07:11.979895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T03:04:17.755968Z digest=sha256:0f95e631a7513f91575fc301e5868351ca29459f2cc58d4d68fd9e6f41930830

Observation efab4aec-579f-4fb9-8c2b-1a25fe507601 · inbound

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? cites this paper.

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:35.274592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:26:35.274592Z digest=sha256:f3ce8a0435aacf5547e05d93aee554747ea209db59737d62d29d4fc8d1e3a90d

Observation 72356ece-1253-4191-8d6d-7e65b1af3de5 · inbound

Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs? cites this paper.

Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.840575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:17:52.610119Z digest=sha256:c0503255fec08dc2a6951fabd720c1b5058ed2f8a36eed1caf109f72209be440

Observation 857d370e-323e-4cbd-9d14-fc454972b0e1 · inbound

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution cites this paper.

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T19:43:15.515948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:43:15.515948Z digest=sha256:75033814970d181c39e5bc89d31966dcf9e461b739cfc1248849036a1b1447c7

Observation c877ba6d-c61a-4fee-9e42-e86dba31844d · inbound

BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? cites this paper.

BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:12:54.855092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:12:54.855092Z digest=sha256:88549eaf7f4f05085fec91c8a3e0e6098c0cf619eed35d919f3289f97e8a69a1

Observation 11f9be90-3195-48cf-9e38-bc5bf5c83ee4 · inbound

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development cites this paper.

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:00:09.494253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:59:29.910200Z digest=sha256:2a48b60e64c20a3b00fe12a088e6edddbb80e74a81d8dddc0e268018f7c4eb2a

Observation ac9ac8a3-bcd2-4802-82da-6601d83b922a · inbound

SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution cites this paper.

SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:20:29.593108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:20:29.593108Z digest=sha256:fa6c869fa3f925864c45bedac70bd0b57ff6e911197184f452a2a56a42b5f2a4

Observation 962a08fc-1c02-42ef-9fb9-073405967f2e · inbound

Effective Strategies for Asynchronous Software Engineering Agents cites this paper.

Effective Strategies for Asynchronous Software Engineering Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T20:49:09.477849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:49:09.477849Z digest=sha256:6586bdbef04a5e878dd047ae5eb380c7b36cab6c6b36eeb43d0b308ed75241fa

Observation 1ca76b21-4b78-47cd-b3b0-9599ad8d659b · inbound

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks cites this paper.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:4eff21d8198f43b29bf2b17096138c0c43aa204a7d64f0d9dd2e22ac7308d9c8

Observation a319e9c7-0805-4589-ac66-9ef7b6cdb150 · inbound

Agent psychometrics: Task-level performance prediction in agentic coding benchmarks cites this paper.

Agent psychometrics: Task-level performance prediction in agentic coding benchmarks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:04.619944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:05:04.619944Z digest=sha256:5b7ea41945e50c741538cc02e1b1db110c1822a80ab77e11075b84a225b2fb45

Observation 10f2fbd2-f229-4841-b32c-c2e956e3fbb9 · inbound

Beyond Resolution Rates: Behavioral Drivers of Coding Agent Success and Failure cites this paper.

Beyond Resolution Rates: Behavioral Drivers of Coding Agent Success and Failure SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:33:16.411149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:30:40.124360Z digest=sha256:2ca9e3ed7afc318933f959a0273e9d1aa28b24a904ec76d6ebbada2671b941e9

Observation 5fe1d313-cf97-414c-9b64-77cc0c915ceb · inbound

AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents cites this paper.

AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:43:11.189056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T19:42:53.778980Z digest=sha256:acfe45b0ebc3817ef9cc39c506baf3c64b1d218ab2a6436cb7a56c6429b1d406

Observation 63be1b4f-e3d6-4aba-8384-0395a19e2e57 · inbound

Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures cites this paper.

Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:08:15.842530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T18:03:18.475429Z digest=sha256:c94000c6a502b460719229b9865ec1eaa381d7d1a077ab81d6d28374cec0141f

Observation 58ff9215-614e-4895-9c28-56911a4cc6cf · inbound

Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution cites this paper.

Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:45:06.682021Z digest=sha256:be2cdcdf53de8fb2bcd415d40a7bb6fc70351a9f30cc2b42aaa22b1c8a215ab3

Observation 03ad9450-e42e-466d-8228-29516e394f06 · inbound

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments cites this paper.

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:07:46.077831Z digest=sha256:d746834c67b92193f74f58a7bb710c72a8a2313877b320a4583e0e6b0c9f7adb

Observation b677eb3d-a0bf-4cce-a88f-80d918952e7e · inbound

REAgent: Requirement-Driven LLM Agents for Software Issue Resolution cites this paper.

REAgent: Requirement-Driven LLM Agents for Software Issue Resolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:56:32.201591Z digest=sha256:de9b2da19b96aa542ab8fbcc49233df156a585381ec93189eec3b0adefdd0bee

Observation d6077b97-22f5-4326-a9c8-a3fd6521f465 · inbound

ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents cites this paper.

ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T00:15:09.034899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:15:09.034899Z digest=sha256:c0a3b5facf485bf2cb781c3eee8484ff8c1d0b0943bb9bd589688621a3f17956

Observation 5eab2639-01f2-41c0-9071-b72e9123a2fd · inbound

HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help? cites this paper.

HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T18:20:22.151216Z digest=sha256:33d8ef769f3136a1e0273f95b22326b7fa5a3a0ec159703a332c35bf1b5c6be6

Observation 8a5df470-1c62-4ded-8643-a3592547b717 · inbound

Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems cites this paper.

Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:07:35.395872Z digest=sha256:2be2d2b813cf8ad33a9f5aced4433fcba014f3aabe8f892096fba0a94bb02a1e

Observation 2f0d715d-d53d-4fd3-8e5d-d106104f7ea0 · inbound

Evaluating Plan Compliance in Autonomous Programming Agents cites this paper.

Evaluating Plan Compliance in Autonomous Programming Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T14:52:51.349446Z digest=sha256:8430bfa31dfdc4a553dc94b60f75032dc37ec6d763a86e4bf67275bae380e17e

Observation dd6c239f-0d6f-4699-bbdc-5a86a72e6a27 · inbound

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks cites this paper.

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:27:11.257817Z digest=sha256:065fa358cfce7a017e915d9d0a2dbd9c808061407a89d36d30e0acfd013eebe9

Observation c4af8321-575e-4797-8840-69e3b113f30e · inbound

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents cites this paper.

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T06:13:32.434201Z digest=sha256:3969629626ae78478b4eb2fffedd7884e67097424c15afd8261ac997681fc94a

Observation f8bd41b9-d737-4592-9b4c-5503b1b9a1b9 · inbound

What Should Frontier AI Developers Disclose About Internal Deployments? cites this paper.

What Should Frontier AI Developers Disclose About Internal Deployments? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T09:33:47.869030Z digest=sha256:0cea336dc16a7ef92880147a38ec960e45ff9abc4132fbda3a32b40506520237

Observation 69b70e61-fb90-4039-96e5-803f5cd20d3f · inbound

KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant cites this paper.

KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T05:57:20.823210Z digest=sha256:82411a1aaefd77fa52206d100515704cb2eda621e822441d55611ea415c63a6d

Observation d410cc8c-814f-4668-841f-cd6609322f23 · inbound

KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant cites this paper.

KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:05:36.451754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T09:03:16.808596Z digest=sha256:693f2f3860022f3989a034a684241a74a5e41deae3443211c47181866c0bd5f9

Observation 9cd76b4c-6bd2-4e90-bdea-b5b53fe87986 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:2a6abd569fd0da2875205462043063927399ae3a837d58d9c6a8ff19ae60e2eb

Observation 8d381dd1-bc04-424a-b780-bf6687025065 · inbound

From Threads to Trajectories: A Multi-LLM Pipeline for Community Knowledge Extraction from GitHub Issue Discussions cites this paper.

From Threads to Trajectories: A Multi-LLM Pipeline for Community Knowledge Extraction from GitHub Issue Discussions SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T15:51:47.520468Z digest=sha256:d205c5ce17722b80d2f25dad2ae1a3560472129a47f7a7203cee6b68823d76e9

Observation 5dcb735e-0ec5-4e3c-b9af-e2542d42a03d · inbound

An Empirical Study of Speculative Decoding on Software Engineering Tasks cites this paper.

An Empirical Study of Speculative Decoding on Software Engineering Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T13:35:40.422686Z digest=sha256:587ffeaff014b69d295e9a7f6c7bc71b5217cfbad3f435551fce171bf6ff7e17

Observation 5a413685-774a-4426-aa15-6a61bee37eb5 · inbound

The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents cites this paper.

The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T18:30:21.856024Z digest=sha256:06388db73dc21539d89e25755ab8ac106e17d656425513534dabd3ce29ac5502

Observation 5226126a-18da-4d8e-827e-b9eaa1bf7171 · inbound

Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks? cites this paper.

Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T18:09:12.419652Z digest=sha256:3d700fbb158c0abebb870b875c2044a1d8cdb6e8be5bba1cf6f1d1f4a49cdf0d

Observation 95b77cf5-352a-4122-95bd-7aadd0001c70 · inbound

Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks? cites this paper.

Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T17:42:51.275272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:42:51.275272Z digest=sha256:0f46de93c2a285c7586f0eee70402a69296a73e156498d30cdcb1c95b10f1dfc

Observation 3ebf38f3-0b04-491a-b912-460c64803631 · inbound

ProgramBench: Can Language Models Rebuild Programs From Scratch? cites this paper.

ProgramBench: Can Language Models Rebuild Programs From Scratch? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T16:02:16.598404Z digest=sha256:efa23986f4f6bbb2abfd6f25611f3574cabc5a412c3ec5ac2e00eca869e69a85

Observation 2252a6bc-f4b8-4c1d-8e49-8054c3554788 · inbound

Reproduction Test Generation for Java SWE Issues cites this paper.

Reproduction Test Generation for Java SWE Issues SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T16:56:33.346935Z digest=sha256:69f85fb6995eeddbcfca0db6eb57af043a786f34b136f6fc495e7d9c7c8ceb47

Observation f5e6106c-2853-4b7c-8a7c-7c119cba2f7c · inbound

Reproduction Test Generation for Java SWE Issues cites this paper.

Reproduction Test Generation for Java SWE Issues SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:48:28.794195Z digest=sha256:1c9e2fb2733ff507742d178116e9eb5a342ee0579f5551cdaf0ca820b5083bef

Observation e5097a72-fd71-451d-b850-25b82c93f511 · inbound

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution cites this paper.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:d184e402731a0a2626e3e43b6ebbf0e0e2ff04a325d43aa2f135ecc0c87f815d

Observation f7cafb1d-2016-4127-9ca4-ea4caad9df08 · inbound

Constraint Decay: The Fragility of LLM Agents in Backend Code Generation cites this paper.

Constraint Decay: The Fragility of LLM Agents in Backend Code Generation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T08:48:36.848425Z digest=sha256:269b6af08361270d0a5bd1f0e4c1d273067bcfac512d60c301191f3be96d0385

Observation 54fcd544-adc8-4f04-942c-471ea125fde7 · inbound

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents? cites this paper.

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T03:24:10.130356Z digest=sha256:6d44dca310457a68c48b3a184a054c4b424f6e37a3014faa7d07c9f4bc24283d

Observation 77150a33-e2b1-46f7-813a-1d98e59558ff · inbound

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution cites this paper.

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T02:28:07.557119Z digest=sha256:1b73d524c05bc9221f82ddab7c428869b83b637b1d6f0c791646a13b9eeccb3a

Observation a504efaa-d9c0-4462-a00c-eda78f883fed · inbound

LLM Agents Already Know When to Call Tools -- Even Without Reasoning cites this paper.

LLM Agents Already Know When to Call Tools -- Even Without Reasoning SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T02:44:00.329276Z digest=sha256:bf09632e7a1baab7f4c1346c7a2580eb2ce507d737ae3d739dbd58e239a1c457

Observation e8ef8e8c-9e34-4861-b2a7-a19be5da0d1f · inbound

LLM Agents Already Know When to Call Tools -- Even Without Reasoning cites this paper.

LLM Agents Already Know When to Call Tools -- Even Without Reasoning SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:56:25.958471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-22T10:55:14.037226Z digest=sha256:a7c012bb5d40d091baaf455ef293b715f0e5ba9e6afb3f140d2caa25100b3a96

Observation a84ee52e-36b6-4a4d-be4e-32d46abc0950 · inbound

PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation cites this paper.

PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T02:36:40.696567Z digest=sha256:903d69284e1eb47a266fed335702bbb3b36800caf02f3ec312b775c2bf4fec53

Observation d83f093e-718b-40f0-bcea-3fed8c8fd736 · inbound

The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents cites this paper.

The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:03:58.419364Z digest=sha256:5f70e354ec3df7ea8480b1d59233fa4367ebc6e4258232b8ecdefbb316e92b21

Observation eb042689-a284-479a-8f12-f6a3a55d4fd3 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.742365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:bb04546f60c361e00f9472835e0a128ee2f93c74a0053e0bb47d11e487527564

Observation c8ade44b-5959-4279-a90b-bffac6b7c730 · inbound

Revisiting DAgger in the Era of LLM-Agents cites this paper.

Revisiting DAgger in the Era of LLM-Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:57:53.403116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:56:06.762156Z digest=sha256:b3aa4f6884c2d66c47becfa1596c2f5fbbd8d397fee43723f57638cbbc93396b

Observation 06d7c0a3-9c97-4644-b93b-5a6fc3091519 · inbound

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle cites this paper.

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:37:35.647810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T18:34:39.997353Z digest=sha256:dddb9bf1c032a04a290cf09c6dc4b659c7962f2ec09e9ecc877573c4e9df20c0

Observation b59972cb-0de7-4587-8e8a-2fac2b9f984d · inbound

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation cites this paper.

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:38:27.489459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T01:38:15.112910Z digest=sha256:3367ea0a409ae4c77620799007022df43912cb5174e031d3804c1aa82074c2c7

Observation 0ab82e84-cc86-4718-b255-fdc308fd9099 · inbound

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation cites this paper.

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T07:56:37.112844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:56:37.112844Z digest=sha256:6c00514c9e3ecbb35531c440b4421cf3aff3fdab2e1e4ddb212d5a7424a3a4b2

Observation 7d198de4-fee1-4182-9719-31488dd2f51d · inbound

SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering cites this paper.

SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T22:43:14.940231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T22:43:13.251812Z digest=sha256:731343ced2ab35aba38db14df6180b93ab4ea026c18d7e4c1e3ebd5618f144e1

Observation 9a80e3eb-b345-4428-ab23-02f0038ead57 · inbound

SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution cites this paper.

SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:43:15.136412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:40:45.397038Z digest=sha256:87151bdc66be463e62c73fd8572d78863cac3bd59983bc76fc17d45db6217e2f

Observation d7bc1a6d-a13d-4e66-8202-8bc9bd28b69f · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:33:12.743474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T10:30:50.256635Z digest=sha256:93bc65563f9d62137e303549aa10cbaeaf876a37128c873ba4eded41e5037d6e

Observation 311134f0-7ca3-4a19-90ff-6b762e1931c3 · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:32.945342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:32.945342Z digest=sha256:9a68313aa97dec105e0a67ac3f21a7c84256c3f33497c9417d0ef12b7123f14a

Observation 92e16258-0172-4c27-9d02-66cab342ed7b · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:39:43.728663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:efbe26496a081440246e247fe1ca4f5761f521428aa8ea42cd56baf1f54150ea

Observation 0f26fd98-371a-42f8-9caf-d8ac484d49e4 · inbound

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding cites this paper.

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:53:54.613591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T01:51:38.671365Z digest=sha256:2b426e8f5a8852dc9766acd65fa298d3a3b67bf1a32e472e83c8ed9c443cc3f9

Observation 5bee8a10-af87-431e-842f-9cf15f3f0aa1 · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:23:57.583837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:20:43.780849Z digest=sha256:a4b3c344a1f11a6b450394867a71fa093adbd09c45b43f52e76a2a353c1cac76

Observation b89c8d84-bfd6-4425-9fde-090fade0186f · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:51:21.923553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:46:26.124683Z digest=sha256:b7ef6f36cd9a2ade9db0b099b4eb14b34c5d9360f58e830679b775fdd66485ec

Observation a92ef363-b45f-4990-b02c-2b60d60f0cff · inbound

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents cites this paper.

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.654428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T17:31:25.538977Z digest=sha256:4f8639a342949b38c4d8ea69e6265b9ee93233c15c88e434dcb10026c753e3a2

Observation f0114c10-1b66-4e8e-8580-10434f5800a5 · inbound

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents cites this paper.

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T03:09:28.093912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T03:07:18.501526Z digest=sha256:1cfa09f45cdc05a49057d8766281d367c784f3f18d172d48ce89ad621e20e4d3

Observation 90a487ad-326a-4460-8376-7407687d6d3c · inbound

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate cites this paper.

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:29:34.576114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:27:25.041652Z digest=sha256:a97ec42af68f3f72b5dec252b2387ca723fda7d03c0af4eecaf1657c96a658c1

Observation c10373ab-a14e-46ad-9db8-3d85178d6f65 · inbound

From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents cites this paper.

From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T05:14:37.631743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T05:13:26.057531Z digest=sha256:99b85eb0afcbbc5f1a630b9326c60e50c07442607fe1dcb00300e8a73d74ab98

Observation 05de791f-d537-42cb-b540-d2f9f442d50c · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 113

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T04:40:22.646987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:102114be96b377d6ce4606ec111917c601989a262cfcaaef432fcc43073cd9a7

Observation b5985207-77b9-45da-80a5-780d2bacb82c · inbound

Stop Comparing LLM Agents Without Disclosing the Harness cites this paper.

Stop Comparing LLM Agents Without Disclosing the Harness SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:25:45.870677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T23:15:37.160073Z digest=sha256:d185e3f9f53fdda9b2434de7af8a676407e9a822f04a0f9f059894840e83bfa8

Observation fc23a8af-ad45-4bc1-8394-9afb0e7626ed · inbound

RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations cites this paper.

RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:53:57.818067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T20:53:30.382870Z digest=sha256:3c8bc696fdaa40734330022627e17c83f777f92826d3734db4d684353f732c5f

Observation 3d95fc5b-f528-4013-b57c-824d67e1275a · inbound

Agentic AI Workload Characteristics cites this paper.

Agentic AI Workload Characteristics SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:13:58.835453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T20:11:50.722787Z digest=sha256:b7d70f0bf32bcc7b6ffe27375322378e873d0fc64ba803ab32d1c5fd10bba6ee

Observation 74362d7c-e805-4960-93e9-4499ec8a3d65 · inbound

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? cites this paper.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:33:45.250290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T17:29:36.006340Z digest=sha256:e87623ce8cb4498a8c05d20065e502414b6a525b345347b8aaca87e424919953

Observation 6cd2488a-f694-4ced-898c-f6f4d4c8cbf8 · inbound

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? cites this paper.

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:10:39.062654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:10:39.062654Z digest=sha256:6ddf11b805d12dcf3200b42a5122521ad44cf05509b6aeed9427b5895b01f6d6

Observation b6315f07-7902-4aa9-8442-d029cbd258d9 · inbound

TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems cites this paper.

TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:13:35.844714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T16:13:30.977123Z digest=sha256:b6b6fa87add937c4cadbab9d7b59a9e79e88d12b84544187a66131d65dc66217

Observation 46ca2011-88b9-40a1-9473-b1e2431d08b9 · inbound

TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems cites this paper.

TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T13:09:14.215805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:09:14.215805Z digest=sha256:d3e7d35579eb46c69072cce44e4a95d36d97d82cc0af67edcfbc26089e2c2e4a

Observation 38d2dc92-3a74-47c3-b03e-2a31eefbb616 · inbound

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems cites this paper.

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:33:32.782352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T15:32:21.737028Z digest=sha256:437ed5c3f62a98e6bd2d0ea7a78deb13998b1d54e0ce1f0da92bdab7e581a15c

Observation 51a6c18d-b57e-404d-ab47-1d6f4aac611c · inbound

PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game? cites this paper.

PTCG-Bench: Can LLM Agents Master Pok\'emon Trading Card Game? SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:23:12.590793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T07:21:49.763994Z digest=sha256:4a9c17e5186319faaef1637bbdaa030d629ab999130c9860159153fbe4fa18bd

Observation 037da2a1-4226-4704-b11d-4c336c4f9fc1 · inbound

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems cites this paper.

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.174304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T00:05:31.780655Z digest=sha256:4101ca65ab0624fde210c55da9199a7548bc6f9337d7ccfed45af7ec2a0bc41f

Observation 46bad182-e56a-4676-ad27-cfd6422c3c78 · inbound

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets cites this paper.

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:06:12.323633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T21:35:45.197423Z digest=sha256:c44014db9fdfec91e510fc6340c53564af1a874a07511dcb0f270eb7e3d14056

Observation 6bceb751-5100-452e-a60a-4ac3614bccdb · inbound

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI cites this paper.

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:06:13.588462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:30:56.324289Z digest=sha256:782a775a10c0b2af802490318b02af6416ee1ae0644d5316fdfdb4eee5b570b4

Observation 3c6ac21d-2210-41c6-a001-7b6bae2bc7d9 · inbound

DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration cites this paper.

DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:56:28.775459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:33:20.920760Z digest=sha256:fae88771178e0fe86be49acb6e1c978ae56b6796a18211a620390b79ce681962

Observation bc911c64-2378-4cb3-8db6-81ffbc66ac0f · inbound

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework cites this paper.

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:36:57.385694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:55:59.002171Z digest=sha256:f6296d797cf256c4d7601e075eadf3aa18ad14795e1f6fbdaa6f375a94739104

Observation fe56401d-1c7d-4c4b-a87f-6c00174200d9 · inbound

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments cites this paper.

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:56:56.839653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:45:28.693098Z digest=sha256:24f34a6ec639c3068fb370448824973f682d08ff1f9ce0c6ee42dd7a4dddd9c1

Observation ebdbb5be-f447-4acd-ab0e-212b6ac70a40 · inbound

Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference cites this paper.

Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:56:57.375937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:41:32.686349Z digest=sha256:f5c625ff7eb8b350ee9935060efb2ca4654666f54c4ed7bf4db4f2d3f8427482

Observation 7f14fcfe-164d-46dc-a39a-0f60100073ef · inbound

SWE-Explore: Benchmarking How Coding Agents Explore Repositories cites this paper.

SWE-Explore: Benchmarking How Coding Agents Explore Repositories SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:19.218108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:21:43.468977Z digest=sha256:a5603c9ea671af550d8b78b4d08483db5614a82434d1785a97f6cc66dd587b33

Observation 0656c1f3-66af-4007-a018-7be7e408ce6f · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 151

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:50:48.437287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:678a3d65ac645268bd3b3bda8473ed2a41598dfc6f990cacc8d32601e0cd7774

Observation 54d18632-e300-4c7e-a44d-55bdc2d6c79f · inbound

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks cites this paper.

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:57:47.632014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:39:51.630229Z digest=sha256:0f10ce271193cb5e8441d14184a0775d1ec0f0c9302b754df56ea8ce6c9ee4c7

Observation eeb84929-671d-412e-88bc-453e6662a8e7 · inbound

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility cites this paper.

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-27T07:10:41.772430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:41:41.799596Z digest=sha256:679055404835ae1c880e775a9a0fb7542f8f126e988de2be4b60406572fa7eac

Observation b428d440-dbb6-4f23-94ef-820c10d84236 · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:17:25.678621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:451437a466d0f8bb04286973e89b8ff02b82574766734dcd88dc11ee9b0aaffa

Observation 4de31687-758a-479c-8de5-17dd8a57e262 · inbound

Dissecting model behavior through agent trajectories cites this paper.

Dissecting model behavior through agent trajectories SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.927589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:27:39.812496Z digest=sha256:e50ba5ead11e7482eba9d8db71d0ead3b5e1d751aacc2d1af2f108d35c98785c

Observation 08baac0d-7a43-4972-a0ba-a83fc0b01753 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:08:58.889649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T23:48:58.497927Z digest=sha256:90b7892b467c2ed197bcc67397497db08bd166ccd47ea8206b5d7ab15847d5ad

Observation 3a91eb18-0a35-4798-8dd1-c3ad097ae358 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:03.064744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:03.064744Z digest=sha256:7ce4c04a36ea12cbecb87e2c39a6e94f9965fc825999feec8062cb8e9231432a

Observation 3a4a2b6b-4888-46f3-913f-78879420a5a9 · inbound

A Framework for Evaluating Agentic Skills at Scale cites this paper.

A Framework for Evaluating Agentic Skills at Scale SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:08:59.894232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T23:42:37.325511Z digest=sha256:466eed5a05dca7cdb2fa6539c7dee238ceda3e51529247ab8c11929376c82697

Observation fc0d4a20-cd2a-42be-8c3a-c3dd72d3102b · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T08:57:48.108196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:9c476149ad0ff1e50da4ff36e9c1b8f7a8a5a08193ee2ce0aa5aa2cd7e58b651

Observation 90e7165b-426d-408f-9739-52eba309fccd · inbound

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play cites this paper.

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.510064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T20:57:49.840546Z digest=sha256:e495ff3a3a03ee501637483f58ac6afa09a3ff43277316d216f4cd450965d2cc

Observation 68aedaf1-b5a4-4b14-a4ef-e54b0d753528 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:19:23.944512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:de73feaad5e09a6d9dff79b2597faf3c8aa7ce80b5b135c01e2fe9712e596654

Observation 97151aee-e408-496f-acda-bd857ebc1ccf · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:36.919614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:ded2052800878db70515a71b25ca4810adf44d9c5c0a8bf2b27a2ff791a8ad38

Observation 4bd2557e-bdf9-4d9c-935d-b65da3acb0d0 · inbound

SwarmX: Agentic Scheduling for Low-Latency Agentic Systems cites this paper.

SwarmX: Agentic Scheduling for Low-Latency Agentic Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:39:39.048754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T13:10:48.369285Z digest=sha256:99ce627ef9c8e435cd313ac6b6dfd62199b27ba0716c81c6bac31c6695074ac1

Observation 750246b3-5d42-40d2-b85e-48eda54ea020 · inbound

SwarmX: Agentic Scheduling for Low-Latency Agentic Systems cites this paper.

SwarmX: Agentic Scheduling for Low-Latency Agentic Systems SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:44:36.554229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T10:40:21.294555Z digest=sha256:c82e03f465d0fd9f06333cad23cccf14375e139cf8dc6f6d1aa3b4fbbf957d50

Observation e1f62e3e-8627-456f-bd3e-a213102da0cd · inbound

CFAgentBench: A Reproducible Environment and Benchmark for Autonomous Construction-Finance Agents cites this paper.

CFAgentBench: A Reproducible Environment and Benchmark for Autonomous Construction-Finance Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:19:44.527905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:50:14.868505Z digest=sha256:bcee702abe54109bc1be51ab29e1b7034cddac8a72d549097a93b18405251a11

Observation 91914b4e-2ab9-4738-a16c-95041df088df · inbound

Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent cites this paper.

Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:49:41.790205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T10:59:52.756992Z digest=sha256:e90830f5eae826da464cde1a516fea250b36a60fe8ce2fc65474f112938c609b

Observation 5fef82c6-8b56-48dc-8757-99ded1855f77 · inbound

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation cites this paper.

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-06-26T08:49:15.248983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T08:39:41.575494Z digest=sha256:443bc8c91b7b369e076a1d6ecad9b1d4c97309b7f512cc1a8842166033431ad1