Pith. sign in

Paper Citation Record · LEDGER

SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2410.03859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03859 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:00:17.468272Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 569e9854-eada-48dd-817c-1d3e64809d0a · inbound

KernelBench: Can LLMs Write Efficient GPU Kernels? cites this paper.

KernelBench: Can LLMs Write Efficient GPU Kernels? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:55:02.102197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:55:01.976356Z digest=sha256:be351f9df6304e3150db7ee6305d8c4e5faec8718abd5823c1464eb5d8b6529b

Observation 0182c42b-11f8-4bfa-9994-6139a69d8496 · inbound

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution cites this paper.

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:27:56.346568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T10:27:56.185943Z digest=sha256:3ac754e5621142dea8ea01c52cd4ffa8a84c701c43fd371ce5f720192cc90a18

Observation bb7b2022-6b02-4d15-b491-9a50cc08c345 · inbound

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments cites this paper.

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:57:16.238337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:57:08.314088Z digest=sha256:b57566e3f371568b6266af759cfc3bc73435a0c763f13c143f8d21d25d9b2688

Observation c65b9d81-33ba-4eb1-ae19-eae98d3585b1 · inbound

SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design cites this paper.

SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:15.130136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:15.130136Z digest=sha256:1bd731ae94ec6fc6c6ecd80385e918b4d4676b5148dcd0a74d2f01701ecb2e49

Observation 9a21d675-1c13-4335-81f9-84ac9f795873 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:17.468272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:17.468272Z digest=sha256:58001c05a2616106d9cb4f38462b0296c4266adb1b18803c77f4fa3f23d12ad6

Observation fe719371-92cc-4834-a8a2-9e075993215e · inbound

Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows cites this paper.

Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:34:31.618617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:34:31.618617Z digest=sha256:9c98d51f79346e7b1dd442c493e523f4ca26bd73709215e59286748c5f70a8c1

Observation 9ae7f4e5-42dd-471d-b023-015998974436 · inbound

Multilingual Multimodal Software Developer for Code Generation cites this paper.

Multilingual Multimodal Software Developer for Code Generation SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.357081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.357081Z digest=sha256:1b7d3d4680c744e68a15ecd90f303db8dcdfb450ed2df199c85ebe80db9a8930

Observation 3789f81f-780c-4954-b3d2-29b734b7aab6 · inbound

The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering cites this paper.

The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:38:55.375974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:38:55.159673Z digest=sha256:8e279f1683ea95977ebb40f2aa1bcbb74f740c32819c18a25abb5c6bc326f477

Observation 0521b7dd-b328-4137-8a6e-81d155a0e877 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.799098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:f1c69825a6399da1bda28397852fd64bd0c306269bf7fe160cfb5a6c0a39acf4

Observation ad34d2a1-8cf0-47ca-92de-de061c9c6bc5 · inbound

Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair cites this paper.

Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:29:20.689824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:29:20.689824Z digest=sha256:3898221367008d285444f2c8b7d321ec8f3f1b57d9eeb5cedfdf4928d226c5f8

Observation c816a82b-f825-466f-9d70-8f6afb83e9bc · inbound

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? cites this paper.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.774713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:bfbb3330af65ad60073179783ebc9df915a931691af34d4312462068cf938101

Observation 95d2e266-5e50-487d-bc5f-4125e7126727 · inbound

How can we assess human-agent interactions? Case studies in software agent design cites this paper.

How can we assess human-agent interactions? Case studies in software agent design SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:40.623619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:36:40.623619Z digest=sha256:a923077f4fa8f81fb57c3095945c05f12bb3277dcfb60a239469c6db2c99fa06

Observation 1039a3cb-282b-4028-9ccb-226845f49d14 · inbound

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios cites this paper.

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:28:24.528054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T20:24:40.939455Z digest=sha256:bb295e92880d8d392218b5f3a36ad7f7725ca50a20111d76ee4c261b572aa677

Observation 66cdd8a2-8d5a-4c27-bc17-cbdd02e06549 · inbound

Compass vs Railway Tracks: Unpacking User Mental Models for Communicating Long-Horizon Work to Humans vs. AI cites this paper.

Compass vs Railway Tracks: Unpacking User Mental Models for Communicating Long-Horizon Work to Humans vs. AI SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:11:01.384484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T14:09:33.786576Z digest=sha256:8f013ad0f0b216a05139a32acf25eca625d467bd6bb4d053d4fa37d417e39811

Observation d3ebc9cc-6ce8-49bf-a4b2-957fe8b7779d · inbound

Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective cites this paper.

Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T23:31:17.421699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:31:17.421699Z digest=sha256:c796afa3f8ef3b5ddff2cb1945cbb3026ef16fecb93b5dd2e17e620c070aa6f5

Observation 6e31f8ef-6b5a-44c4-8f5d-6d244db49fe9 · inbound

AlphaEval: Evaluating Agents in Production cites this paper.

AlphaEval: Evaluating Agents in Production SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.877213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:b39154a9d531f6e39fa7a2b0d06df8fabd0514b1511ad5544caea3454bdd187a

Observation 7d13fd6b-bae3-476c-9a8b-788fc466c5fd · inbound

Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering cites this paper.

Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:56:26.738521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T13:26:26.648410Z digest=sha256:f0c836ebb665ffcd822aa61a6f6788db993fe095f4fcde4712b75f73937d4ea1

Observation bd71bb2e-2a45-426d-a00d-50930f3ba827 · inbound

The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents cites this paper.

The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:25:39.708557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T18:30:21.856024Z digest=sha256:b2fe797b27f31d1333acc5addebf59ae5e8a2d364b9ef8ea942f554e613a4a25

Observation 5356b90d-e7b9-4a66-a319-e9d47b33bc26 · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:21:26.539858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:21:44.087943Z digest=sha256:6bce883adf48ffc779a2652de9d3e6d48545d75726ab9b7e9a6a7d28c965693a

Observation 085178c1-03f3-4b6f-956a-f03900be35ec · inbound

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning cites this paper.

SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:32:59.504330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:30:42.766390Z digest=sha256:884c92f2f9ea8eba3339d7e865f3e9be3afd768ae86fb68768054b85a96eb529

Observation 7a779983-db03-424a-b2f2-10e14520202f · inbound

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades cites this paper.

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:33:32.348124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T02:31:18.183715Z digest=sha256:227cf63c702a99a58788939186438f14883201ea7ba22b1ef21e29857a50d949

Observation 41f1a7c4-cf12-47f7-ad24-997f0c3ee081 · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:43.750853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:02f9aa3d63eb856e3eda5e6e55955e30fc5341abb1b489242776c4eada16e872

Observation b5bdc529-81b9-402c-8587-8a1d835b3434 · inbound

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents cites this paper.

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:22:51.863309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:06:57.377183Z digest=sha256:75edd23624aec42151202f678f9cf6bf4bf1577ea1e586ea194c87939ca94aea

Observation ccc90743-de5f-4748-b3b5-bdb7d7df450e · inbound

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications cites this paper.

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:42:36.256892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T18:53:18.645984Z digest=sha256:708908437d6fdedf13474a972fcb6eb83607673b6485cfa4361072fe5afb88cd

Observation 33d5f3d5-9b40-45c6-b172-dab231b14e54 · inbound

Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations cites this paper.

Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:22:37.804536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T18:44:12.893087Z digest=sha256:b76ecd01e6519646b42c03322575ea70f4ade8cc23ef361a7befbed95d6601a0

Observation bb2696f8-df19-4e7c-90af-a0223aacb4d4 · inbound

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions cites this paper.

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:29.215914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T09:56:36.860369Z digest=sha256:cd5887e54083343b38ed52baf312dd531427cce2488e6e68bb97cdd107ca2d6d

Observation 38fa32d3-1c5a-4cff-b2b0-c8185b33f454 · inbound

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement cites this paper.

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:27:04.283039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T00:26:22.041924Z digest=sha256:7942febaec226785c62c4981716c4a224ba8014bbaefabca7860101173196ef8

Observation 19d84601-11dd-49dc-aa1c-5df8de9866df · inbound

What makes a harness a harness: necessary and sufficient conditions for an agent harness cites this paper.

What makes a harness a harness: necessary and sufficient conditions for an agent harness SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-27T15:21:00.801996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:15:57.372858Z digest=sha256:11d5bbdaa77e06f94fd51eccf5a7c907feff84b77e24e4bd13038f7374dd1120

Observation b70e4fc8-eb15-4695-8bf6-f249c480ea16 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 282

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.639775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:e4e8ccff34c063b6f76cb8881f7244b57b55223ff6c97376aecca22142f2e287

Observation 8d02e1bb-f5eb-4d85-a741-0e422d144021 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:02.881675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:6d18c5bfb4ead2664328dd63ed1f74cd82d1130e55ffee9a62a9a4dfd07f4fa2

Observation 3c8df488-9523-4581-9792-dc4616f0332a · inbound

Dissecting model behavior through agent trajectories cites this paper.

Dissecting model behavior through agent trajectories SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.805422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:27:39.812496Z digest=sha256:df5d570ac0ce8685ccfeaa80739043c050ea87b5256e3e7416a73b37ebd218e4

Observation 8a038349-a69d-4a13-9a3b-b18c1670b277 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:58.945492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T23:48:58.497927Z digest=sha256:909e413b9aca5802f111ebe6718052633654579f68654de282bdba9511879667

Observation 2f80db8c-a37a-4106-98e8-666af8c28a14 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:03.963144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:03.963144Z digest=sha256:120cc627e3fbd4e96ba6271304b817f048b11cee1e098a0ebeef5063c424bf2a

Observation 41139610-86fb-4ddb-b748-6ef934e23745 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.236521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:f2f1793f6022da4e9194ec59241154c2b0033b5a3ccc4d73fda491461220539f

Observation e12773f3-bcc7-4168-9eae-0b672d7bcb66 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.939382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:a36a9c4fa1e53f0d98dc4596e151a778b95d88835c3c376bd9e9bba4c0d3bc07

Observation 59593680-b4e8-4ac0-adfe-660c1024b766 · inbound

Unlocking Model Potentials Through Adaptive Multi-Agent Scaffolding for Efficient Issue Resolution cites this paper.

Unlocking Model Potentials Through Adaptive Multi-Agent Scaffolding for Efficient Issue Resolution SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.769919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:15:00.517571Z digest=sha256:12db7bc6f67089129705194a5cc00bc91fda98264d98cd628fa59ce830a29850

Observation 6c9d3edf-dcc3-4f85-8f1b-787e72c9847a · inbound

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks cites this paper.

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:37:16.506209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-02T18:19:43.146102Z digest=sha256:b243a53a51d4ee803e9d8e4aa93a6f8403ca9ee64175d0ce7b3cef81c8d1aa0a

Observation f4605d47-9cd8-480f-9cc4-c8952400f1ea · inbound

LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent cites this paper.

LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:46.348120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:46.348120Z digest=sha256:d17567757f11f45e112270023632e56fba10613128aa84e1eedd990e5330e090

Observation c428be79-3082-463a-a539-fe217ca6d6cf · inbound

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification cites this paper.

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T18:29:10.457483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:29:10.457483Z digest=sha256:9f818fcc08e3cc223117b182263cfcc176f6c27c2bf1f37b4639c78e9cfe74d4

Observation 474521f1-6871-4316-9674-9fed3c956337 · inbound

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports cites this paper.

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:40.658366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:02:40.658366Z digest=sha256:a28cd556d7c8ba2511d012e65b9954171c22ea33716ebd2623635bf705fb1b9d