Pith. sign in

Paper Citation Record · LEDGER

SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2504.08703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.08703 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:32:07.773449Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ec09632-f2d9-4c6a-95be-8947c0d99554 · inbound

MigrationBench: Repository-Level Code Migration Benchmark from Java 8 cites this paper.

MigrationBench: Repository-Level Code Migration Benchmark from Java 8 SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:07.773449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:07.773449Z digest=sha256:8dcc8b138b9313e1889a92cc292baa800a567586870b3a159cf26b4b5c159605

Observation c6840061-8020-4f91-9307-9c76666be5f1 · inbound

CoRet: Improved Retriever for Code Editing cites this paper.

CoRet: Improved Retriever for Code Editing SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:33.519926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:33.519926Z digest=sha256:182a34f5f50ddeb436a97291204251c624c59165b7aab2a06d86b5e74c2dd50e

Observation 8d5a436d-2c7d-4f6c-a0c4-cc0393dca03d · inbound

SemAgent: A Semantics Aware Program Repair Agent cites this paper.

SemAgent: A Semantics Aware Program Repair Agent SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:26:51.172895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:26:51.172895Z digest=sha256:bbd3b4f69be4c9e263ed915b39af52de250c90abc0c631d692f1d3e99ac7cb1f

Observation 89ce5fba-1f0f-4c4c-a4b2-0eb064c6aa2e · inbound

Is Your Automated Software Engineer Trustworthy? cites this paper.

Is Your Automated Software Engineer Trustworthy? SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:06:53.865883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:06:53.865883Z digest=sha256:703a5f0845c781af3892b1c3f9b39b159f8ab3f26260de9dff6a049394bc9f3d

Observation f959a0d6-0aac-4caa-8f88-4f6a70d79a16 · inbound

AI-Assisted Fixes to Code Review Comments at Scale cites this paper.

AI-Assisted Fixes to Code Review Comments at Scale SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:29:48.231955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:29:48.231955Z digest=sha256:da35f727640602d2238b9219e00fc5d37c74e4791559d77e15c7804be8c15174

Observation 9db39d24-7605-4fab-ab05-23178030fac3 · inbound

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition cites this paper.

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:14.337616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:14.337616Z digest=sha256:142fa4b16abe34f97cb74b39a0bd82e832cb7420b424ed900b2c8581bcc351df

Observation 5a38868c-0153-4821-b29e-ffea33ada3bb · inbound

RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models cites this paper.

RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:28:37.987551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:28:37.987551Z digest=sha256:9b6aa66352ec911c6c2147ae0ad695d38721c5af48e117544543d3cb3e571ce3

Observation 556750d0-1a52-4700-a53a-c74ffca308bc · inbound

Vulcan: Instance-specialized, Verifiable Systems Heuristics Through LLM-driven Search cites this paper.

Vulcan: Instance-specialized, Verifiable Systems Heuristics Through LLM-driven Search SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T13:15:17.244078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:15:17.244078Z digest=sha256:ec9772cee11d95fa195a6b1cfe8129ec1d9feb6765f46cc6667e0f5d4b8b6293

Observation 7275fef0-9400-493e-bb0f-73e710672053 · inbound

BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? cites this paper.

BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T19:12:54.875615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:12:54.875615Z digest=sha256:e7987e93cf8d6e0dae5ba2b6d205a060cd22ef075d1f0639a089da91b2f5161a

Observation 4de5c26f-12ea-44f2-a967-ffdc32a0d7ec · inbound

Reproduction Test Generation for Java SWE Issues cites this paper.

Reproduction Test Generation for Java SWE Issues SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.513480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T16:56:33.346935Z digest=sha256:ebdff7829e0ebbdd9293e4ab2576081797fa22a9931a0cfd64593131ba05ca1e

Observation 084b3ac7-c203-40f1-b9a1-b5cc75fc3470 · inbound

Reproduction Test Generation for Java SWE Issues cites this paper.

Reproduction Test Generation for Java SWE Issues SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.895264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:48:28.794195Z digest=sha256:6cc6037a3750071eafbb34edcdc07ea4e520ecca79f2cb0fadc64d2081934d85

Observation dd5f6290-f84e-4999-9947-6ed841d1a0eb · inbound

Constraint Decay: The Fragility of LLM Agents in Backend Code Generation cites this paper.

Constraint Decay: The Fragility of LLM Agents in Backend Code Generation SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:31:11.916648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T08:48:36.848425Z digest=sha256:2bec5418aa9ed59de795a2a04f677b1ab4b9f32751f9771b64adfee7da6d08d9

Observation 79e4c14f-be29-40e1-b775-334c6b6e7ec4 · inbound

SkillMaster: Toward Autonomous Skill Mastery in LLM Agents cites this paper.

SkillMaster: Toward Autonomous Skill Mastery in LLM Agents SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:24.072276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T00:59:03.361037Z digest=sha256:3f47f0eae7da949c3c66ab8b6ca8eb6a6a2b1308d3a70f32781557d609c696aa

Observation 0ce48223-aa1f-4691-8e5e-fed59c3bf9d9 · inbound

SkillMaster: Toward Autonomous Skill Mastery in LLM Agents cites this paper.

SkillMaster: Toward Autonomous Skill Mastery in LLM Agents SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.879949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:30:09.985448Z digest=sha256:c9dc9f867b2953da5b228cf5f9cb37dc9e61edd6dd38705d5b39ff4cdf01be0e

Observation d2c908b6-d85d-4bed-ae64-dfdcbab49330 · inbound

Do Coding Agents Understand Least-Privilege Authorization? cites this paper.

Do Coding Agents Understand Least-Privilege Authorization? SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.990907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T16:34:14.379419Z digest=sha256:a2e72340dc12c2cc0e557154c9a852884181ad609c0559ba8c2875971f4ea933

Observation 7a29ded5-7b9f-49f1-89a8-621a29370407 · inbound

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models cites this paper.

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:13:58.857087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T20:11:45.644494Z digest=sha256:f0aa5bd488cae4e90fcb1144e99bffee3d7168a631f49fb7354774d22989fc12

Observation ef2aa17a-230b-4cf0-bbe1-411dee9dd333 · inbound

RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations cites this paper.

RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:53:57.810161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T20:53:30.382870Z digest=sha256:526e86d0f8442658081371e5b835f24857385476b9fd1418b18c4120f219aff9

Observation be5bfb7f-fc18-41eb-91f2-77ad0d2cecc9 · inbound

HARP: Measuring Harm Amplification in Multi-Agent LLM Systems cites this paper.

HARP: Measuring Harm Amplification in Multi-Agent LLM Systems SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:44.861699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:20:09.363041Z digest=sha256:2952faa12e58e749536b32d9812be219a7e4867d3f07bb03e06b324fb86fd4c6

Observation 4ce69601-6eb9-402b-b649-8167127c7905 · inbound

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code cites this paper.

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T09:56:51.737060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T05:21:34.772199Z digest=sha256:f5159bb758444a1fe8ada7e634c2996ebf8916e4b7cc0b5ed8601313ec11071c

Observation 40079cff-c810-4807-af2b-5e86041df481 · inbound

Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey cites this paper.

Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:29.843586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T18:18:41.453580Z digest=sha256:24246615ea4160c9c116c4c2b617c93edb7bd4e5c605ce464ce0c6a9a1190cdf

Observation 3db99ded-e4dc-4bbd-be11-530475791d76 · inbound

Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent cites this paper.

Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.745699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T10:59:52.756992Z digest=sha256:69104036737dfffe3ad489a82179ac88008f0e789595519b1fe14fb41600bd98

Observation 2f60a935-f5ec-46a0-a1ec-b17d58d41293 · inbound

Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents cites this paper.

Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:19:43.615477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T10:13:31.405238Z digest=sha256:78d977fca38d04b7ca796512f450f9d396a2d0df11815ad168577d043e0ea1ee

Observation a0e49560-a412-4918-ac6c-83696027dce1 · inbound

Unlocking Model Potentials Through Adaptive Multi-Agent Scaffolding for Efficient Issue Resolution cites this paper.

Unlocking Model Potentials Through Adaptive Multi-Agent Scaffolding for Efficient Issue Resolution SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.786589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-25T20:15:00.517571Z digest=sha256:fab861571ea346dddf4cd29075106afab533ac01cd3c1a77c3e4ee2b1e7d9914

Observation a4370db8-2b61-4139-983d-3fd729037edb · inbound

LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution cites this paper.

LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.588686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T08:42:10.658469Z digest=sha256:d4a6fcc59278b8435bc1cffd01a7dbb29f4665fc13df1391c8cee9adc0768571

Observation d6b33905-2d98-4e07-9a77-ff3abf94e719 · inbound

SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests cites this paper.

SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:26:47.285372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T08:22:54.966922Z digest=sha256:48ae7b147e02adc4322de5f84b5df385697496fb9839f67dac0da9104cbb6fb2

Observation 1a25496d-0ef0-4a49-8363-9c8639d30965 · inbound

What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents cites this paper.

What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:25:02.410420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T14:16:25.641101Z digest=sha256:e0336cd97ef495ca1b356bdf3a58807ee73e95932dad017d11f905c7469830c5

Observation 38dcf282-46d8-40dc-b7b4-5cbd9d787ffa · inbound

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs cites this paper.

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 149

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T13:57:06.670054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-10T13:49:17.343893Z digest=sha256:1bec99529a1ff59a35d1b4bedc03b244f395a7b4a1b8d3ad2421309a8d81a311

Observation 868d74f9-c2a1-4362-a2f8-dc85f2d96a19 · inbound

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation cites this paper.

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T10:46:01.272433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:46:01.272433Z digest=sha256:b25c370e847883a22bf47ec546ad588171da20d82b4480cd4b2445cb919d4da8

Observation c6dbbf03-e191-4648-b221-062e31cd8bef · inbound

Bespoke Visual Assistance: What and How do Blind and Low-Vision People Create with Agentic Programming? cites this paper.

Bespoke Visual Assistance: What and How do Blind and Low-Vision People Create with Agentic Programming? SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:47.258312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:48:47.258312Z digest=sha256:eea57eda86fec641318485c00fdec0985f8efd324d3bf6972c49d92f66dbb6a2

Observation 93ad891a-740d-4833-80ac-59b502d4e4f4 · inbound

ExplainBench: Evaluating Code Explanations from Agents cites this paper.

ExplainBench: Evaluating Code Explanations from Agents SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T15:36:24.159083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:36:24.159083Z digest=sha256:7a6de13484c6bdc88b8ae627ad05ba6b0533262641469c836d680326d3a21ebb

Observation 6f6eecbf-26b4-4811-83cd-9892d8dafca8 · inbound

SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements cites this paper.

SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T07:55:25.299066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:55:25.299066Z digest=sha256:8dec264c1d5c39870b3dbd9e70f9892f66da1e78f019e95f36b987fec5e7f8a3

Observation dbfa0fc3-3702-4796-b194-091c90ce97f7 · inbound

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks cites this paper.

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T03:01:05.590423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:01:05.590423Z digest=sha256:ee247b215bfde78f9bbcbd8669850ab971be79a0f0db944be1b56aa8cfc0130f

Observation c876ea63-a64d-4a70-b5fb-2e4fe5bf4988 · inbound

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks cites this paper.

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T04:24:46.096618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:24:46.096618Z digest=sha256:7ac300d10b8fcaea5080b35b01e49c11044f06fdafea6ffd0b7a8d7afd5265c9

Observation a70a37d2-dabb-4c9b-9b66-d9a51647561f · inbound

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring cites this paper.

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T10:33:05.142802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:33:05.142802Z digest=sha256:eb87cf5da9bdf657796900235a72ee1a4f89f97623bf9e25e428ded9d98ca04f

Observation babcfd37-37e3-40c3-8eda-c32d9e406937 · inbound

One Recipe, Many Harnesses: What Self-Evolution Encodes Across Languages and Models cites this paper.

One Recipe, Many Harnesses: What Self-Evolution Encodes Across Languages and Models SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Reference 2025

Resolution
malformed identifier
no resolver link, observed 2026-08-14T04:18:57.350641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:57.350641Z digest=sha256:5350ab8b5bded6f2ff08bd91f225de3d86f248c3b0b18afb8b2e3c3d51ac4c7f