Pith. sign in

Paper Citation Record · LEDGER

RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.07832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07832 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:55:46.635284Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c9ca7b5-642e-46a3-8476-9a6b21702808 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.452526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:d008a1b1e12b8cbf5370f87f88b6405488d2f99ac072ff7a84196391501a1fbc

Observation 6934c22c-1f5c-4c9e-8eb2-ce0e23850d7a · inbound

Evaluating Code Reasoning Abilities of Large Language Models Under Real-World Settings cites this paper.

Evaluating Code Reasoning Abilities of Large Language Models Under Real-World Settings RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:34.165493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:23:44.762007Z digest=sha256:a53ac6ff380a9aa2b78f0319cd282e7da9d28f5f31088e79b23e0a2302ed465e

Observation 9c7a203f-a7cb-40dd-bd56-b021951e9c3b · inbound

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair cites this paper.

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:49.880043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:05:32.117271Z digest=sha256:3a5d4f697a540688745958182e540f3822e8c75baeb29b0817efb1d80f306e42

Observation cdd6732f-5408-437a-87e4-3ba143d72931 · inbound

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair cites this paper.

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.871897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:58:41.120953Z digest=sha256:b10f261d305859ef8ebb8ac743be0ae5cde817f3fe5660f8efe65e060bf39e54

Observation 9229fe0f-68cf-44af-b4c6-d7cc00c0e079 · inbound

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution cites this paper.

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:56.777009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:28:07.557119Z digest=sha256:5b976536c1fce9049c38f36b5dcb2098a37e75a11e13f248a27c5cca38d513dc

Observation 5d898c85-7161-4c56-9e36-4a9da96ff680 · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:26.109425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:47229dbdc35d1596ae2c141ea5f34bede56bd61f95f3803e5679c71f58565c8c

Observation cab455ad-1a6a-4a53-b9a2-d30927ee470b · inbound

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search cites this paper.

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.931493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:51:59.101930Z digest=sha256:07c3d4600f0dbc05956f924140e677f87a0651c7ae105a52061179d68cf105a5

Observation d9c2560a-8ebe-49ee-8a6d-6bb4b26e09fb · inbound

SmellBench: Towards Fine-Grained Evaluation of Code Agents on Refactoring Tasks cites this paper.

SmellBench: Towards Fine-Grained Evaluation of Code Agents on Refactoring Tasks RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:56:59.492628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T00:50:43.941064Z digest=sha256:c6f6cf682dc16d157bc10564206ce3b98b5382464816cba3df1e47760249d025

Observation 6c53dfeb-cd32-439a-ad2f-0b8bd088aaea · inbound

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation cites this paper.

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T05:40:26.231504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:40:26.231504Z digest=sha256:5db1d801cf4207bebe5009c9f453d38478aeea77e13468da67c9b5d40bc5f65e

Observation 87664a74-9708-45fa-8451-4e797bf6d067 · inbound

LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent cites this paper.

LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:46.635284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:46.635284Z digest=sha256:817914173ddd1b54ee925848230f438cd351fb745c5dffb6aa2c46d0305e9f3b