Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2607.01360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.01360 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-03T19:23:50.577675Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T10:52:17.791456Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-31T12:16:09.833992Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact7
  • verified fuzzy18
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b104bec-47e9-47f2-b16c-fad90c66df98 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Evaluating Large Language Models Trained on Code

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:28:51.853327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:d92559c07be686c75af571deb1fe722089a2a8d1bf8c8b47e53d80c414f59f3f

Observation 238fd308-0870-4db3-a60e-caeee6dcbfe5 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Livecodebench: Holistic and contamination free evaluation of large language models for code,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.474479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:2aac48a83fc59e2018aefc147c5b2b4f02220bac5dd5e21f918783d7881c6871

Observation 7181e20a-21d8-4fbc-bf2e-9af7d2a2b46a · outbound

This paper cites Available: https://openreview.net/forum?id=chfJJYC3iL.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Available: https://openreview.net/forum?id=chfJJYC3iL

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.472657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:1b94ff0e48637b3de54b0b86983055a78354288edf5cea79ad05fee40cae027c

Observation 9a61d9d4-3092-4501-97f7-6005e33af4d2 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues?.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback SWE-bench: Can language models resolve real-world github issues?

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.476465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:a65302b53d3e0320daa40bef4dd7a7d7372adcfaf2090c84958c444b2558c2cd

Observation 211008f7-c255-4e3f-938b-edde3869310e · outbound

This paper cites In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vienna, Austria) (ISSTA 2024).

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vienna, Austria) (ISSTA 2024)

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.391555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:db001d35acd680248b01e8afacb27f0b60f672bd6160cf84e2abc741d3b5c3b6

Observation 63672f72-25d8-4cc5-a4fb-1b04d029f994 · outbound

This paper cites Agentless: Demystifying llm-based software engineering agents,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Agentless: Demystifying llm-based software engineering agents,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.485329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:9de6d806cb7b80d0f59020359b5106a9ad1b618e398e6b6ebf26391a37be6012

Observation 944e2715-b456-45b5-90ab-049c8b511a2c · outbound

This paper cites K., Barr, E.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback K., Barr, E

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.396131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:bc25a0f96a46c7aae4cce8dcc4cc48b8025708671fa05f7f3302393cf778cf92

Observation 0525c222-92df-4326-a555-5af6b5b84fee · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.859271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:6e756344c8a16fce8333905ff409a3b7c2d8d7fa493e5c9fd3e98ebc425798b8

Observation f9d4c3f7-8477-4a86-ab49-eb5f2c3cf23c · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.387866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:25627daf794a661cb689c1e424ecf9d19712690f88e29c47bd8692ed8f95ae3c

Observation 91093192-025f-4f59-84e6-92ef1f35b9b0 · outbound

This paper cites Introducing SWE-bench verified,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Introducing SWE-bench verified,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.460584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:4d0a090104807d26405e7541ca51acf84519c6f6d70c6f75df7237f2782b9994

Observation 673c6e0f-2a8e-4be6-b917-52fb9ed00a85 · outbound

This paper cites Graph-based, self-supervised program repair from diagnostic feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Graph-based, self-supervised program repair from diagnostic feedback,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.467339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:8860fd9105b8e58b29c752fcb86ed120f131f67adf1cb7fc608c212d99065b43

Observation 5f356c19-650a-4d27-a5bf-362d4f1fb968 · outbound

This paper cites Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.856434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:058959e6b56e9560c302397bcbd82a2e44ed0e3bd49eeda90859faa2d0c562fe

Observation 433e75ac-23f1-4466-bd7c-27d7b5882ffd · outbound

This paper cites Self-edit: Fault-aware code editor for code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Self-edit: Fault-aware code editor for code generation,

Reference 13

Resolution
verified exact
doi, observed 2026-07-03T19:28:51.389944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:e8d3da0774e3e8593397ef7623deeec84c55ceb18cb4ba08b78ab52e21e32727

Observation 7f9ce8dc-90c4-49bd-b33e-477909ad5e02 · outbound

This paper cites Teaching large language models to self-debug,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Teaching large language models to self-debug,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.469176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:d03bd9f831a0b8aa9bafb6fb61106a03f7b7c481e5539782cbd822e0290e0198

Observation f80555c1-b40b-4d1a-aad1-05dade6dd969 · outbound

This paper cites Large Language Model Guided Self-Debugging Code Generation.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Large Language Model Guided Self-Debugging Code Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.382722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:81a23220c607d7cf9534c7ef6537a19229f252cc9688c803828fe8c8f96569dc

Observation eb871329-fe48-4fe4-b4e5-425382ddab16 · outbound

This paper cites ConvCodeWorld: Benchmarking conversational code generation in reproducible feedback environments,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback ConvCodeWorld: Benchmarking conversational code generation in reproducible feedback environments,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.465607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:992976ef11b4eddf599887e741e6e03827ff03c15d12aeb80b67a130ca61db97

Observation 86e9ecba-ede8-4cbf-b4f4-7a8224d1ddb4 · outbound

This paper cites When benchmarks talk: Re-evaluating code LLMs with interactive feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback When benchmarks talk: Re-evaluating code LLMs with interactive feedback,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.490698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:cd83172cbf68b1183e8b6b954f499b1654fe73dd3ef17a20fd676f3a2ac3458f

Observation 20ec7849-b1de-4a85-98c2-8431755ea605 · outbound

This paper cites Available: https://pairbench.site.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Available: https://pairbench.site

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.478192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:02ca129e5d3ef22f88ae70d2976116d21cbbba1a03367f3856f892c4bbff2064

Observation 77bb7840-4d74-452a-a413-580400cd71c8 · outbound

This paper cites Program Synthesis with Large Language Models.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Program Synthesis with Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:28:51.849550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:f3894a241a9e790e921224f5db29e3da448a21f2a6b45c1a343482448f803706

Observation d6f5e396-2f71-41c9-ae06-8e4e6cb56def · outbound

This paper cites InterCode: Stan- dardizing and benchmarking interactive coding with execution feed- back,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback InterCode: Stan- dardizing and benchmarking interactive coding with execution feed- back,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.479965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:0967e86416e74e392f443e451f245d96cbec3b606adac32dfb2a26fb83922407

Observation a305db9c-3781-4498-83f0-c12467b17b6c · outbound

This paper cites Measuring coding challenge competence with APPS,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Measuring coding challenge competence with APPS,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.492617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:d49b836eb865d5989cebb12fcd15123e614aa5e2addf6df30a62de6879b8e63f

Observation c5ab4027-6899-41ec-8c59-51ec2b5bdba4 · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.483590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:e4501ed9dbe5429d09c544d64ad079e248ffcfc286b80540f89b865aac183a36

Observation 26ca6827-d294-488c-b529-6aad35550b00 · outbound

This paper cites Evaluating language models for efficient code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Evaluating language models for efficient code generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.470869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:0856d72d13e0eb5041445ea9a6b1a12bbb4f559262756eb414c64d637f37fdfd

Observation 1e886f20-d2a3-4955-b8b6-c837745430af · outbound

This paper cites Ernst, Reid Holmes, and Gordon Fraser.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Ernst, Reid Holmes, and Gordon Fraser

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.399385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:47288a953aaa814ed97d6cc7063daa21ffca0d0f9dae08fea3691982b95c4ded

Observation 8e737cee-a16b-4f6d-aba8-7595810b137f · outbound

This paper cites Yiheng Xiong, Ting Su, Jue Wang, Jingling Sun, Geguang Pu, and Zhendong Su.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Yiheng Xiong, Ting Su, Jue Wang, Jingling Sun, Geguang Pu, and Zhendong Su

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.397826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:7afce4ef5abfd898c546e81488fb635d3199012b59359d945afa98e002631cad

Observation 516ffc33-0f2f-4f51-bfdf-fefeb65ed89e · outbound

This paper cites Quixbugs: A multi-lingual program repair benchmark set based on the quixey challenge.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Quixbugs: A multi-lingual program repair benchmark set based on the quixey challenge

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.392860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:4d4ba547bca928887719d7803ddea0a4cf991f545b4abe338172c361d2179738

Observation 41db9097-3227-4ddb-8186-1431dc74e1fa · outbound

This paper cites The power of feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback The power of feedback,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.462274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:565e7f87f62af52d984acc1fb87059583bffd394e556f28d2c6a3c96d7ce4645

Observation db73f802-0a32-4d85-8e48-d02ee76f06d7 · outbound

This paper cites Focus on formative feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Focus on formative feedback,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.488965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:dd75ec2e69b7bed3755813e240a53ada373966d641b82d4936c84898feab5dec

Observation ce7d2be9-fec2-4b9a-aa1d-144e0292bc4d · outbound

This paper cites The role of tutoring in problem solving.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback The role of tutoring in problem solving

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.463916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:d471062b6c737b3f9ff5c789886d970b490803c4d6230fe9fd66c677d12bf152

Observation 4d17eba8-5ab9-41e5-86ab-e89874eab233 · outbound

This paper cites an unresolved cited work.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-07-05T03:00:39.481731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:9c9c4ee0f8dc90018bea9c6a08d55745516c047646ebfe8d5db21064efb40398

Observation 2159dfc9-8ec1-479c-add0-a48e31fa5206 · outbound

This paper cites Codeforces-python-submissions,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Codeforces-python-submissions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.487111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:60c67ae604d09f94594676024377579ecee3cb7f03666b4b3c7e758c08caa397

Pith citing papers

Observation eac9c070-c0f6-4b28-a262-7a09080005d4 · inbound

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair cites this paper.

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-31T10:56:25.047648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-31T10:52:17.791456Z digest=sha256:13b8a8228d410ea40c8634302d273140299e99ee4177a3ab01423282e315525f