Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback

As of 23 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2607.01360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.01360 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-03T19:23:50.577675Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T10:52:17.791456Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-31T12:16:09.833992Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact7
  • verified fuzzy18
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b104bec-47e9-47f2-b16c-fad90c66df98 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Evaluating Large Language Models Trained on Code

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:28:51.853327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:3e45e0fbd155345e8c80486cc28be1b91bad42864823b74258b99d7670e4f1ae

Observation 238fd308-0870-4db3-a60e-caeee6dcbfe5 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Livecodebench: Holistic and contamination free evaluation of large language models for code,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.474479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:4ad39f699de90b3519d743a0a2dbb4734fe0387cb8559b170dc14fb6cf214b83

Observation 7181e20a-21d8-4fbc-bf2e-9af7d2a2b46a · outbound

This paper cites Available: https://openreview.net/forum?id=chfJJYC3iL.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Available: https://openreview.net/forum?id=chfJJYC3iL

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.472657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:db7a617810c988cca7df1ea67300d4958dbf928decd10ca263655c9268e0e9ac

Observation 9a61d9d4-3092-4501-97f7-6005e33af4d2 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues?.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback SWE-bench: Can language models resolve real-world github issues?

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.476465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:52dded9b0acaa49432f6ca2863b2f610ae907826f8b38f8d7ee77457571d7102

Observation 211008f7-c255-4e3f-938b-edde3869310e · outbound

This paper cites In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vienna, Austria) (ISSTA 2024).

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vienna, Austria) (ISSTA 2024)

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.391555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:f27bbf41f1beba9c31279153b52e2d74352ce178812f2c74795c018aff4de222

Observation 63672f72-25d8-4cc5-a4fb-1b04d029f994 · outbound

This paper cites Agentless: Demystifying llm-based software engineering agents,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Agentless: Demystifying llm-based software engineering agents,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.485329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:4bb8ee99f53157886271e02e12c69b553cb8c31e67c37650e61cc9a3a216fb09

Observation 944e2715-b456-45b5-90ab-049c8b511a2c · outbound

This paper cites K., Barr, E.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback K., Barr, E

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.396131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:1161088341416f9a222f48bf7b32d99e0d694c6a35a884efb6f4a0d83b5442c1

Observation 0525c222-92df-4326-a555-5af6b5b84fee · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.859271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:5b840588417f9cc3044920224be9652cfa33e388c7227975a69acffcd253fcbf

Observation f9d4c3f7-8477-4a86-ab49-eb5f2c3cf23c · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.387866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:8b2e4d1af23d67b991712820e9b99225243e854e1de18d18efa01c6bc4d34922

Observation 91093192-025f-4f59-84e6-92ef1f35b9b0 · outbound

This paper cites Introducing SWE-bench verified,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Introducing SWE-bench verified,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.460584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:ebddeb05abd4bed5cf053a6fd90d8663923b9086aa487871db2b2e954f65dd64

Observation 673c6e0f-2a8e-4be6-b917-52fb9ed00a85 · outbound

This paper cites Graph-based, self-supervised program repair from diagnostic feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Graph-based, self-supervised program repair from diagnostic feedback,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.467339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:6e70fb90591d5399513d4848a4d49eb800c6709c557dd306edff71d442d80c8b

Observation 5f356c19-650a-4d27-a5bf-362d4f1fb968 · outbound

This paper cites Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.856434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:ba5097db726dda2c1f8eedf836704cc6e9016435dfb127a164cd84d1b6bb1fff

Observation 433e75ac-23f1-4466-bd7c-27d7b5882ffd · outbound

This paper cites Self-edit: Fault-aware code editor for code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Self-edit: Fault-aware code editor for code generation,

Reference 13

Resolution
verified exact
doi, observed 2026-07-03T19:28:51.389944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:9850aa7f5142dbb29cd313a2bf431494e2ca794db28a3f133b007e4302098add

Observation 7f9ce8dc-90c4-49bd-b33e-477909ad5e02 · outbound

This paper cites Teaching large language models to self-debug,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Teaching large language models to self-debug,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.469176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:c70593cd70b472b670e3ebb18c1a4f43a20ce5787535f8f787d8c16772d82190

Observation f80555c1-b40b-4d1a-aad1-05dade6dd969 · outbound

This paper cites Large Language Model Guided Self-Debugging Code Generation.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Large Language Model Guided Self-Debugging Code Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.382722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:e1e77990f09232969bd3269c69d3fcf25ff131678354f4d338bce331f7cd0463

Observation eb871329-fe48-4fe4-b4e5-425382ddab16 · outbound

This paper cites ConvCodeWorld: Benchmarking conversational code generation in reproducible feedback environments,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback ConvCodeWorld: Benchmarking conversational code generation in reproducible feedback environments,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.465607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:269d71b6ac5db79fdb5ed1e406b474cbf0b19e39e82ab04478c558a5e3d7f2fd

Observation 86e9ecba-ede8-4cbf-b4f4-7a8224d1ddb4 · outbound

This paper cites When benchmarks talk: Re-evaluating code LLMs with interactive feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback When benchmarks talk: Re-evaluating code LLMs with interactive feedback,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.490698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:da59453a74ac9f8911ff1cef39fdd693e8b14de76c0138ed83f8d4a8c84d68b5

Observation 20ec7849-b1de-4a85-98c2-8431755ea605 · outbound

This paper cites Available: https://pairbench.site.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Available: https://pairbench.site

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.478192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:3aa2ba736ddc9b4ea984b77ca798a37ce4bc14c077791a398049c3e9f9d14d2e

Observation 77bb7840-4d74-452a-a413-580400cd71c8 · outbound

This paper cites Program Synthesis with Large Language Models.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Program Synthesis with Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:28:51.849550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:1794709599dcf918488724c1852455bd942bd6ed229161c366364e82c9725828

Observation d6f5e396-2f71-41c9-ae06-8e4e6cb56def · outbound

This paper cites InterCode: Stan- dardizing and benchmarking interactive coding with execution feed- back,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback InterCode: Stan- dardizing and benchmarking interactive coding with execution feed- back,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.479965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:1c5b4dda49f60897388eb75c8d9c3c04a5ca7a0b609f5823ca3ee54e4182076c

Observation a305db9c-3781-4498-83f0-c12467b17b6c · outbound

This paper cites Measuring coding challenge competence with APPS,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Measuring coding challenge competence with APPS,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.492617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:5f4c5a0ba981cdca11eb925b96c3fe24a21c8ba9f568bcd9c1a4329ce0b8f9f2

Observation c5ab4027-6899-41ec-8c59-51ec2b5bdba4 · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.483590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:b4b9c61373268d862dcaa419e162233a6886adf075d1860d0ee907e06e9a18d1

Observation 26ca6827-d294-488c-b529-6aad35550b00 · outbound

This paper cites Evaluating language models for efficient code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Evaluating language models for efficient code generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.470869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:ec61e4b6ea83c23d4e309b1addea75c37ca06e5468708234098596318f7e93bb

Observation 1e886f20-d2a3-4955-b8b6-c837745430af · outbound

This paper cites Ernst, Reid Holmes, and Gordon Fraser.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Ernst, Reid Holmes, and Gordon Fraser

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.399385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:ef8104d1797fe89d822399b677332b3515adfab04b8903cb19a367c58f777f89

Observation 8e737cee-a16b-4f6d-aba8-7595810b137f · outbound

This paper cites Yiheng Xiong, Ting Su, Jue Wang, Jingling Sun, Geguang Pu, and Zhendong Su.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Yiheng Xiong, Ting Su, Jue Wang, Jingling Sun, Geguang Pu, and Zhendong Su

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.397826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:e07e63993876bad385e617fa024d6f644fe1320a2817d0ee940260b7a6f2b6b3

Observation 516ffc33-0f2f-4f51-bfdf-fefeb65ed89e · outbound

This paper cites Quixbugs: A multi-lingual program repair benchmark set based on the quixey challenge.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Quixbugs: A multi-lingual program repair benchmark set based on the quixey challenge

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.392860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:bffcd531b8e88ef9ce14e18f2a8c870bb0f40b68ef7c72bd6b8f1ebfd42bd0fc

Observation 41db9097-3227-4ddb-8186-1431dc74e1fa · outbound

This paper cites The power of feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback The power of feedback,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.462274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:c2ce01a0b9ce31a0426e73c4ae4c170b4a6182123c847beeb4e74cab10135e46

Observation db73f802-0a32-4d85-8e48-d02ee76f06d7 · outbound

This paper cites Focus on formative feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Focus on formative feedback,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.488965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:86939a25c8df866b1af95238cdfac501b6fd641e9fbc50b25dbd8d541eb0a094

Observation ce7d2be9-fec2-4b9a-aa1d-144e0292bc4d · outbound

This paper cites The role of tutoring in problem solving.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback The role of tutoring in problem solving

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.463916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:edc9a92aad0dce41b4e0c4c07274e3700d1380847dbf81937287ceee62a58b46

Observation 4d17eba8-5ab9-41e5-86ab-e89874eab233 · outbound

This paper cites an unresolved cited work.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-07-05T03:00:39.481731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:0d324fdd2eedbfb75fd4f1cefbb00a23df3e25e4d96dacbb2096b555af63f484

Observation 2159dfc9-8ec1-479c-add0-a48e31fa5206 · outbound

This paper cites Codeforces-python-submissions,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Codeforces-python-submissions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.487111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:dbabfc7f7914c4f780260bf3dff5783cbf8216f50f7f1c98e146c21cec5fafe7

Pith citing papers

Observation eac9c070-c0f6-4b28-a262-7a09080005d4 · inbound

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair cites this paper.

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-31T10:56:25.047648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-31T10:52:17.791456Z digest=sha256:a113f8f33c84d63de674e148ae50e0a0eac40b6ae1dcc1f4c5bfe34ad10605ac