Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2607.01360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.01360 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-03T19:23:50.577675Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T10:52:17.791456Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-31T12:16:09.833992Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact7
  • verified fuzzy18
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b104bec-47e9-47f2-b16c-fad90c66df98 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Evaluating Large Language Models Trained on Code

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:28:51.853327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:4e495abe08bba2c98ce01cc49ea8d7048aacb9eb03249ea80b474255e63cc5fd

Observation 238fd308-0870-4db3-a60e-caeee6dcbfe5 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Livecodebench: Holistic and contamination free evaluation of large language models for code,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.474479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:6b2f752d6ed74cfb5301180ad50a1d997b07b5b92898686cfca459946515ad80

Observation 7181e20a-21d8-4fbc-bf2e-9af7d2a2b46a · outbound

This paper cites Available: https://openreview.net/forum?id=chfJJYC3iL.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Available: https://openreview.net/forum?id=chfJJYC3iL

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.472657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:f3059873944e83ed5275ef4d2c366613af9bcabf6b1597adc6b8935e9807af59

Observation 9a61d9d4-3092-4501-97f7-6005e33af4d2 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues?.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback SWE-bench: Can language models resolve real-world github issues?

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.476465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:c8420ec49a0f0233f21962b58983352ef57056fa1c514cf2f00dfc2df40d3f3f

Observation 211008f7-c255-4e3f-938b-edde3869310e · outbound

This paper cites In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vienna, Austria) (ISSTA 2024).

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vienna, Austria) (ISSTA 2024)

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.391555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:1710a809e8d431c84ee1d8aeab995cf7583c0532a8c21004c9c2cf718051d4b1

Observation 63672f72-25d8-4cc5-a4fb-1b04d029f994 · outbound

This paper cites Agentless: Demystifying llm-based software engineering agents,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Agentless: Demystifying llm-based software engineering agents,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.485329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:5d41ff35943e88901ed59cef77fb956b9a9a1890cd4282b5a8ccb0ca88fe8d5e

Observation 944e2715-b456-45b5-90ab-049c8b511a2c · outbound

This paper cites K., Barr, E.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback K., Barr, E

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.396131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:85dc8d78ea384cc0ca37cc4af0aa7b59ad79c9b122974763d032f46c092a3453

Observation 0525c222-92df-4326-a555-5af6b5b84fee · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.859271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:b11596c5fef75e8773b4305cf35b988e7fb07e8dc5fd4cac95e4279983ef3646

Observation f9d4c3f7-8477-4a86-ab49-eb5f2c3cf23c · outbound

This paper cites Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.387866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:2c1d7ba2f4c6ca2bd180147e852a3cb6d09409deabad5ce3b6f7d0af5c2d6de4

Observation 91093192-025f-4f59-84e6-92ef1f35b9b0 · outbound

This paper cites Introducing SWE-bench verified,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Introducing SWE-bench verified,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.460584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:37557669daea822adf15f2f5c16c73eb29e1fa16691feb5576dbd2ce64960800

Observation 673c6e0f-2a8e-4be6-b917-52fb9ed00a85 · outbound

This paper cites Graph-based, self-supervised program repair from diagnostic feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Graph-based, self-supervised program repair from diagnostic feedback,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.467339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:f34bd098af545c1e14e8ae24d5307510a7329366e1f864fa9ecdc71762060a07

Observation 5f356c19-650a-4d27-a5bf-362d4f1fb968 · outbound

This paper cites Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.856434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:c68d9c8ccdfafdf2109118fc9a31cf5732b90b6c10889a56b60fca207627f5f5

Observation 433e75ac-23f1-4466-bd7c-27d7b5882ffd · outbound

This paper cites Self-edit: Fault-aware code editor for code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Self-edit: Fault-aware code editor for code generation,

Reference 13

Resolution
verified exact
doi, observed 2026-07-03T19:28:51.389944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:f1395c9a09800dba8e4d717e0823e841c84fb5192ddff884742c58a89024e459

Observation 7f9ce8dc-90c4-49bd-b33e-477909ad5e02 · outbound

This paper cites Teaching large language models to self-debug,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Teaching large language models to self-debug,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.469176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:8e85e9f25b5bbb14f6ae715f555b335cc7487f47b74830eb05402bfbdbceb164

Observation f80555c1-b40b-4d1a-aad1-05dade6dd969 · outbound

This paper cites Large Language Model Guided Self-Debugging Code Generation.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Large Language Model Guided Self-Debugging Code Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.382722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:d9bbb573bb25b9a6bd0eff6074a0f5e9c88cc1f19d046e67e492ea392942620c

Observation eb871329-fe48-4fe4-b4e5-425382ddab16 · outbound

This paper cites ConvCodeWorld: Benchmarking conversational code generation in reproducible feedback environments,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback ConvCodeWorld: Benchmarking conversational code generation in reproducible feedback environments,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.465607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:42c305f436f7902eb4d3423b6142c0704aff3cd842989af478ddea03c4c7d598

Observation 86e9ecba-ede8-4cbf-b4f4-7a8224d1ddb4 · outbound

This paper cites When benchmarks talk: Re-evaluating code LLMs with interactive feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback When benchmarks talk: Re-evaluating code LLMs with interactive feedback,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.490698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:02f67bdfa9b133bfcb8477b20a5dfabb2736c3ecdb455a0fea9f23ef53417192

Observation 20ec7849-b1de-4a85-98c2-8431755ea605 · outbound

This paper cites Available: https://pairbench.site.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Available: https://pairbench.site

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.478192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:256453f1c870c7ff8edab3cdb8db7b3c15dc76bf99009921672b52927fc1345a

Observation 77bb7840-4d74-452a-a413-580400cd71c8 · outbound

This paper cites Program Synthesis with Large Language Models.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Program Synthesis with Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:28:51.849550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:66251a23f9d77b4fad1f789e6aae59405a1c1e15956b00a4601e9ea897beca0c

Observation d6f5e396-2f71-41c9-ae06-8e4e6cb56def · outbound

This paper cites InterCode: Stan- dardizing and benchmarking interactive coding with execution feed- back,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback InterCode: Stan- dardizing and benchmarking interactive coding with execution feed- back,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.479965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:8a58773548af88e602e5f3e4be3d04e5965f9394500aa1555bd3d8c735a63bcf

Observation a305db9c-3781-4498-83f0-c12467b17b6c · outbound

This paper cites Measuring coding challenge competence with APPS,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Measuring coding challenge competence with APPS,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.492617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:5591d9e0e6d603913cac303eb1e729542ab522ed701587da59d0615b456007b2

Observation c5ab4027-6899-41ec-8c59-51ec2b5bdba4 · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.483590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:447be234c1473c2512b18086a55646f94c79b9c458d7167ae54b85da52c47210

Observation 26ca6827-d294-488c-b529-6aad35550b00 · outbound

This paper cites Evaluating language models for efficient code generation,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Evaluating language models for efficient code generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.470869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:bf9c4ee7d54d15327b90c0e4de22c026be46d7be02a9139f9433db2ae74c3253

Observation 1e886f20-d2a3-4955-b8b6-c837745430af · outbound

This paper cites Ernst, Reid Holmes, and Gordon Fraser.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Ernst, Reid Holmes, and Gordon Fraser

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.399385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:760e8670a65affd5f93e86dd6f08202ee265d5ee5d0719391c49d7c34f8547ed

Observation 8e737cee-a16b-4f6d-aba8-7595810b137f · outbound

This paper cites Yiheng Xiong, Ting Su, Jue Wang, Jingling Sun, Geguang Pu, and Zhendong Su.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Yiheng Xiong, Ting Su, Jue Wang, Jingling Sun, Geguang Pu, and Zhendong Su

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:51.397826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:32e03f147af04b1879ba6c590ad435d0b68d0be48e782a3acf07b95fb41f39d3

Observation 516ffc33-0f2f-4f51-bfdf-fefeb65ed89e · outbound

This paper cites Quixbugs: A multi-lingual program repair benchmark set based on the quixey challenge.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Quixbugs: A multi-lingual program repair benchmark set based on the quixey challenge

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:28:51.392860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:c6e9d96747b9c720b3c08d23af22f1319bd8060a410d95b69cafd8d0fbea1675

Observation 41db9097-3227-4ddb-8186-1431dc74e1fa · outbound

This paper cites The power of feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback The power of feedback,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.462274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:0f6821dc15f3d94df171484cd65af3ddb8ca6bf001b1ae45533ce4b7fe013955

Observation db73f802-0a32-4d85-8e48-d02ee76f06d7 · outbound

This paper cites Focus on formative feedback,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Focus on formative feedback,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.488965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:c64abf03a8557b4b9ef3493dd0c2e4f76034efcbd1100a54aad34e7f7616cb4d

Observation ce7d2be9-fec2-4b9a-aa1d-144e0292bc4d · outbound

This paper cites The role of tutoring in problem solving.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback The role of tutoring in problem solving

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.463916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:78d210949db0ae7bbedb0a15b49416d4823c93ee6746dc2c3b58184624015ebe

Observation 4d17eba8-5ab9-41e5-86ab-e89874eab233 · outbound

This paper cites an unresolved cited work.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-07-05T03:00:39.481731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:48e2aad440c9b9bc165247a02ef5b7fa997c91f98f75201a1f977a9df02d3917

Observation 2159dfc9-8ec1-479c-add0-a48e31fa5206 · outbound

This paper cites Codeforces-python-submissions,.

Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback Codeforces-python-submissions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T03:00:39.487111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T19:23:50.577675Z digest=sha256:ef871b2dcfaaf368cccfb396686556c5ffc178429e338a888a8799df7249857e

Pith citing papers

Observation eac9c070-c0f6-4b28-a262-7a09080005d4 · inbound

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair cites this paper.

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-31T10:56:25.047648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-31T10:52:17.791456Z digest=sha256:49b3f2ae437ebe9295f0c22dd714ef9ee38a84abdc47d61b5e10742e51846699