Pith. sign in

Paper Citation Record · LEDGER

Specification Grounding Drives Test Effectiveness for LLM Code

As of 21 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.06636.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06636 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T00:51:00.099499Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact5
  • verified fuzzy17
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 18ced049-54cc-44bf-bc4b-dfa73f6dbd16 · outbound

This paper cites Program Synthesis with Large Language Models.

Specification Grounding Drives Test Effectiveness for LLM Code Program Synthesis with Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:57:42.229796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:f02f0a05f51dedbb995f42a98a107696920f3c51ba016279072ac9bc7f6c33ad

Observation 90c7f3e3-64d8-498e-b576-77bada4def34 · outbound

This paper cites Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo.

Specification Grounding Drives Test Effectiveness for LLM Code Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.742097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:fbaf19637e2ba62967d3fab3a6abd9203517c34ed893386985d8ce50b08bc854

Observation 807c3622-8771-43fa-a43e-2ec72e5c3cc3 · outbound

This paper cites CodeT : Code generation with generated tests.

Specification Grounding Drives Test Effectiveness for LLM Code CodeT : Code generation with generated tests

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.703171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:78afc1f70773cba6c3054703662dca2bf93cf4e3155a1f141dafbab78995522d

Observation fe755f8a-0428-42ba-813c-7b741e59448b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Specification Grounding Drives Test Effectiveness for LLM Code Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:57:42.209053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:5e644e93d4ba7564d932167cb6009304b135e81c76c47dbdcfa7deba2b534247

Observation a3ccc7b2-b11d-4d14-bf6c-7c921931fd59 · outbound

This paper cites Teaching large language models to self-debug.

Specification Grounding Drives Test Effectiveness for LLM Code Teaching large language models to self-debug

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.741578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:f8f4576f679b63aa8176ddfb85513cb380d1409a3788a4e3dc4a1d0d9765aa9d

Observation f1ea7651-03d9-4a73-9fb4-9dcaf485e75b · outbound

This paper cites QuickCheck : A lightweight tool for random testing of Haskell programs.

Specification Grounding Drives Test Effectiveness for LLM Code QuickCheck : A lightweight tool for random testing of Haskell programs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.698325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:ff2a618c03acaabcd1b7e6535ca2e408bf62d0e357e70b51d6b059e47fc383de

Observation a8227cb4-ac82-4173-b9ef-77f670e0dccb · outbound

This paper cites Selective classification for deep neural networks.

Specification Grounding Drives Test Effectiveness for LLM Code Selective classification for deep neural networks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.511296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:ea66efd7e0cfd236750d5690a4c2f9231508cc38a5e18d430505ccf0a8f51161

Observation 0fccd5ac-d18c-44ee-8079-a754ac592f49 · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Specification Grounding Drives Test Effectiveness for LLM Code Large language models cannot self-correct reasoning yet

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.762344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:4db90a53c85bb5a211f1ebd861bcaeb4f9f7c48217d4dba6ac89f5c23c081185

Observation ec7bdfc3-b19b-4762-86ec-ac971b864a55 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

Specification Grounding Drives Test Effectiveness for LLM Code Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.634363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:4f22d7c8f0071564215c486e106f36f624b88de47dbe1b720c046ebf07add952

Observation c6a1faa0-07e3-4cef-a5a1-8a6c36dcc17c · outbound

This paper cites Interactive Code Generation via Test-Driven User-Intent Formalization.

Specification Grounding Drives Test Effectiveness for LLM Code Interactive Code Generation via Test-Driven User-Intent Formalization

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T00:57:42.220834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:a116a56b6e76eecf8fc42d7c9542335afc63ee7870399b1e609aad50e80df777

Observation 2a9b4abe-73fb-48cb-a49b-66ffcca98aa0 · outbound

This paper cites Automated program repair.

Specification Grounding Drives Test Effectiveness for LLM Code Automated program repair

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.553040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:cce6236318012b8d67a7912b77ff06982489301fa87bb3eb65bfe39fdb3d76e7

Observation f269747e-b7cb-4d24-beff-d11f6b30e62a · outbound

This paper cites Rustan M.

Specification Grounding Drives Test Effectiveness for LLM Code Rustan M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.784593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:c5ca0dea1c72b815b6da6416f3a29d8d2bab0d5ba158b1622f9590c467f3c50c

Observation 3139d840-7ad2-42c9-ad15-c3628af7a3b6 · outbound

This paper cites Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation.

Specification Grounding Drives Test Effectiveness for LLM Code Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.634836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:946de6775b3fcb4d4e23ff11e100031c2df70277d9d7d3e0f508fb105c324fa4

Observation 6eed79ab-7505-46bd-b8ab-837fab07b990 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Specification Grounding Drives Test Effectiveness for LLM Code Self-refine: Iterative refinement with self-feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.720664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:ded083a8b6e3e1e745790bfd9a85772398f8cea8c4b95663daa1b5cb85d1faca

Observation 2ec58a2c-4985-4b4e-b380-d6db75aa6e2d · outbound

This paper cites Applying ``design by contract''.

Specification Grounding Drives Test Effectiveness for LLM Code Applying ``design by contract''

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.804671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:aaf09b52d61f1f76d58e946da181e987c8ba8a0ec8f82935e72d33fe9e0f4ffe

Observation e853258a-e462-4993-a42f-33d620cb5a45 · outbound

This paper cites Wang, and Xi Victoria Lin.

Specification Grounding Drives Test Effectiveness for LLM Code Wang, and Xi Victoria Lin

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.592151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:7474d818e0c8ee64fd840133e046c4958daab449b1ece171554a0923e8bbfb1b

Observation 4feb195b-af34-4dce-86cf-ab60336b8493 · outbound

This paper cites Code as Agent Harness.

Specification Grounding Drives Test Effectiveness for LLM Code Code as Agent Harness

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:57:42.254847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:7bf18356116f97a5ca360f567b5da3ddad5d4b778a7086e23bdd2d4357a91d2c

Observation 81295b76-5c62-4e3f-b36f-b14811dfe54e · outbound

This paper cites Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama.

Specification Grounding Drives Test Effectiveness for LLM Code Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.526191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:946445c00d8147f9eac658e4d5cdc253e5607813210a28b22bd37f6673c28fb4

Observation 6b10baaa-6987-44a3-8794-10622ec23b9a · outbound

This paper cites Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering.

Specification Grounding Drives Test Effectiveness for LLM Code Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:57:42.166859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:7243d58dfe2df3ad75be98fc7cc5182ed583ffd29ee41cd4c43d4a9157e68241

Observation c8778059-64eb-4452-a5a7-bd69056e1f59 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Specification Grounding Drives Test Effectiveness for LLM Code Code Llama: Open Foundation Models for Code

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:57:42.187173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:ec0d3836b1e6da689c7cba6fb658359794504e5398ec32abb9f4f15549f5381a

Observation 68c6b764-5532-43f4-beb9-f94e534058ed · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Specification Grounding Drives Test Effectiveness for LLM Code Reflexion: Language agents with verbal reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.656740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:3780ba6900719f1db943618a1482c192fd3616a5ffb3e680d93ef98596207e48

Observation d029648e-9111-4ca8-9ccf-0cac2dc62982 · outbound

This paper cites Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press.

Specification Grounding Drives Test Effectiveness for LLM Code Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.762741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:3cb409fdb7d1e8af815bb92ccb34ba753456f61cd65293ce77e8dd5eec0f1093

Observation 420a1065-235d-4494-83c1-cabb7bf434e4 · outbound

This paper cites Judging LLM -as-a-judge with MT -bench and chatbot arena.

Specification Grounding Drives Test Effectiveness for LLM Code Judging LLM -as-a-judge with MT -bench and chatbot arena

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T00:57:44.784278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-11T00:51:00.099499Z digest=sha256:ae33500eedb0444e5d41f4d2299a778275bea6dade35d62ac568cf2d089ce5fe

Pith citing papers

No inbound Pith citation observations are available.