Pith. sign in

Paper Citation Record · LEDGER

Disproving Program Equivalence with LLMs

As of 12 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2502.18473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18473 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:17.861186Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T19:45:34.338092Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a60a836c-9f6b-45d8-a15f-749b4247e61e · outbound

This paper cites Automated unit test improvement using large language models at M eta.

Disproving Program Equivalence with LLMs Automated unit test improvement using large language models at M eta

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:18.105514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.791370Z digest=sha256:dae3b609ecd439fc5feb7625b8597ed5d19c621de59a8b4ac1e6a0aca8272874

Observation 9948a37f-d826-4b94-b4fe-3937186e3c16 · outbound

This paper cites CodeT: Code Generation with Generated Tests.

Disproving Program Equivalence with LLMs CodeT: Code Generation with Generated Tests

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.795544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.795544Z digest=sha256:56fb14f99fd7ba8e105814897988c1cc14158a27978ffeeada6c27ab052b75a9

Observation 3074b179-cd43-42c5-aaa0-de9d0299692e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Disproving Program Equivalence with LLMs Evaluating Large Language Models Trained on Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.799348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.799348Z digest=sha256:0ecfc70f08c7aa253af5ad4980348d926e95f8eb80926a4f604a560a2d2cba6e

Observation 7550de0f-eae8-4dde-9d62-c0284217b314 · outbound

This paper cites Universal Self-Consistency for Large Language Model Generation.

Disproving Program Equivalence with LLMs Universal Self-Consistency for Large Language Model Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.803254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.803254Z digest=sha256:a28e9a0108087223caffa5b06f0d5b167fdf621393f491b72932111fae47bc3e

Observation f19c978a-e492-446c-bace-d01fe1e58f6c · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Disproving Program Equivalence with LLMs Teaching Large Language Models to Self-Debug

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.808996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.808996Z digest=sha256:92c44e564f86e13e0fe323861807bb7ae534bf2c6c5cf8e8abe0136c5ae2be6a

Observation 57993b9c-9aec-4d5a-b4cb-5084f327ae06 · outbound

This paper cites Automated testing of refactoring engines.

Disproving Program Equivalence with LLMs Automated testing of refactoring engines

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:18.093908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.812924Z digest=sha256:a4c530e902bd13f8b9e028a5d056a3d016f730f69d372c3f42eea4f1942eb83c

Observation 66f442b8-373a-4504-9292-1a66a89bfd07 · outbound

This paper cites SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning.

Disproving Program Equivalence with LLMs SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.816673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.816673Z digest=sha256:e72d3230f9302ff29e15a1993858b01cb4a34d0df2e7912666f99de924106138

Observation b8585ea0-7c3c-4c6a-ab67-6de2df08abea · outbound

This paper cites Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?.

Disproving Program Equivalence with LLMs Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.820677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.820677Z digest=sha256:4abd18bca663019138670d04622ad8f1267eb8ad1d8aaf71387cc7f6d6324ed8

Observation 7ce956eb-9b39-49a3-82c4-f5d1e9c5d5a2 · outbound

This paper cites and Bishop, M.

Disproving Program Equivalence with LLMs and Bishop, M

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:18.082704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.824780Z digest=sha256:ac4779181f8e32796175246a7270c3676cedc632e2b916948202b60a9fca0926

Observation e759eb1e-9ffe-415a-b295-11203c7957f3 · outbound

This paper cites ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution.

Disproving Program Equivalence with LLMs ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-09T05:54:17.922488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.828483Z digest=sha256:8d9eb592f60b16867d815e3730d5c7608119f74ec35038f7f8b82aad21a6ab67

Observation 6916095f-6a28-4427-bd0c-bbea6779177c · outbound

This paper cites Self-supervised learning to prove equivalence between straight-line programs via rewrite rules.

Disproving Program Equivalence with LLMs Self-supervised learning to prove equivalence between straight-line programs via rewrite rules

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:18.071092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.832080Z digest=sha256:4fa0a230d0109d244e31983f581b484d01e86b19122a2785db49d0f2cb14f7d3

Observation ebaa91f4-2df0-46f8-85e1-5ed89425a599 · outbound

This paper cites Compiler validation via equivalence modulo inputs.

Disproving Program Equivalence with LLMs Compiler validation via equivalence modulo inputs

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:18.058832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.835729Z digest=sha256:b46ad911021f8ab94b38d4d9f25c154b2b22966326a41abc469d4ebdbb337a59

Observation f4f586be-af10-4f90-a88d-03924d136ac5 · outbound

This paper cites P., Lahiri, S.

Disproving Program Equivalence with LLMs P., Lahiri, S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:18.048023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.839581Z digest=sha256:17cf6061e6d46270b1559b7e66443e7c28c8f7583128132a8461555dd38a66df

Observation 1307cb70-c637-4db1-bc84-0ca6e6045bf2 · outbound

This paper cites S., Wang, Y., and Zhang, L.

Disproving Program Equivalence with LLMs S., Wang, Y., and Zhang, L

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:18.037283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.843035Z digest=sha256:c5afb6f878bf4131bcba0085490c8d4edb6a87a8f2de139b73116d188d8b0c84

Observation e008c9ec-d0dd-42a5-b23d-e94421bec760 · outbound

This paper cites On Leakage of Code Generation Evaluation Datasets.

Disproving Program Equivalence with LLMs On Leakage of Code Generation Evaluation Datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.846372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.846372Z digest=sha256:dcae352591a2e170bd91aa815a2108e7b0413c2ba8d177edcc6008ce282aaa9b

Observation 30f4f378-8d72-497d-8669-895e95d57930 · outbound

This paper cites User interaction models for disambiguation in programming by example.

Disproving Program Equivalence with LLMs User interaction models for disambiguation in programming by example

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:18.027178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.850052Z digest=sha256:b5e0fd15e22c48afbff4412d339e2c8e8e32f70ce18f0260b5320b21b18f5c14

Observation 2595e9de-c4ec-4a1e-bf81-784122c6127c · outbound

This paper cites an unresolved cited work.

Disproving Program Equivalence with LLMs Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T05:54:18.015326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T05:54:17.853618Z digest=sha256:711f54b50029344f69383a03769eb3cc85a35f55506a938628b81bd06d027f82

Observation 3a77ba75-7c70-4189-9b3f-21a534de5227 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Disproving Program Equivalence with LLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.857319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.857319Z digest=sha256:8d08b6ef8c0fc48ad35d9e609f2e044d79b6fb28adfdf3eaed4df70c991cd9b6

Observation 3cbcfe10-2de1-4f59-8ab0-29ba1875fc34 · outbound

This paper cites write newline.

Disproving Program Equivalence with LLMs write newline

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:17.861186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:17.861186Z digest=sha256:5349a1bb244880747266cece5a9138622fb71791e79ecc326665f1981b3aab65

Pith citing papers

Observation 42bc0641-7f62-491c-bc91-cc70d2faa0d8 · inbound

How Robustly do LLMs Understand Execution Semantics? cites this paper.

How Robustly do LLMs Understand Execution Semantics? Disproving Program Equivalence with LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:46:32.914616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T19:45:34.338092Z digest=sha256:6f08ea5886f8a2061fe2a2c603f323e02d2fb8b1fd8fd670f90dd49ef3866119