Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny

As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.22370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22370 v4

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:20.375361Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4c4811e-c837-481f-9b17-92d39f489c0f · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:24.100340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.122363Z digest=sha256:5c025b05982f3a7816c989501d81922b3fba7b5a2be59e9fd2eab4831788b6eb

Observation 61c8ef18-3838-4605-b2eb-af93dbae9874 · outbound

This paper cites Do AI assistants help students write formal specifications? A study with ChatGPT and the B-Method.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Do AI assistants help students write formal specifications? A study with ChatGPT and the B-Method

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:10:20.557101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.146590Z digest=sha256:08f2cefd1118aa28d4fcaaa4ddf3a55096aec7418d62ab2cfcfe5ef581bf52fc

Observation 2b587b74-2640-4695-97a5-38c078019567 · outbound

This paper cites In: Proceedings of the 23rd ACM International Workshop on Formal Techniques for Java-like Programs.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 23rd ACM International Workshop on Formal Techniques for Java-like Programs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.900773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.210768Z digest=sha256:ed6659415e47146660a1353f138eb7c7bc4069d17dc980aab7fb4b9acafe80c8

Observation 54a65751-e22c-4207-afe0-22fc08e0a5ed · outbound

This paper cites Computers and Education: Artificial Intelligence7, 100290 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Computers and Education: Artificial Intelligence7, 100290 (2024)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.768504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.243840Z digest=sha256:9713430ca65200a4163d0853dd5e3607965b2628ff16370d00ef23a632b828ef

Observation b76a109a-0889-4aa4-9422-14de82073c15 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:23.588382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.290210Z digest=sha256:29b9a1aebf6a048ab3c68550e64fff473e85c291104ba6858637b8bcb8408320

Observation 4370b81c-38ba-4fd1-9bf5-a84598731bb6 · outbound

This paper cites Applied Sciences14(10), 4115 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Applied Sciences14(10), 4115 (2024)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.410110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.378703Z digest=sha256:418299250ff50a3b34ac51e9721faa75b0067a2dceabe075dc5e3d05fc2939f1

Observation c0889314-3feb-4250-8d20-554ca277e91f · outbound

This paper cites In: International conference on logic for programming artificial intelligence and reasoning.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International conference on logic for programming artificial intelligence and reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.451424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.451424Z digest=sha256:6b8865a84f227ff8578f1ec0b167164fa78f048fedb403639f35f7a67c3c7ff9

Observation 56e87d08-1a2d-41db-ba05-793a0798aa78 · outbound

This paper cites DafnyBench: A Benchmark for Formal Software Verification.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny DafnyBench: A Benchmark for Formal Software Verification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.488681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.488681Z digest=sha256:62735b31bff0fec39ec7da5c507004c31dd43b7f670ab5bbe0ef41f8e7424c2a

Observation 75f1cf6d-cdbf-4b22-aade-a14bb249f638 · outbound

This paper cites In: Proceedings of the 56th ACM Technical Symposium on Computer Science Education V.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 56th ACM Technical Symposium on Computer Science Education V

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:23.156920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.493600Z digest=sha256:94af206c1708477dea23482fac8383e762aa60b4845e19c2284eb21439997650

Observation 6073b738-31f2-4d1c-9042-a9549d87fde2 · outbound

This paper cites Assured Automatic Programming via Large Language Models.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Assured Automatic Programming via Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.502790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.502790Z digest=sha256:6c049ff76b9e88242b01203da07239d44f17583c7e9ee1e53a96156c95eb21da

Observation b3f59be7-b167-40a1-a408-850c86a0c474 · outbound

This paper cites Proceedings of the ACM on Software Engineering1(FSE), 812–835 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Proceedings of the ACM on Software Engineering1(FSE), 812–835 (2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.970377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.547781Z digest=sha256:046daa719c80e901658f047f399cde3dc16921e190869d36474e9fff422bf759

Observation 46ed1b5e-4f0f-4645-aa9d-219e279763e4 · outbound

This paper cites Proceedings of the ACM on Programming Languages9(OOPSLA1), 1519–1545 (2025).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Proceedings of the ACM on Programming Languages9(OOPSLA1), 1519–1545 (2025)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.803786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.614436Z digest=sha256:e823dcab9394b28ebe881856d5f380141d9df788195470676128ab5d81c38ee2

Observation 97463543-45bf-4e0f-bb42-b083cbb0bec9 · outbound

This paper cites In: NASA Formal Methods Symposium.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: NASA Formal Methods Symposium

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.635095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.712977Z digest=sha256:d11464c826316e7e00bef55175ee338be005c143ed226648304a2b9cbfe87c5b

Observation 68cd8d5a-b1e2-4995-9a80-49f1b69cee0c · outbound

This paper cites In: International Conference on Fundamentals of Software Engineering.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International Conference on Fundamentals of Software Engineering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.475039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:19.840113Z digest=sha256:8c3bffbcc7897bb8edc1755192e5cb5c343b888b1541c539f393b17f6e5f9543

Observation 77098d97-8aa7-4a23-9456-0ae2eec6f4d3 · outbound

This paper cites dafny-annotator: AI-Assisted Verification of Dafny Programs.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny dafny-annotator: AI-Assisted Verification of Dafny Programs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:19.945974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:19.945974Z digest=sha256:09e8b9cc497173251e946bfec20878e80cfbefa1cfa6fa4c892ef447114a27f0

Observation 6e4b1156-728a-455f-a3d4-d2cd697e98b6 · outbound

This paper cites In: Proceedings of the ACM Con- ference on Global Computing Education Vol 1.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the ACM Con- ference on Global Computing Education Vol 1

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.307291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.024034Z digest=sha256:2124acd5f5e4c7fb422b426c890e3aacbc8ba65ff593f42c4e97c5ca5d9c79c2

Observation 810191bd-857c-4856-9448-11c0964c1342 · outbound

This paper cites It’s weird that it knows what I want.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny It’s weird that it knows what I want

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.151705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.103786Z digest=sha256:a846e94164ebaefbee8c4be2ac7b03de9c904c4d65c24b836a5cda56d3331724

Observation 6ba07279-6064-41a6-9c1e-752edd9fbfa4 · outbound

This paper cites In: Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:22.026977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.159230Z digest=sha256:f44583319dccf01d1badf34fc4eae8461a3075be205010014924c8882d7bbb90

Observation e952621b-b581-44f9-b11e-b93d7decd9b3 · outbound

This paper cites Exploring the Use of ChatGPT as a Tool for Learning and Assessment in Undergraduate Computer Science Curriculum: Opportunities and Challenges.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Exploring the Use of ChatGPT as a Tool for Learning and Assessment in Undergraduate Computer Science Curriculum: Opportunities and Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.208512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.208512Z digest=sha256:3adc24753496c281e91f0a671e2198c3744928ea99624c61142a2d5d3d121deb

Observation 1c3fc999-e5e6-4e00-81fe-42c42ecf9885 · outbound

This paper cites an unresolved cited work.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:21.812585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.225656Z digest=sha256:9411db1f9334d9bc152e65df80cd555cb0a7818bf5dd819b0bc555cfde8bfbea

Observation ec95da06-b675-4a8f-a4f9-072de1460dcd · outbound

This paper cites In: Proceedings of the 2024 IEEE/ACM 12th International Conference on Formal Methods in Software Engineering (FormaliSE).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 2024 IEEE/ACM 12th International Conference on Formal Methods in Software Engineering (FormaliSE)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.643639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.255691Z digest=sha256:8776a471244ecbed25deff41587ca63bb8afa022296394afd109661ea96cfffe

Observation bf8e2275-0295-4037-b569-d1d6f8c44caa · outbound

This paper cites In: International Symposium on AI Verification.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: International Symposium on AI Verification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.420615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.273290Z digest=sha256:3326132fc083fc751d50e63ce14179bac2b3c0b975c34f2d5b8389c38e82ddd2

Observation 0b6970fd-49df-42d5-98c6-229f88c64605 · outbound

This paper cites International Journal of Educational Technology in Higher Education21(1), 14 (2024).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny International Journal of Educational Technology in Higher Education21(1), 14 (2024)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:21.182525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.292400Z digest=sha256:a548bce0ce002dca7c52ea24b2137c77b4a607ea0ed4678b2a359328bdf06624

Observation e6684581-ff6d-4104-8e4d-7a10a94d377d · outbound

This paper cites In: 23rd International Conference on Software Engineering and Formal Methods (SEFM) (2025).

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: 23rd International Conference on Software Engineering and Formal Methods (SEFM) (2025)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:20.989008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.317739Z digest=sha256:42b3aab83e5b68e4da6d7ac5143093963772238e55fbb18ab098ff7cbc0e50af

Observation f0e85018-93e9-4385-a92c-0cef8b8d3ce6 · outbound

This paper cites In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training.

Can Large Language Models Help Students Prove Software Correctness? An Experimental Study with Dafny In: Proceedings of the 46th International Conference on Software Engineering: Software Engineering Education and Training

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:20.839689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:10:20.375361Z digest=sha256:5a981699a975f6a237a754a65a49b9251837e07452589349918deeda90606751

Pith citing papers

No inbound Pith citation observations are available.