Pith. sign in

Paper Citation Record · LEDGER

Comparative Evaluation of Large Language Models for Test-Skeleton Generation

As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2509.04644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04644 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:59:44.334749Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a56c354b-4765-4186-9b04-75a0ca420820 · outbound

This paper cites • DeepSeek-Chat: A domain-specific model optimized for developer tasks, with enhanced per formance on programming-related queries and structural code outputs.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation • DeepSeek-Chat: A domain-specific model optimized for developer tasks, with enhanced per formance on programming-related queries and structural code outputs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.977040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:41.344770Z digest=sha256:69ef28132ab080bad02aa11fde39d50b9b49d58fe141c174a2fa423815979337

Observation 1f18e828-bd31-46e6-8bd6-5b05e1cf722d · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.969241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:41.410462Z digest=sha256:8b116ed2898b98449acc463febf47b01d30063589d89b7a50d4eda3ae1b8ec42

Observation 48b3af5c-2e3f-45bc-8270-30cbed800f3d · outbound

This paper cites This metric captures how completely the model covered the methods of the class.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation This metric captures how completely the model covered the methods of the class

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.961628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:41.507627Z digest=sha256:ce930193d9a65b228e9b51a23f0b55ca6cffa5ccdf5e563a239607fa37331f38

Observation 486d227c-6b63-4476-8e9e-092fa731a294 · outbound

This paper cites The reviewer assessed each generated skeleton using six dimensions, scoring on a 1–5 scale:.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation The reviewer assessed each generated skeleton using six dimensions, scoring on a 1–5 scale:

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.954259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:41.630098Z digest=sha256:5ccb22250f0a1c441ae85a8ddad678cc0a33a2814cccd241c87ff096481d65b8

Observation 2aa9f794-00e1-49da-bd61-c0556f896be5 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.946549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:41.722768Z digest=sha256:747055628e591b6dc9fc567d9ccbf6eed72a092441b21f6a8db4a9d28bd5d05e

Observation 23c1419b-ba06-44f3-89b2-18278065f3e9 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.938308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:41.807785Z digest=sha256:b3f0bf8386301852df322e55f4a9baff08e6db828e67168e9e3095b6938c1c03

Observation be35ffbd-2315-4430-b548-def0f3e66a13 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.931073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:41.881018Z digest=sha256:21b9ca943902caebf6d6f1cfe3ed721e2e04ebc450fbd495ab3fcbc98bf7a212

Observation 54b6d876-ec38-424e-a6c7-00f8b0975f9f · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.923339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:41.972822Z digest=sha256:c532e29f12fe909c15e9a5c877dc478fba5b18ee33f70f577a149f47c402ed81

Observation 7968bccd-2a55-494f-a818-0eeb245f31bb · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.915965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.063694Z digest=sha256:886b23f668945566b85d8c453d401b7f65d9c26186c5550f0642b117e1b62d03

Observation 3cdd46fe-230b-403c-81ec-48642f7b9676 · outbound

This paper cites The expert review focused on identifying semantic misinterpretations, structural flaws, and usability within the skeletons generated by the models.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation The expert review focused on identifying semantic misinterpretations, structural flaws, and usability within the skeletons generated by the models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.908762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.131373Z digest=sha256:500edd1d76da9ef53ae4a0665b7d761bc6a14bb2a7701d7facb4cb2f5f3feb43

Observation b85c90f0-a531-4c50-8718-a442114582a2 · outbound

This paper cites This included appropriate usage of RSpec describe blocks and clear distinction between instance and class methods.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation This included appropriate usage of RSpec describe blocks and clear distinction between instance and class methods

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.900460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.194904Z digest=sha256:78b2ddcb93a8014dac371038211c9aadf91211bfb0e8fd1332c21f872d8bde21

Observation 0f6af92f-122c-406a-bab3-00aeff581380 · outbound

This paper cites Its test skeleton was praised for its well -organized structure and strong alignment with RSpec idioms.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Its test skeleton was praised for its well -organized structure and strong alignment with RSpec idioms

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.891842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.268192Z digest=sha256:bfc3ffa97152371e2138a7600eedb8f31fde89b4aff535d9ca504e92f848c92a

Observation 98978fa1-f833-40e1-a54b-a47155aff6f8 · outbound

This paper cites Llama’s was the cleanest and easiest to maintain.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Llama’s was the cleanest and easiest to maintain

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.884101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.376257Z digest=sha256:9ba118e97f5c4cf817cf1cdd83e2fdede93f09e50f0d484121859818e8191060

Observation 3f890ebc-a1d5-4762-b228-9b6fe29c5b27 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.876794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.510525Z digest=sha256:6cee42482364a0e2485883fda8373b3b7a2ab20d1f1821dab0fc6f7e4812a028

Observation e4329897-e4db-4b47-9278-f75e5cae2c60 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.869064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.605560Z digest=sha256:84e439f8a500558d60e83f11e28696b70452ad32ba2f8a2c356c23f0cda3833c

Observation 29b6aaa1-5873-40a0-91d1-64f80750467b · outbound

This paper cites Models that produced clean, readable code, like Llama4, were deemed more suitable for collaborative workflows.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Models that produced clean, readable code, like Llama4, were deemed more suitable for collaborative workflows

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.860914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.740858Z digest=sha256:84dd5fec9ad39aaec8ee95b42299785d02bc33aabe330d95d37475746e83256a

Observation e0e3f8a9-0755-4317-9ddf-74f8d609fb38 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.852855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.846810Z digest=sha256:22e94bb908ef0b71ad4baf3f8419b8d2ba9bd852cf4534e792c9934927b3d9b3

Observation 7a433ddd-0f37-47ab-bfae-26a6d7f05a1e · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.845038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:42.940550Z digest=sha256:7840c47f8b3c1dda1cb8177eb589dd8988dbb52677ee624b8447f263b51bf240

Observation 87812ec4-1834-44d7-8c1e-46d95612554a · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.837008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:43.062723Z digest=sha256:f3b94c0ab7339d0bb86994984950c266253bf04d953a380606dc44aa4a9898fd

Observation 862e439c-4c30-465b-8671-195159d70a0c · outbound

This paper cites While some models generated verbose skeletons with full coverage, these outputs were often dense, unstructured, or syntactically flawed.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation While some models generated verbose skeletons with full coverage, these outputs were often dense, unstructured, or syntactically flawed

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.793706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:43.203333Z digest=sha256:29d2b3e43b3d3d5ad562a1257fe8486fa036382d14b162a4fec508e9dabd3f0d

Observation 142814e7-915c-4348-9883-7f7ede077fad · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.673151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:43.295770Z digest=sha256:9f5f4286a742da758ed796b36ab8410e4f5e04894e1755c91618470e11773ebd

Observation 88317820-8781-4781-877a-8906c4741ca8 · outbound

This paper cites Automation of Test Skeletons Within Test-Driven Development Projects,.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Automation of Test Skeletons Within Test-Driven Development Projects,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.381587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.381587Z digest=sha256:a5da078fe0d8c31b54fcae75b65fcdcb54707d455e2e2077cbe93073a13f6e8d

Observation c1ecfe73-34cb-4784-bcef-bebe459a42ed · outbound

This paper cites Software Testing with Large Language Models: Survey, Landscape, and Vision.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Software Testing with Large Language Models: Survey, Landscape, and Vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.507568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.507568Z digest=sha256:7a508b81b3d7ded572daa0958172d440b49897d81986217ca489651ea94db763

Observation c57c1dd8-5430-40dc-aee1-fa33e674ced4 · outbound

This paper cites Large Language Models for Software Engineering: A Systematic Literature Review.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Large Language Models for Software Engineering: A Systematic Literature Review

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.512605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:43.653363Z digest=sha256:16685b37ef93ae2c5136e03b4b8002d9353cc405ecdb25240b4daf83b2f4698e

Observation 6baeb417-25f0-4f31-bc11-249234aed318 · outbound

This paper cites Intelligent Software Testing: Harnessing Machine Learning to Automate Test Case Generation and Defect Prediction.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Intelligent Software Testing: Harnessing Machine Learning to Automate Test Case Generation and Defect Prediction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.371492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:43.741012Z digest=sha256:1019c9232e625a785808f662988813584da149c9760fef9bb56883ac3f88032d

Observation f1ef20a3-8da8-4d4d-8f12-4e68ed204cb4 · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.828934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.828934Z digest=sha256:c80b78d09da47925cfd01c43ed7d846ce1df70a9c899865662e92ae24d548b03

Observation 9c1dff1c-6bb4-492c-94a6-4b25e24507f4 · outbound

This paper cites Evaluating Large Language Models for Software Testing. Science of Computer Programming.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Evaluating Large Language Models for Software Testing. Science of Computer Programming

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.264082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:43.963435Z digest=sha256:e6c8aa7fa88e45d2693b4747790e3128b3a9296a3917b734d05ec11367b99f7d

Observation b804f479-7dab-4b25-a69b-f7f736243ff1 · outbound

This paper cites EvoSuite: Automatic Test Suite Generation for Object- Oriented Software.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation EvoSuite: Automatic Test Suite Generation for Object- Oriented Software

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.091153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:44.085058Z digest=sha256:402294a05e2a9a9ea07b88782718e3273c1e73b320fdb077ec04aa3de7b8a55a

Observation 158ed5fb-e8d8-41a8-ac15-c6a5a0a0941b · outbound

This paper cites CUTE: A Concolic Unit Testing Engine for C.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation CUTE: A Concolic Unit Testing Engine for C

Reference 30

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T05:59:44.786201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:44.229379Z digest=sha256:1dc948cd5df18eacf53c17e2d2690ab35d7c1c58762e0f4aeeeafa6579d2362f

Observation 780e796e-02ff-4c64-8b1d-df185da2c10d · outbound

This paper cites Improving Automated Test Case Generation with Learning-to-Rank Techniques.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Improving Automated Test Case Generation with Learning-to-Rank Techniques

Reference 31

Resolution
verified exact
doi, observed 2026-08-05T05:59:44.462803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:44.294754Z digest=sha256:f37325fd328307e83b9f46ff4fffdb1f2be7a0ba02323894204e8f8e8c22bea0

Observation 57918f8e-d878-48fc-9b2d-54d82ff63d7b · outbound

This paper cites Intrinsic Nonlinear Hall Detection of the N\'eel Vector for Two-Dimensional Antiferromagnetic Spintronics.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Intrinsic Nonlinear Hall Detection of the N\'eel Vector for Two-Dimensional Antiferromagnetic Spintronics

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T05:59:44.620488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T05:59:44.334749Z digest=sha256:ec342d57394901e7d670dbd6098defd7dbbb50fad5d5f91adb4c599f178657f9

Observation eec26237-07d0-495b-a655-4b8974e6c9ba · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 419

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:44.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:44.163389Z digest=sha256:a4e6d3f6f00a0fcede06b1f0c7b36806541863fdb06a051910f7ad83a1dad897

Pith citing papers

No inbound Pith citation observations are available.