Pith. sign in

Paper Citation Record · LEDGER

Comparative Evaluation of Large Language Models for Test-Skeleton Generation

As of 23 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2509.04644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04644 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:59:44.334749Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a56c354b-4765-4186-9b04-75a0ca420820 · outbound

This paper cites • DeepSeek-Chat: A domain-specific model optimized for developer tasks, with enhanced per formance on programming-related queries and structural code outputs.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation • DeepSeek-Chat: A domain-specific model optimized for developer tasks, with enhanced per formance on programming-related queries and structural code outputs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.977040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:41.344770Z digest=sha256:ae904da97737c5dec9cfbd96ac6f899bde9b405551c9e68bfb1ce0b54e634fc3

Observation 1f18e828-bd31-46e6-8bd6-5b05e1cf722d · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.969241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:41.410462Z digest=sha256:102ff2af3dee0034437f3e6a1b973212b6b46f2ebbaa1b5f1be3ca95d1239736

Observation 48b3af5c-2e3f-45bc-8270-30cbed800f3d · outbound

This paper cites This metric captures how completely the model covered the methods of the class.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation This metric captures how completely the model covered the methods of the class

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.961628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:41.507627Z digest=sha256:d77beff16fffd3c42eb27069b3afe7af5ee7c705c42b2df541b7240bf6daf0ca

Observation 486d227c-6b63-4476-8e9e-092fa731a294 · outbound

This paper cites The reviewer assessed each generated skeleton using six dimensions, scoring on a 1–5 scale:.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation The reviewer assessed each generated skeleton using six dimensions, scoring on a 1–5 scale:

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.954259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:41.630098Z digest=sha256:bc8b8980038bbc458d44f9d6cf5393a4365a91f1c39d2cfed4b8b0110cfaee19

Observation 2aa9f794-00e1-49da-bd61-c0556f896be5 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.946549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:41.722768Z digest=sha256:f8de40ed555268b0a786a79d53f9b3dfc2325f449832332b38e65fa87ff767a8

Observation 23c1419b-ba06-44f3-89b2-18278065f3e9 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.938308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:41.807785Z digest=sha256:939c211d1cbbd4b1f811531327770dfbe6bfd74ed67848fd95fea095876e7ed6

Observation be35ffbd-2315-4430-b548-def0f3e66a13 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.931073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:41.881018Z digest=sha256:2befc68e99a4375300ff4507741ca4d22a537ce1140e89971eece57a8392428d

Observation 54b6d876-ec38-424e-a6c7-00f8b0975f9f · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.923339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:41.972822Z digest=sha256:6be57986c780b2207fcaf055051a8c8ea1e7982f6dc652fae68bde71d5e98396

Observation 7968bccd-2a55-494f-a818-0eeb245f31bb · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.915965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.063694Z digest=sha256:13238dec8712de5fa14cc4f4a39b8bd3f509197f3541c6e695bf11d8d3ff182c

Observation 3cdd46fe-230b-403c-81ec-48642f7b9676 · outbound

This paper cites The expert review focused on identifying semantic misinterpretations, structural flaws, and usability within the skeletons generated by the models.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation The expert review focused on identifying semantic misinterpretations, structural flaws, and usability within the skeletons generated by the models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.908762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.131373Z digest=sha256:04ec47615f92368b703413961ea184f6e1288a8479e2772ea5fcc2018263c75d

Observation b85c90f0-a531-4c50-8718-a442114582a2 · outbound

This paper cites This included appropriate usage of RSpec describe blocks and clear distinction between instance and class methods.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation This included appropriate usage of RSpec describe blocks and clear distinction between instance and class methods

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.900460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.194904Z digest=sha256:ceb93861cb29592535dc4bc910e264ea040467da136858ff25343c9623228287

Observation 0f6af92f-122c-406a-bab3-00aeff581380 · outbound

This paper cites Its test skeleton was praised for its well -organized structure and strong alignment with RSpec idioms.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Its test skeleton was praised for its well -organized structure and strong alignment with RSpec idioms

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.891842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.268192Z digest=sha256:be377582d1fcad3350567132087f3bf2fb336868e2f5ea1ab7bfbe03c584f837

Observation 98978fa1-f833-40e1-a54b-a47155aff6f8 · outbound

This paper cites Llama’s was the cleanest and easiest to maintain.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Llama’s was the cleanest and easiest to maintain

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.884101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.376257Z digest=sha256:eb6d9da2b44034ba0d4f6c7f8caafd217c4cdcc04f4d8aaa9b48729b90a5aa7f

Observation 3f890ebc-a1d5-4762-b228-9b6fe29c5b27 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.876794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.510525Z digest=sha256:c4490dc07425b7b8b3e934c7ecf7c53db8da8e742afd2fc09b4eb8e6aa97f921

Observation e4329897-e4db-4b47-9278-f75e5cae2c60 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.869064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.605560Z digest=sha256:1cc583dfed1a56f0ff2b281ad86a92eb751e98a4dc053d38f7b3ded9f9c77893

Observation 29b6aaa1-5873-40a0-91d1-64f80750467b · outbound

This paper cites Models that produced clean, readable code, like Llama4, were deemed more suitable for collaborative workflows.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Models that produced clean, readable code, like Llama4, were deemed more suitable for collaborative workflows

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.860914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.740858Z digest=sha256:619c0f4c860d68cc226017ef79675c0bac018cb0fc690ba91f656e96d07651f6

Observation e0e3f8a9-0755-4317-9ddf-74f8d609fb38 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.852855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.846810Z digest=sha256:ac1fe82f6dc4a22aa83bf079ae1df318d993dd82e9f2dd232322fc75992b3373

Observation 7a433ddd-0f37-47ab-bfae-26a6d7f05a1e · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.845038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:42.940550Z digest=sha256:15f39d65fee7530d18ef2b3d3a9ba746a6fe9a65fada2c9475bb2019bf5dd752

Observation 87812ec4-1834-44d7-8c1e-46d95612554a · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.837008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:43.062723Z digest=sha256:36f1cd6524f37f2cb8b9ac52b72caaf75078a25f869a23b412133a8400c17a08

Observation 862e439c-4c30-465b-8671-195159d70a0c · outbound

This paper cites While some models generated verbose skeletons with full coverage, these outputs were often dense, unstructured, or syntactically flawed.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation While some models generated verbose skeletons with full coverage, these outputs were often dense, unstructured, or syntactically flawed

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.793706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:43.203333Z digest=sha256:5962c0198fb257257e6de37e7e55ee52400e2376bacb3781c1dbf447a296c6c2

Observation 142814e7-915c-4348-9883-7f7ede077fad · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.673151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:43.295770Z digest=sha256:5e016ecd917bddb7805e838fb1787b19b7b13d61574e32f9c2e2656a904ac62b

Observation 88317820-8781-4781-877a-8906c4741ca8 · outbound

This paper cites Automation of Test Skeletons Within Test-Driven Development Projects,.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Automation of Test Skeletons Within Test-Driven Development Projects,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.381587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.381587Z digest=sha256:1906282ac26f072b0ddd9fcc5b20f135fd8e0fcac072f0aa0d275ff1454fd406

Observation c1ecfe73-34cb-4784-bcef-bebe459a42ed · outbound

This paper cites Software Testing with Large Language Models: Survey, Landscape, and Vision.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Software Testing with Large Language Models: Survey, Landscape, and Vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.507568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.507568Z digest=sha256:50b084084bd7158b63cb03efa01d110c30bceac3f93b32f94621bd19d4eed24e

Observation c57c1dd8-5430-40dc-aee1-fa33e674ced4 · outbound

This paper cites Large Language Models for Software Engineering: A Systematic Literature Review.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Large Language Models for Software Engineering: A Systematic Literature Review

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.512605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:43.653363Z digest=sha256:f8dc30270dcb510f79f959878b8c07d68dc59b09762c0112c20a599738dc91ad

Observation 6baeb417-25f0-4f31-bc11-249234aed318 · outbound

This paper cites Intelligent Software Testing: Harnessing Machine Learning to Automate Test Case Generation and Defect Prediction.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Intelligent Software Testing: Harnessing Machine Learning to Automate Test Case Generation and Defect Prediction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.371492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:43.741012Z digest=sha256:025d3506ad306ec66a0f86b496165e6717061b2176355693c58afef6b2063d56

Observation f1ef20a3-8da8-4d4d-8f12-4e68ed204cb4 · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.828934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.828934Z digest=sha256:88af33e5efc0a1103338e7b22e9e5eef7e7ac9c23b608f43b9283a742ff300d5

Observation 9c1dff1c-6bb4-492c-94a6-4b25e24507f4 · outbound

This paper cites Evaluating Large Language Models for Software Testing. Science of Computer Programming.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Evaluating Large Language Models for Software Testing. Science of Computer Programming

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.264082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:43.963435Z digest=sha256:9ed5d6dfc08b2cfd9f774021f76fbda3bbc8aea55df10af5cdf4115877cb328e

Observation b804f479-7dab-4b25-a69b-f7f736243ff1 · outbound

This paper cites EvoSuite: Automatic Test Suite Generation for Object- Oriented Software.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation EvoSuite: Automatic Test Suite Generation for Object- Oriented Software

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.091153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:44.085058Z digest=sha256:7487117f987ef03682e2ffb60fc57e2664d740e51460c321b515dfa741be74f9

Observation 158ed5fb-e8d8-41a8-ac15-c6a5a0a0941b · outbound

This paper cites CUTE: A Concolic Unit Testing Engine for C.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation CUTE: A Concolic Unit Testing Engine for C

Reference 30

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T05:59:44.786201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:44.229379Z digest=sha256:957e9ee14271615ea2a3a56a8f82b45f8ad33edd24eb31b4a53d3e63fbfbe9e5

Observation 780e796e-02ff-4c64-8b1d-df185da2c10d · outbound

This paper cites Improving Automated Test Case Generation with Learning-to-Rank Techniques.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Improving Automated Test Case Generation with Learning-to-Rank Techniques

Reference 31

Resolution
verified exact
doi, observed 2026-08-05T05:59:44.462803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:44.294754Z digest=sha256:c078b3067d68b4062dcd699cc5df9b0cd336da7d7e1887888cbc7992052d9a76

Observation 57918f8e-d878-48fc-9b2d-54d82ff63d7b · outbound

This paper cites Intrinsic Nonlinear Hall Detection of the N\'eel Vector for Two-Dimensional Antiferromagnetic Spintronics.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Intrinsic Nonlinear Hall Detection of the N\'eel Vector for Two-Dimensional Antiferromagnetic Spintronics

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T05:59:44.620488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T05:59:44.334749Z digest=sha256:4d840a0c2ee6b2d9c0289ce7bd1f9f654fc1034c005e6b6f1090e3a217229454

Observation eec26237-07d0-495b-a655-4b8974e6c9ba · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 419

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:44.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:44.163389Z digest=sha256:83f345c30ad022e438e5c06cf0e044ca203262523ccae2efc98afe3fae86597a

Pith citing papers

No inbound Pith citation observations are available.