Pith. sign in

Paper Citation Record · LEDGER

TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2409.17561.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.17561 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:44:09.729270Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:07:15.220856Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 077b3721-7aaa-40b5-b744-ca9cdbd15e8e · inbound

A Large-scale Empirical Study on Fine-tuning Large Language Models for Unit Testing cites this paper.

A Large-scale Empirical Study on Fine-tuning Large Language Models for Unit Testing TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T10:27:18.108917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:27:18.108917Z digest=sha256:97f095ac861fd16ebe3ff2a0fe5f8d9fafa1ac1cc03074aa6bd3c8ff14c0c011

Observation 8be7cd65-24b6-47a7-ae15-563f1939239d · inbound

CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification cites this paper.

CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:26.107547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:26.107547Z digest=sha256:b167c6aab1e49b1f13dfdeb93a10cdeefcc938264816122e9629776318a67058

Observation 05460c55-7a88-42d1-a9f5-469868613ccd · inbound

HardTests: Synthesizing High-Quality Test Cases for LLM Coding cites this paper.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:58.370843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:58.370843Z digest=sha256:74e35f1cf5f16886305fa6aa562df4d57e4ea29a14f7341d241e6a16176c586e

Observation df8f2f9d-b5fb-4c36-93a6-44936b199a5e · inbound

Mutation-Guided Unit Test Generation with a Large Language Model cites this paper.

Mutation-Guided Unit Test Generation with a Large Language Model TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:07:15.222383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T11:05:15.898419Z digest=sha256:8ec5d9be29e8d5133decd596c5d7e58b00ada5ea5bc6abc00c1411fd3dabc0c6

Observation 232477cf-05a4-4641-ad9a-aeb4b312fac2 · inbound

SAGE:Specification-Aware Grammar Extraction for Automated Test Case Generation with LLMs cites this paper.

SAGE:Specification-Aware Grammar Extraction for Automated Test Case Generation with LLMs TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:46.160055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:46.160055Z digest=sha256:2a05de3ec438f66de91ddc8950f40ae639a888df5f410fa7fd9aece79f1e6d38

Observation e47c74e6-fa6e-453d-a0a9-88cff76ef09e · inbound

Large Language Models for Unit Testing: A Systematic Literature Review cites this paper.

Large Language Models for Unit Testing: A Systematic Literature Review TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-15T19:44:09.729270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:44:09.729270Z digest=sha256:f87b551490686f954990cb3954c70cd4114289fd9c5aa6b945e2aa0b6cc6b6ee

Observation 8bc6679c-be36-4654-8ab4-9c2bbf5c259b · inbound

PSearch: Search-based Patch Generation in the Era of LLM-based Automated Program Repair cites this paper.

PSearch: Search-based Patch Generation in the Era of LLM-based Automated Program Repair TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:44.819046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:44.819046Z digest=sha256:48444dc3eea294e89ee02d87d192119060b9651caa4348ec7fa2eb15e64d28f0

Observation cf2f71a3-4925-48f5-82d3-fce2de10f34b · inbound

Benchmarking LLMs for Unit Test Generation from Real-World Functions cites this paper.

Benchmarking LLMs for Unit Test Generation from Real-World Functions TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.245997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.245997Z digest=sha256:a03277c1f541c3f879771ceae1da2f31cb75456bbe0d1744a2f251b73b7ba835

Observation 4f82533b-509b-4629-81db-f10e666e8a85 · inbound

Call-Chain-Aware LLM-Based Test Generation for Java Projects cites this paper.

Call-Chain-Aware LLM-Based Test Generation for Java Projects TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:51:10.180283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T20:56:02.634264Z digest=sha256:c5f9f9fdc5ada944c3df04a6a2cc5ec8475bea9752b919e89e6fff81e24820e6

Observation dfc5257c-2e6a-4d11-be43-ef2c55e8144c · inbound

PPO guided Agentic Pipeline for Adaptive Prompt Selection and Test Case Generation cites this paper.

PPO guided Agentic Pipeline for Adaptive Prompt Selection and Test Case Generation TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:40.093414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T19:17:56.164183Z digest=sha256:9fdbf57c53140d3ee0478928bebe12523a9320078d3afbc548fd079f554dfb93

Observation 3ce5ed74-816e-4630-ba96-dce6a3070db8 · inbound

FeedbackLLM: Metadata driven Multi-Agentic Language Agnostic Test Case Generator with Evolving prompt and Coverage Feedback cites this paper.

FeedbackLLM: Metadata driven Multi-Agentic Language Agnostic Test Case Generator with Evolving prompt and Coverage Feedback TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:07.945625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T15:04:59.096221Z digest=sha256:dba676d75c1939cb29daceb79cc8f7f4aa514f7ee46a18b1b91addac90e30fdc

Observation f8335f12-4e0a-4900-b0a6-95a4163b995b · inbound

Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study) cites this paper.

Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study) TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T15:31:40.417126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:31:40.417126Z digest=sha256:ea00fbe8c8a0a38a7eb6f38b5c473d8f07579e274f51303713606e96da8a684c