Pith. sign in

Paper Citation Record · LEDGER

Addressing Data Leakage in HumanEval Using Combinatorial Test Design

As of 13 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.01526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01526 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:20:21.524087Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact6
  • verified fuzzy6
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36a47dbc-3bde-46d7-bf17-f69ce5f5ccef · outbound

This paper cites Large language models for software engineering: Sur- vey and open problems,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Large language models for software engineering: Sur- vey and open problems,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.759238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.426420Z digest=sha256:1db4146e95903fb05584d144c001b40129c966649a3f5085f57009837e330393

Observation 0dba7e71-d769-4c94-987a-01f9d7561e66 · outbound

This paper cites Software testing with large language models: Survey, landscape, and vision,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Software testing with large language models: Survey, landscape, and vision,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.431387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.431387Z digest=sha256:24941ae656b7fe8aca5233272ef91b49c43e872e36626985e7499b4526f27dfd

Observation 0012f093-1f51-4be2-94c7-f904e3847ed0 · outbound

This paper cites Large language models for software engineering: A systematic literature review,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Large language models for software engineering: A systematic literature review,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.435724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.435724Z digest=sha256:92a933c7b831098557f05ec313fa052aa92a7f26e5ca31f5e5f916e125afce4b

Observation f48b6488-b426-477f-bc63-905a50902254 · outbound

This paper cites A Survey on Evaluation of Large Language Models,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design A Survey on Evaluation of Large Language Models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.440549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.440549Z digest=sha256:a91d6a7a1bcaaf4bcf5ff8f9fa95786d0d509456020af70aa886e92149e96a07

Observation 068aba4a-d1a9-4192-8403-7350dcc76267 · outbound

This paper cites A Software Engineering Perspective on Testing Large Language Models: Research, Practice, Tools and Benchmarks.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design A Software Engineering Perspective on Testing Large Language Models: Research, Practice, Tools and Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.445014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.445014Z digest=sha256:0acc8019c7da450cb96dd2deed076e245f261a2e9f6e498780b67c7ecca2971b

Observation 4555590f-4a22-452e-84b2-8e682f5bb20f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.449441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.449441Z digest=sha256:2a53011d19895f465ad6ea18677894fcabec5248a468a75538584277debc1ecc

Observation 236712a5-0d82-495a-a02b-1859d596513c · outbound

This paper cites Codereval: A benchmark of pragmatic code generation with generative pre-trained models,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Codereval: A benchmark of pragmatic code generation with generative pre-trained models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.454253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.454253Z digest=sha256:3a1203308ad11570d613006334c6669a3b81299bec0857babaab47c5bcbf9328

Observation 676a281d-38e1-4338-a028-8d24e3328ef6 · outbound

This paper cites Towards standarized benchmarks of LLMs in software modeling tasks: a conceptual framework,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Towards standarized benchmarks of LLMs in software modeling tasks: a conceptual framework,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.464228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.464228Z digest=sha256:05757eb7495f58f7f7449b3d0c012c64defe213cda51141c2b0a73dceb58932a

Observation 2b3d364f-f575-4ba1-9924-bdbbe3c5d0ad · outbound

This paper cites Using benchmarking infrastructure to evaluate llm performance on cs concept inventories: Challenges, opportunities, and critiques,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Using benchmarking infrastructure to evaluate llm performance on cs concept inventories: Challenges, opportunities, and critiques,

Reference 9

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:22.463999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.468618Z digest=sha256:73373e666eb5e5b6af1da0a8c6acd1bab3295c3360d4e85beda294edff3cf953

Observation 149a2248-9216-4ffd-8f0b-ba279108ddea · outbound

This paper cites Condefects: A complementary dataset to address the data leakage concern for llm-based fault localization and program repair,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Condefects: A complementary dataset to address the data leakage concern for llm-based fault localization and program repair,

Reference 10

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:22.285490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.472665Z digest=sha256:59604c1dd99f95c0a85a9ed410d30a71e912b75a48dc02c110fceb82190258ed

Observation d67b6f95-605e-4ca9-836a-4789e64b2353 · outbound

This paper cites BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation,

Reference 11

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:22.096588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.476747Z digest=sha256:5bb4a2fa27eff553dae2a0c6eb66887e8cc5e0a1d885a0f4beb38e7c562774b2

Observation ffd2df8d-b59f-4695-9cc1-eb5b565deece · outbound

This paper cites Holistic Evaluation of Language Models.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Holistic Evaluation of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.483043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.483043Z digest=sha256:f8187751b0aa9a819ef4e109ad1bfc35d5235826844596174b171e5a34a77035

Observation 5e8413bb-b11a-41b9-b492-e82ad20b5500 · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.487565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.487565Z digest=sha256:e61e1554e620c8c35e12d8525679cfd475e8a00729f4657af840f1cf46f72174

Observation 06b07fee-2001-47b6-9caf-33dbff040a34 · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.730332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.491898Z digest=sha256:79664c8dbbae65a1ac7d1db40033185fe11cbfdd91faa3e843a6fcd6c1cce2cb

Observation d241f699-5b41-4b59-9583-b9b5290bfc30 · outbound

This paper cites Task Contamination: Language Models May Not Be Few-Shot Anymore,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Task Contamination: Language Models May Not Be Few-Shot Anymore,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.717836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.495785Z digest=sha256:23f8e9fd436ba4e2e0db40da629465f340964f463af24dc1ca35c0b3daba4f89

Observation 06bc2373-2751-4211-9c62-f2b9a7b2ea26 · outbound

This paper cites A test generation strategy for pairwise testing,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design A test generation strategy for pairwise testing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.705319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.499796Z digest=sha256:7f3e2b14de64f46a490923419d97c17ce729058f6f109aea108363b6003aac96

Observation 42dce237-fb1f-419d-beb3-d164c71fa7ce · outbound

This paper cites Combinatorial test design in practice,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Combinatorial test design in practice,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.692897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.504311Z digest=sha256:a3c66d90735bea23998b8d59a4120af14a1fb2ac69474d9cc4d58010b9da0552

Observation 9944ad66-216d-4120-8a31-196d1054daef · outbound

This paper cites Using benchmarking to advance research: a challenge to software engineering,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Using benchmarking to advance research: a challenge to software engineering,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-12T04:20:21.926313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.508384Z digest=sha256:005bade30588b44590adc9be56eb8097a6ab1cd579086e0cb9064dd1bb5f99ad

Observation babcd247-2587-4e4a-bdec-24a7767e85fd · outbound

This paper cites Using combinatorial benchmark construction to improve the assessment of concurrency bug detection tools,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Using combinatorial benchmark construction to improve the assessment of concurrency bug detection tools,

Reference 19

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:21.789763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.512271Z digest=sha256:9ffc0b7b246785620e6b05ffe3940833a7d3c838e041914eba472e87338989df

Observation 49c0fa33-9bfe-4b02-a60e-1b8da0d396f9 · outbound

This paper cites Applying pairwise combinatorial testing to large language model testing,.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Applying pairwise combinatorial testing to large language model testing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:20:22.679251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.516299Z digest=sha256:a249a2049180aafd954c0133d3f0c3c53b41d7435ec76b3ee76ca9940a0dc74f

Observation a62219f4-63e0-4c44-b634-98e80ad385fb · outbound

This paper cites Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.520097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.520097Z digest=sha256:b0ca93b79d7373196d2c7497935df98d9771f0bb00f7158a418aef3be3e61245

Observation 41996491-cd86-4db3-8569-cc51fdad9ac3 · outbound

This paper cites Dynamic Evaluation of Large Language Models by Meta Probing Agents.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:20:21.524087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:20:21.524087Z digest=sha256:86fca7f6cfbcc635543de840b4a720b18f0de6f2899dc2b5f0157d1571ed74cb

Observation 8e08a151-0be0-48f5-9029-1100fa38c204 · outbound

This paper cites Available: https://doi.org/10.1145/3597503.3623316.

Addressing Data Leakage in HumanEval Using Combinatorial Test Design Available: https://doi.org/10.1145/3597503.3623316

Reference 2024

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-12T04:20:22.642217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:20:21.459075Z digest=sha256:97197047adf2fd3a17d56fdfb87723e07a4297dcb286e77d390c42309a74f0f8

Pith citing papers

No inbound Pith citation observations are available.