Pith. sign in

Paper Citation Record · LEDGER

Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2409.14781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.14781 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:13:35.269339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T20:57:23.314837Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0a93cd73-86ba-4d1e-9729-1e33082a474d · inbound

Investigating the Impact of Data Selection Strategies on Language Model Performance cites this paper.

Investigating the Impact of Data Selection Strategies on Language Model Performance Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.286703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.286703Z digest=sha256:b191cd629aca322e69eebc8db9b39a121df9bba744d36379642f565ded194062

Observation fd847f73-7102-49d5-ac6f-f371f12ea372 · inbound

STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings cites this paper.

STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:35.269339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:13:35.269339Z digest=sha256:86f749812f1d91defabcf2257321892bbfc051b54df8dd5ec9fa49cadcb43362

Observation 49876dfe-dece-44e3-b0c7-4a02b78e18f9 · inbound

Automatic Calibration for Membership Inference Attack on Large Language Models cites this paper.

Automatic Calibration for Membership Inference Attack on Large Language Models Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-15T23:59:23.388368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:59:23.388368Z digest=sha256:3ab1b18558e0586088d274460a14906eac8e037beae235a6dbb337ca6b80bded

Observation 9e7b95e0-4ba4-4d80-8a4b-3bf523483d1a · inbound

Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering cites this paper.

Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.161581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.161581Z digest=sha256:e795baba841f789d71e342d052d7cf9bf3733428340edfcc11fc216a1950e6a2

Observation 637cefb0-2249-407d-885f-c228f6cf9551 · inbound

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework cites this paper.

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T15:15:21.454509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:15:21.454509Z digest=sha256:ef40fe5881482c51bce119ba0cff97d6e6bcc0f92f71ebc673f617222fbe8354

Observation 7fc4b8d2-cf88-4f0a-b520-a16a19c6c187 · inbound

Investigating Training Data Detection in AI Coders cites this paper.

Investigating Training Data Detection in AI Coders Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.845119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.845119Z digest=sha256:e750d9bef5d3a442e3cf5117a0991eb9cdff7ae9820b3135ae5d21ae357faeb9

Observation fe8683d4-a5ee-4a1b-8324-728973be40e7 · inbound

Filling the Gaps: Selective Knowledge Augmentation for LLM Recommenders cites this paper.

Filling the Gaps: Selective Knowledge Augmentation for LLM Recommenders Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:31:01.654876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:36:36.680191Z digest=sha256:52b0f71a2496606594e8f9a373b518de7c946c9d1cd07966a796970a22251dbc

Observation c8181a8a-0d7b-4f30-96d3-61fca2a7470a · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:23.317104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:0977255b4b6b2a57645029e5702b75be42061f574cdb6fdebdf004aacbf70040