Pith. sign in

Paper Citation Record · LEDGER

Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2404.02936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.02936 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:11:35.067065Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f43d9dd4-aadf-4745-abd3-8350063f0700 · inbound

How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence cites this paper.

How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T18:11:35.067065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:11:35.067065Z digest=sha256:c1c6da71e4065a73996df16c2d07fe7d2607b777e7282133b04bc4363bf5e7bc

Observation b20cd380-7302-4846-bcfb-40d6757d1a3e · inbound

DocMIA: Document-Level Membership Inference Attacks against DocVQA Models cites this paper.

DocMIA: Document-Level Membership Inference Attacks against DocVQA Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T04:11:59.861601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:11:59.861601Z digest=sha256:a46ce985d79900a2fb22ba7de362dd45d3695cb3f86ee585e475e4e3c9c47efd

Observation 5207c1a4-dda9-4edf-9908-bf130e0cc506 · inbound

DUSK: Do Not Unlearn Shared Knowledge cites this paper.

DUSK: Do Not Unlearn Shared Knowledge Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:25.147132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:25.147132Z digest=sha256:5a17623c847b3b8790317282549a72ca3b956231c5c0c72440156cd54add0738

Observation 548de2d1-4350-4c21-b212-3e1c67f612a8 · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:06.177607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:06.177607Z digest=sha256:080349286fd02b4bc423cbd5b136c12e01590ab2ddfb97a144106e132d611b7c

Observation 1d64ba19-d40c-4db7-b05e-9791aa2bd424 · inbound

Membership Inference Attacks on Sequence Models cites this paper.

Membership Inference Attacks on Sequence Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:30.967563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:30.967563Z digest=sha256:2dbcfa05dd4f27aeabb99ce2d1fde572a97d95109b87e35757df6dd40e6bae7f

Observation 54e7395b-2eba-4243-8bec-cb458cf5e60b · inbound

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems cites this paper.

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:16.451624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:16.451624Z digest=sha256:26641d3005d93e87c7b8c830466e39dbdd4e65bfb5008d22907abb8e975868d9

Observation 8c09842f-e619-420a-9d09-bbaf013e431f · inbound

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks cites this paper.

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:17.063199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:17.063199Z digest=sha256:9407d92d870ff6ba131b076c60f73afc8cf4410e44ec1f32bb026fab5e98dac8

Observation 27c22f6c-8553-4cf2-9002-a8896860bd67 · inbound

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework cites this paper.

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:15:21.428831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:15:21.428831Z digest=sha256:b92304e4ef211e7b19a0ff74b6db146a02ef4b3b79af65742e864c7ff6c1a62c

Observation 34c00687-a372-46be-8ae7-bcb5347cb0e3 · inbound

Investigating Training Data Detection in AI Coders cites this paper.

Investigating Training Data Detection in AI Coders Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.839169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.839169Z digest=sha256:ea102409f6b50aa446b379b73fc17177109511c687f38680f230d8ef2a5cce5a

Observation 034fc69a-6222-4535-ab99-2cf42a1c8c0d · inbound

Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis cites this paper.

Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:29:27.157115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:29:27.157115Z digest=sha256:b71c5c5d1a92b6ea22a1d46d7ee05d2cda73d4636878d017091e521c392c95e7

Observation 8052ae0b-864b-4d1b-a8ad-db208d0d0cdc · inbound

Membership Inference Attacks on Tokenizers of Large Language Models cites this paper.

Membership Inference Attacks on Tokenizers of Large Language Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:22.454938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:22.454938Z digest=sha256:8afc8d91c21969105347d9cc92028f98c0936b45faf01417063de5ccf24e15c0

Observation 36584a79-e6c1-4784-b819-5ccbfe15c863 · inbound

LLM generation novelty through the lens of semantic similarity cites this paper.

LLM generation novelty through the lens of semantic similarity Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.311136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.311136Z digest=sha256:559f9d1d89ca932a2ea5c6c23af19c06415cd2b6ee6f3dbe7d673948bf844fe0

Observation 625fbecc-31a9-4d15-8ad8-89a18e076274 · inbound

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards cites this paper.

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:40:17.766089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:37:54.702010Z digest=sha256:0f539320aeace5513625c1d915df328f2bb1faaee34523da09840df120d55c77

Observation ceca5c65-e301-4a5f-a529-51ac3e4385ff · inbound

Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Models cites this paper.

Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:44:50.219620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:42:18.076004Z digest=sha256:e1beb786e73fa9dbc011d513986a9d48846bfb01580220ddef6f97f692a7e3cd

Observation a7adb287-207d-4551-b810-bc240ae08fef · inbound

Learning the Signature of Memorization in Autoregressive Language Models cites this paper.

Learning the Signature of Memorization in Autoregressive Language Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:53:11.478145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T19:53:10.396785Z digest=sha256:a3ba86900ea7f3d5f4a944e68fcc69aa6eb2e73bcc5d866f882ce07a4cc15d7f

Observation df7fb1b4-698c-4156-b89d-f663c5452565 · inbound

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs cites this paper.

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.938818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T15:32:05.976335Z digest=sha256:f5eaaceff91212bf6f37c4b7b4f18a8892dd85ba6ea23e73f942d35273dc82fb

Observation 00cc933f-07da-414e-b0a8-d626ea453f1b · inbound

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning cites this paper.

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:56:27.596710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T11:27:02.902720Z digest=sha256:06f7bc20d14ff403e00dfecfe513a20d9bfb4866b883b2afda46a9889686de5d

Observation e07449e6-5615-4e5d-830a-d4a80a3db36c · inbound

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning cites this paper.

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T12:32:56.389581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:32:56.389581Z digest=sha256:573dcbc768af77be2370f99adc0fd834ad614393c21f2033a7be7d136936c7f0

Observation a0c2d51d-8867-4505-9a6a-2d199d838521 · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:23.329479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:93589e1fa1863dd3ea9764e07b605386d51efb3d2456783b22e096f830032890

Observation f02e6a18-0bbf-4e9c-b300-0dbdb7bc3c4f · inbound

Amplifying Membership Signal Through Chained Regeneration cites this paper.

Amplifying Membership Signal Through Chained Regeneration Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:25:26.760412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T06:22:08.140403Z digest=sha256:f0948be57215d4418befdbbc3b6af5a5fed6d56dda057bc1e6ff9e416927bd07

Observation 3d34a930-1de0-4218-a169-0891cc3ef470 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 188

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:37.193909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:37.193909Z digest=sha256:f436ef77c09f5b76ff6a69e91a25f874c6dc5f608aff54d21db052de3aa2b112

Observation 3c0af4f0-f221-4eca-aa36-a9b188f75ecc · inbound

Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration cites this paper.

Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:35:50.513439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:35:50.513439Z digest=sha256:64aa558d6cf5a8eadc4b785a538013ef930775d7e6bfad8ef5093e332955862c