Pith. sign in

Paper Citation Record · LEDGER

Extracting Training Data from Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2012.07805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.07805 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:37:25.072025Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

275
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f249adc-febb-4812-a104-ac684d731332 · inbound

The Pile: An 800GB Dataset of Diverse Text for Language Modeling cites this paper.

The Pile: An 800GB Dataset of Diverse Text for Language Modeling Extracting Training Data from Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:35:18.749655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T21:35:18.513342Z digest=sha256:40fa73959798280c688affc484f7a8e9be4cd35c8807bf35d492d9f23c7dd8db

Observation a001a767-4dd2-40a8-909e-e78e492c4507 · inbound

Deduplicating Training Data Makes Language Models Better cites this paper.

Deduplicating Training Data Makes Language Models Better Extracting Training Data from Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T13:39:31.714001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T13:36:55.210708Z digest=sha256:586a68c0b2ab23486d744f4f963cbee7e81b62eaacf7f53148218e4faf2b3f00

Observation 5d360aae-b923-4573-b2ad-60a12b397e5b · inbound

Ethical and social risks of harm from Language Models cites this paper.

Ethical and social risks of harm from Language Models Extracting Training Data from Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:24:29.952856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T18:24:28.835688Z digest=sha256:731f574b4b3f4488f23bd84f6954b6fa588f1aec51d45d4ef33411072f76185b

Observation 3da7c7d4-035e-4215-a99a-aa3bd6df49c2 · inbound

LaMDA: Language Models for Dialog Applications cites this paper.

LaMDA: Language Models for Dialog Applications Extracting Training Data from Large Language Models

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:17:32.518336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:17:32.272353Z digest=sha256:f972b4483762cb8ff1298dc9175e7fb8e42b0c9cbeda1be0a8f1e00a8439813a

Observation 0e12e445-3aea-4195-bc41-e52d725cb488 · inbound

Quantifying Memorization Across Neural Language Models cites this paper.

Quantifying Memorization Across Neural Language Models Extracting Training Data from Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:04:59.731563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:04:59.678438Z digest=sha256:52efc9e5e83dafd6b03c86dac0265f477156716b62c692250b5768c85c6543a0

Observation 78f2a4f4-10c0-4085-a8cd-f4364edeca01 · inbound

Scaling Laws and Interpretability of Learning from Repeated Data cites this paper.

Scaling Laws and Interpretability of Learning from Repeated Data Extracting Training Data from Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:52:40.461771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T15:52:40.335080Z digest=sha256:d77dcb743fd6bd4ca077e2f96c1a4fd1944f53a804cd17a6a7c6a36d859c08de

Observation 862123b6-3250-4082-8f30-c978da3b3fc5 · inbound

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cites this paper.

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned Extracting Training Data from Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:38:08.468054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:08.362920Z digest=sha256:72a3f3b6eafaceb97167eba792acca6e5b9fcb7d2276ab30c9da2b016fc2ff01

Observation 54ae7179-ff65-4628-80c9-4d5bd6586ed9 · inbound

MusicLM: Generating Music From Text cites this paper.

MusicLM: Generating Music From Text Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:27:06.256783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T14:27:05.198402Z digest=sha256:7f657c06268d8107b8b834d7d438cecaa9452eea2a3b752d8002dcee157337d5

Observation ae0ef567-9f81-41fc-a125-e94baeaec035 · inbound

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions cites this paper.

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions Extracting Training Data from Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:23:52.895916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T04:21:49.775278Z digest=sha256:998e3d6934d90e95b34f7f3690696141fc913de3fcc78ef26c68b08a307f982e

Observation 8fd242df-c7e3-41f4-83a3-3354df4306fa · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Extracting Training Data from Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.740850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:ddb987b4e4f4fcb2120087f1eb6bb15ae8d5090d82d13d68e05ea1dbc73e6891

Observation 07b45183-d87e-42d3-aee1-ba78389b23f1 · inbound

Towards the Anonymization of the Language Modeling cites this paper.

Towards the Anonymization of the Language Modeling Extracting Training Data from Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:32:39.434987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:28:16.975305Z digest=sha256:382dc3428745f434b4df7ef5fbad0571b390076e3d5e2a54e284b0e4d9b51634

Observation 62fcb607-fd74-4d3c-ad40-67c0d7f50894 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Extracting Training Data from Large Language Models

Reference 253

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:11:37.303275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:9a3c7505409b2db03c047f8b1a0f6765d20a3acbd4c59e53b96b3105336afdad

Observation 73256e9c-8ad7-463d-8010-e673d7e5b027 · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning Extracting Training Data from Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:25.072025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:25.072025Z digest=sha256:0b4a468a5f506db37fee94780f56535db15f13a0cabac498ecfc37e51dced012

Observation 65626d80-068e-4c69-a336-5968e501f862 · inbound

Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation cites this paper.

Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation Extracting Training Data from Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:34:30.348594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T01:34:21.657999Z digest=sha256:236142cf112452d13d54cbbd8e0b2d2e531daf0e03a4a753b7201a7e46bee4ba

Observation 63487873-dc3e-4722-88fe-87cac685c658 · inbound

Approximating Language Model Training Data from Weights cites this paper.

Approximating Language Model Training Data from Weights Extracting Training Data from Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:58.403904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:58:58.403904Z digest=sha256:27278043d8987d4fa93f0d0273c609b584f633c7ec2ecbef25b0ddbb0d2df3c0

Observation 337899db-b046-4a2b-9411-196bf7d67a2e · inbound

From Teacher to Student: Tracking Memorization Through Model Distillation cites this paper.

From Teacher to Student: Tracking Memorization Through Model Distillation Extracting Training Data from Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:10.559721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:10.559721Z digest=sha256:6c9dc6c4ebe700384af8ba3f270cab21454e9707d76dc5aba7e6f70d41093ace

Observation 42bf4c06-3c92-4a00-9aa0-b9d9be79d119 · inbound

Low-Perplexity LLM-Generated Sequences and Where To Find Them cites this paper.

Low-Perplexity LLM-Generated Sequences and Where To Find Them Extracting Training Data from Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:58.634485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:44:58.634485Z digest=sha256:02981a16ec4b40e16db0f8947474475a20653be16e912fb5e317c183706a43b8

Observation 6ea55d67-c336-4e53-831c-47d55a3e9ae5 · inbound

Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models cites this paper.

Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models Extracting Training Data from Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:48:55.141023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:48:55.141023Z digest=sha256:0a692eb14e4c1aff2c777a91fa7f12439a676bfec3f9e84b77b40b448aaf292a

Observation e941e346-f608-4269-aa6c-193a4c956cbc · inbound

Adversarial Machine Learning Attacks on Financial Reporting via Maximum Violated Multi-Objective Attack cites this paper.

Adversarial Machine Learning Attacks on Financial Reporting via Maximum Violated Multi-Objective Attack Extracting Training Data from Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:00.983161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:31:00.983161Z digest=sha256:4e2f247d4097a93ff4a2ce5cb610d2a7f1ce466bd55b9a402687edd9e330d11b

Observation 77948df8-f2ae-4a6a-9eeb-3f0a4ab7d5c2 · inbound

On the Performance of Differentially Private Optimization with Heavy-Tail Class Imbalance cites this paper.

On the Performance of Differentially Private Optimization with Heavy-Tail Class Imbalance Extracting Training Data from Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:00.426132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:00.426132Z digest=sha256:ecf04aea5a26f32164ea4ebdb7913a03a3005610b5a2af9c632027468c7a7486

Observation aa9c0ed5-5f41-4d48-941c-dccbe7fe0fdc · inbound

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI cites this paper.

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI Extracting Training Data from Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:55.618154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:55.618154Z digest=sha256:d9191ef598f150dff5da657cce2c7f81afd044a74e6ba4d6fe252d3d3ccb80a4

Observation 45843fde-c633-4720-a317-5c86e2576e11 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Extracting Training Data from Large Language Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.822425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.822425Z digest=sha256:877b3599887b42067ade6655a54b3b291e9c742114c665269e00450bcbd0ce08

Observation cd504875-9749-41af-a9f8-8cfebb784e8d · inbound

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage cites this paper.

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage Extracting Training Data from Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:50:19.605008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:50:19.605008Z digest=sha256:7b4e5082fa6bfda668d5cfb1d480c7ef82acdd6653791b6b001db6b7808d4774

Observation 5a4074d2-053f-4fce-b7a7-fa03fb5fe38e · inbound

Evaluating Differentially Private Generation of Domain-Specific Text cites this paper.

Evaluating Differentially Private Generation of Domain-Specific Text Extracting Training Data from Large Language Models

Reference 2650

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:03.247046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:09:03.247046Z digest=sha256:dfe89199272c8e144df22a87688f13bd41727227c1f7cd8a3a4baae6a4d21cb6

Observation 178f3650-d0a3-4e33-8ccd-087eef77de7b · inbound

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants cites this paper.

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants Extracting Training Data from Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:46:41.159734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:46:41.159734Z digest=sha256:206456f9836d14463614c79cc78988f8071685fcae29786b884601df487d2bca

Observation ffe1c536-0a58-4f78-b42e-25b0318b13be · inbound

SynBench: A Benchmark for Differentially Private Text Generation cites this paper.

SynBench: A Benchmark for Differentially Private Text Generation Extracting Training Data from Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:46:37.468783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:46:22.479895Z digest=sha256:d5488f604d0929fa006f5c2747a26ba6c1f5f0a611670c62024f4c37c2879dfe

Observation 1b70cc15-f9fc-40f1-864b-0fc5c61e68b4 · inbound

When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation cites this paper.

When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation Extracting Training Data from Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.851547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:47:31.667427Z digest=sha256:c0493fef823e12e5dec90231e2e56ab7ac20347dfe4b49f44905304cad78b027

Observation a00801a1-c0c0-4373-9b48-aa540bf3dae9 · inbound

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts cites this paper.

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts Extracting Training Data from Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:48.805002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:44:26.931344Z digest=sha256:e26808251d291dc06e4e4dabc729484113a4266a7b4c27120a9717c9a7fbe5fa

Observation 6dc63095-16dc-4a9c-bda3-7ae3027a0fea · inbound

Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies cites this paper.

Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies Extracting Training Data from Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:34:07.373969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T22:26:34.487400Z digest=sha256:1b6c70d71975444fb77aa8ef915b03494a933a83e412f95b3183b048c12da4dd

Observation 5b4fee91-8675-4ad1-9c0a-12203b86194f · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Extracting Training Data from Large Language Models

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:09.389604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:d918ebbf349ce42489db2752c378ceccdfd765940b05b6d010f190c2840e0417

Observation 89517f86-bac3-429d-b231-14029185edba · inbound

Making AI-Assisted Grant Evaluation Auditable without Exposing the Model cites this paper.

Making AI-Assisted Grant Evaluation Auditable without Exposing the Model Extracting Training Data from Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:51:17.879632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:12:20.534893Z digest=sha256:ebaa12290ce72dcf7ecfe74b725cb857ae7f66fc111b67689999a1a3e4e2570c

Observation 4f74794c-9c4d-4e90-b0d3-ef5718da6854 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Extracting Training Data from Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.288375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:862b5a78496c3fe4e8fdb36f96b63ec670453f60e3c8064fc87754af07d9d095

Observation 57b34014-a152-497f-914f-558789c1199b · inbound

LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems cites this paper.

LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems Extracting Training Data from Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:01:05.960115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T04:58:37.780195Z digest=sha256:6dc7dca174e4405a698d78572d79e5eb177662f2b214f9cc4b09194cd05748cf

Observation f9fbb4f0-4061-4a72-9d45-69a78e5d0c7e · inbound

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing cites this paper.

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:44:01.664060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T22:37:02.708284Z digest=sha256:c235afed53ee1057bd7e111122ffd38514f0705ef499ed10cf0e4b75ae3e7d45

Observation 40c45581-7f6a-4603-a973-ad6231aaf9b3 · inbound

MRMMIA: Membership Inference Attacks on Memory in Chat Agents cites this paper.

MRMMIA: Membership Inference Attacks on Memory in Chat Agents Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:27.072292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:04:53.511344Z digest=sha256:36424f3dff6505f2224ee815f1a90a2d21760db7dfb01c0b9992240419541ffa

Observation d64b9315-cf10-4cbb-b788-e426d461e0a9 · inbound

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing cites this paper.

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:56:23.234607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T13:49:47.542697Z digest=sha256:18d91fa7c239223331aaf8fdc6c449c4f3cae4c1f1e633db8c0f433f63b97ca8

Observation 1468ffd0-0d5b-4fbc-8846-251fc95b7cbf · inbound

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails cites this paper.

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.057543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:06:31.341519Z digest=sha256:cd36bef0ed2e488a699111684a716090099a7dcb8f2e11b3f6fa98c3456b8d85

Observation 80ad5aa8-0703-40ed-8ccb-81d3bd7d3de1 · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Extracting Training Data from Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.468052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:ebe1a3f7457126bb74c2860935d9652095c9290b505bed1e89f061d46236f896

Observation f30c694f-27ce-46d4-9fc0-2cd79110ae11 · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Extracting Training Data from Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.357949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:d898226c6c48e7af4e3626f3fbb7bc4926e8896ad14bc3b4f495bab26b53045b

Observation b6dac66e-c769-457e-8edf-e64ccde2034f · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs Extracting Training Data from Large Language Models

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:07:30.231259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:31ae3b157aaab33adc100347f3f2ece2c63f64162b31ac06067045e288784197

Observation e2d36494-b0a8-4b0d-8507-d9c8ae0a36cb · inbound

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans cites this paper.

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans Extracting Training Data from Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:05:36.687299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T09:00:23.558245Z digest=sha256:3cc8df2dea79019a6f78963819b149fcbdc5346c2b11ab312feae6b2ba705f06

Observation 48cfd250-1281-4f89-a3b2-72bd4189fdcf · inbound

OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents cites this paper.

OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents Extracting Training Data from Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:58:06.926713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:11:42.204778Z digest=sha256:9d35857bdee53849e8f7a4ea25c2d61dd4f5b829dfc693ae5346c113d9ff7021

Observation fb4f126f-1954-4a98-80d7-4e31d24dede2 · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:58:47.182482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:748404b433d9d4396579f4781ce118515b8b15bf644b313c885021bf32b871ac

Observation 99fef884-5b0d-4d97-9e02-13ed9e2c4cf6 · inbound

Exposing the Illusion of Erasure in Knowledge Editing for LLMs cites this paper.

Exposing the Illusion of Erasure in Knowledge Editing for LLMs Extracting Training Data from Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.276219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:10:39.422141Z digest=sha256:d98e8499727723d8230afc9fb2f2d87b2b40c3c141253007b46e712ca9639991

Observation ec1a872e-ec30-4beb-b1a3-9eb6d7dd83c7 · inbound

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents cites this paper.

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents Extracting Training Data from Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:09:53.281660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T04:29:16.386339Z digest=sha256:578b2bf886e5608cf2c595735c9328d7f50ba4569252a6fb445258550293e6e3

Observation ffdbb1a8-433a-472a-900d-5d4097d0fa7d · inbound

AI Native Games: A Survey and Roadmap cites this paper.

AI Native Games: A Survey and Roadmap Extracting Training Data from Large Language Models

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:06:58.758903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T12:59:20.908659Z digest=sha256:a5a3ff2e8d5d2980781edd5ae3111928d8fb9f4cbb039a28251dc1cf61690ce7

Observation 4d8c8d02-19f8-4a42-88a7-cde6e538eab4 · inbound

AI Native Games: A Survey and Roadmap cites this paper.

AI Native Games: A Survey and Roadmap Extracting Training Data from Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-12T09:30:07.729486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:30:07.729486Z digest=sha256:f0035c4ab8141986f2dfb2ab0e24eaba1e9b2cc255b25d734e7ab9ead58db7d4

Observation 41d6405e-1aa9-43f1-8a6a-e244df868e0e · inbound

Auditing Forgetting in Limited Memory Language Models cites this paper.

Auditing Forgetting in Limited Memory Language Models Extracting Training Data from Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.312170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T13:09:34.491580Z digest=sha256:d6bed69816b4b31848107262d00569c2236158045a4aa0f19e986c6ac3b0f576

Observation 724b7aaa-782d-4928-b56a-e74cdf712e32 · inbound

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification cites this paper.

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification Extracting Training Data from Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T14:11:43.482286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:11:43.482286Z digest=sha256:0e3d1442fada4fd3c278a3a1c6bf02ee5a9c699b7a62bb9dfa0315c21088509f

Observation 9c1c3df4-b145-41bb-a08f-62af760ed21c · inbound

DECAF: De-Clustering for Adaptive Representational Unlearning cites this paper.

DECAF: De-Clustering for Adaptive Representational Unlearning Extracting Training Data from Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:33:49.808882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:33:49.808882Z digest=sha256:21f9f90ea1f7f5dc85a5ec6d5f630e98de99759ad94330ef8bf5e493cc0db2a7

Observation c08c5f00-e512-4bec-91e8-c1bbee4adaf8 · inbound

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization cites this paper.

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization Extracting Training Data from Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T02:28:36.726398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:28:36.726398Z digest=sha256:72f77a90b398067de62ea7ccd19d9d7240485f977713e467025deebdff0c5cc1