Pith. sign in

Paper Citation Record · LEDGER

Extracting Training Data from Large Language Models

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 54 inbound Pith citation observations for arXiv:2012.07805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.07805 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 54 of 54 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:51:05.639460Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

275
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f249adc-febb-4812-a104-ac684d731332 · inbound

The Pile: An 800GB Dataset of Diverse Text for Language Modeling cites this paper.

The Pile: An 800GB Dataset of Diverse Text for Language Modeling Extracting Training Data from Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:35:18.749655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T21:35:18.513342Z digest=sha256:621b4c063e1226b05b8916cb96215caab995b1a32e770d99a788b2b487e9dea4

Observation a001a767-4dd2-40a8-909e-e78e492c4507 · inbound

Deduplicating Training Data Makes Language Models Better cites this paper.

Deduplicating Training Data Makes Language Models Better Extracting Training Data from Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T13:39:31.714001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-24T13:36:55.210708Z digest=sha256:60bffcf5834e9da92c2175ca969ccd57ad042d731f5dbb83391fac0fcdcc1a29

Observation 5d360aae-b923-4573-b2ad-60a12b397e5b · inbound

Ethical and social risks of harm from Language Models cites this paper.

Ethical and social risks of harm from Language Models Extracting Training Data from Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:24:29.952856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-11T18:24:28.835688Z digest=sha256:b3c14b7e85ebeba1a0efa9bdc63df76f51e9c6cb1d9d34b8e48b64ddf6cb0010

Observation 3da7c7d4-035e-4215-a99a-aa3bd6df49c2 · inbound

LaMDA: Language Models for Dialog Applications cites this paper.

LaMDA: Language Models for Dialog Applications Extracting Training Data from Large Language Models

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:17:32.518336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:17:32.272353Z digest=sha256:68634ac2d0c0c82e97a6fa2cedab7f530bb1d4002d2285941f9efb1a9553805c

Observation 0e12e445-3aea-4195-bc41-e52d725cb488 · inbound

Quantifying Memorization Across Neural Language Models cites this paper.

Quantifying Memorization Across Neural Language Models Extracting Training Data from Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:04:59.731563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T22:04:59.678438Z digest=sha256:1c227e37c52cb3be916e653a272e98277ad4e263a994a0314549617bffa8645b

Observation 78f2a4f4-10c0-4085-a8cd-f4364edeca01 · inbound

Scaling Laws and Interpretability of Learning from Repeated Data cites this paper.

Scaling Laws and Interpretability of Learning from Repeated Data Extracting Training Data from Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:52:40.461771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-17T15:52:40.335080Z digest=sha256:53687b2ddaaf127d395a6d714da38be438a2a64c3aea2bb711c6cf6fe8677a7d

Observation 862123b6-3250-4082-8f30-c978da3b3fc5 · inbound

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cites this paper.

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned Extracting Training Data from Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:38:08.468054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T01:38:08.362920Z digest=sha256:66e12620abf4dd02bef5244a65c74b8d489515045f0f6a60e211162d4facf6e9

Observation 54ae7179-ff65-4628-80c9-4d5bd6586ed9 · inbound

MusicLM: Generating Music From Text cites this paper.

MusicLM: Generating Music From Text Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:27:06.256783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T14:27:05.198402Z digest=sha256:0e858a66f3b27bae259dfa96ab5f4e372e9309c369b55f1c0dd4cd583b53fee5

Observation ae0ef567-9f81-41fc-a125-e94baeaec035 · inbound

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions cites this paper.

Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions Extracting Training Data from Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:23:52.895916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-24T04:21:49.775278Z digest=sha256:2d60334775bcd64d15ecbda93d7e12b02d68ac634b6b1c94127b7108a0902570

Observation 8fd242df-c7e3-41f4-83a3-3354df4306fa · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Extracting Training Data from Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.740850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:2b8c07e996767aaf483f6cadd092a799190b2b9ed0bd7b4babdcef5887f19fa5

Observation 791924ef-9c4a-4656-b0f7-6e84274e76fc · inbound

A Social Outcomes and Priorities centered (SOP) Framework for AI policy cites this paper.

A Social Outcomes and Priorities centered (SOP) Framework for AI policy Extracting Training Data from Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:51:05.639460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:51:05.639460Z digest=sha256:e32a0c1d9bbdad9469387012872db6e0e8baf3c7087706f853bfffb80d9041bb

Observation d60d1838-244d-47d7-b20c-fca6d297796d · inbound

Detecting Memorization in Large Language Models cites this paper.

Detecting Memorization in Large Language Models Extracting Training Data from Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T04:50:15.626636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:50:15.626636Z digest=sha256:f54852c2c81669f574f55fc5f6245144838dc3925bd5ca55d8a009d49815ec9b

Observation e37d055d-91eb-41eb-967c-83c4a0bf7aa9 · inbound

Trust & Safety of LLMs and LLMs in Trust & Safety cites this paper.

Trust & Safety of LLMs and LLMs in Trust & Safety Extracting Training Data from Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T23:52:33.159458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:52:33.159458Z digest=sha256:0b3b17516b22022cc5dc1d6e822bc17bb73b918713052fe3181ac0cf812dd3bd

Observation 07b45183-d87e-42d3-aee1-ba78389b23f1 · inbound

Towards the Anonymization of the Language Modeling cites this paper.

Towards the Anonymization of the Language Modeling Extracting Training Data from Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:32:39.434987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T06:28:16.975305Z digest=sha256:d8d73bc3a2f93ba2b6900ccd4656b95b8c96594867829d03fb46c7701ac7507b

Observation 62fcb607-fd74-4d3c-ad40-67c0d7f50894 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Extracting Training Data from Large Language Models

Reference 253

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:11:37.303275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:b1aff73d7f5e57a35063a54171321c5534e412d9cf592d57f7528c8c94d7a4ad

Observation 73256e9c-8ad7-463d-8010-e673d7e5b027 · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning Extracting Training Data from Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:25.072025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:25.072025Z digest=sha256:7cc0240a10dd941ff17b6bd8b91db2fa6f681cbe00a27a24d799bfd2496e43b7

Observation 65626d80-068e-4c69-a336-5968e501f862 · inbound

Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation cites this paper.

Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation Extracting Training Data from Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:34:30.348594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T01:34:21.657999Z digest=sha256:bba52f4295ff67eea9f4682f84279a889c7f9b22f64b5567956a162ea68b897e

Observation 63487873-dc3e-4722-88fe-87cac685c658 · inbound

Approximating Language Model Training Data from Weights cites this paper.

Approximating Language Model Training Data from Weights Extracting Training Data from Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:58.403904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:58:58.403904Z digest=sha256:9b0b2b862e776f9dcc76e93e8b5085baf7c50249c91404a8d6843572d91156e7

Observation 337899db-b046-4a2b-9411-196bf7d67a2e · inbound

From Teacher to Student: Tracking Memorization Through Model Distillation cites this paper.

From Teacher to Student: Tracking Memorization Through Model Distillation Extracting Training Data from Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:10.559721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:10.559721Z digest=sha256:bde6993f5b73bd226bb8f4d3bbcfa8d980846db7171dc9f68bcc3f833a406833

Observation 42bf4c06-3c92-4a00-9aa0-b9d9be79d119 · inbound

Low-Perplexity LLM-Generated Sequences and Where To Find Them cites this paper.

Low-Perplexity LLM-Generated Sequences and Where To Find Them Extracting Training Data from Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:58.634485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:44:58.634485Z digest=sha256:ee4a1ce448b84983007fae77531a3d904852b75dca54821f76b2fc3e65c5cb51

Observation 6ea55d67-c336-4e53-831c-47d55a3e9ae5 · inbound

Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models cites this paper.

Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models Extracting Training Data from Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:48:55.141023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:48:55.141023Z digest=sha256:4a8f4aeb403d48d7cdc1900fb16e2dc66d6e7ad9018d3e1b27e051c5aab48bef

Observation e941e346-f608-4269-aa6c-193a4c956cbc · inbound

Adversarial Machine Learning Attacks on Financial Reporting via Maximum Violated Multi-Objective Attack cites this paper.

Adversarial Machine Learning Attacks on Financial Reporting via Maximum Violated Multi-Objective Attack Extracting Training Data from Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:00.983161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:31:00.983161Z digest=sha256:93698186f1ad02ece87dc559c297d7720b2d869247ab456de08af05fc8905d9f

Observation 77948df8-f2ae-4a6a-9eeb-3f0a4ab7d5c2 · inbound

On the Performance of Differentially Private Optimization with Heavy-Tail Class Imbalance cites this paper.

On the Performance of Differentially Private Optimization with Heavy-Tail Class Imbalance Extracting Training Data from Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:00.426132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:00.426132Z digest=sha256:5a8cea3e08941b0bba93749cd519bffab4c6891828f292147225f6c63423f44c

Observation aa9c0ed5-5f41-4d48-941c-dccbe7fe0fdc · inbound

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI cites this paper.

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI Extracting Training Data from Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:55.618154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:55.618154Z digest=sha256:f6710ba3b7e4dc83d8185b8db95423ddd62f94f71402813bbb922efb5868bc67

Observation 45843fde-c633-4720-a317-5c86e2576e11 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Extracting Training Data from Large Language Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.822425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.822425Z digest=sha256:6d7898ed3fcb4e10cf317159d3c693c486e8d6d731b74c7d6220359ad7f126e3

Observation cd504875-9749-41af-a9f8-8cfebb784e8d · inbound

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage cites this paper.

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage Extracting Training Data from Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:50:19.605008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:50:19.605008Z digest=sha256:6204b7372a3296a42137f9088fae5bda432b576cfe001c43e247ffb424db41f0

Observation 5a4074d2-053f-4fce-b7a7-fa03fb5fe38e · inbound

Evaluating Differentially Private Generation of Domain-Specific Text cites this paper.

Evaluating Differentially Private Generation of Domain-Specific Text Extracting Training Data from Large Language Models

Reference 2650

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:03.247046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:09:03.247046Z digest=sha256:f9b77ff32bb096fe4576610f1ac28257e29ac88c4a4e7eef3da2f1106aea24e4

Observation 178f3650-d0a3-4e33-8ccd-087eef77de7b · inbound

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants cites this paper.

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants Extracting Training Data from Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:46:41.159734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:46:41.159734Z digest=sha256:a3f48caf852f4de1b096b58fee5aa4585912c0ba5b21b9f122b13f3c0101add3

Observation ffe1c536-0a58-4f78-b42e-25b0318b13be · inbound

SynBench: A Benchmark for Differentially Private Text Generation cites this paper.

SynBench: A Benchmark for Differentially Private Text Generation Extracting Training Data from Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:46:37.468783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T16:46:22.479895Z digest=sha256:11ad6236c5e2568707415314bc3d9c3419d4b76bf0a65b502e6eb5494422621a

Observation 1b70cc15-f9fc-40f1-864b-0fc5c61e68b4 · inbound

When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation cites this paper.

When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation Extracting Training Data from Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.851547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T23:47:31.667427Z digest=sha256:97ec6b7b4c81581542a72d1c6fddf227c29cef8f84908d38ce7fb22b9744376e

Observation a00801a1-c0c0-4373-9b48-aa540bf3dae9 · inbound

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts cites this paper.

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts Extracting Training Data from Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:48.805002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T19:44:26.931344Z digest=sha256:3693daa2ef47eab908ff3e47df4e288630efb6594b0a531af0e4e13575c63b01

Observation 6dc63095-16dc-4a9c-bda3-7ae3027a0fea · inbound

Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies cites this paper.

Separable Expert Architecture: Toward Privacy-Preserving LLM Personalization via Composable Adapters and Deletable User Proxies Extracting Training Data from Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T22:34:07.373969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-09T22:26:34.487400Z digest=sha256:45bbbaf9192f2e33c0f8f0f13248a636f6906a937969a12d0f946d3c39cb9d8c

Observation 5b4fee91-8675-4ad1-9c0a-12203b86194f · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Extracting Training Data from Large Language Models

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:09.389604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:0af779713d2165b2321b5ee0cfd7226c5f98b176f36a6200364f4f1a0ef5b49a

Observation 89517f86-bac3-429d-b231-14029185edba · inbound

Making AI-Assisted Grant Evaluation Auditable without Exposing the Model cites this paper.

Making AI-Assisted Grant Evaluation Auditable without Exposing the Model Extracting Training Data from Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:51:17.879632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T16:12:20.534893Z digest=sha256:e4454a6c08bb5b492b6a94b83327d6a2db92468e58aec67ad74da1e27a4dbb26

Observation 4f74794c-9c4d-4e90-b0d3-ef5718da6854 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Extracting Training Data from Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.288375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:1185798322d3eeeed0a08c151f6b10fb369c0e3baaf77779ab4a989e2dd5fc4c

Observation 57b34014-a152-497f-914f-558789c1199b · inbound

LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems cites this paper.

LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems Extracting Training Data from Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:01:05.960115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T04:58:37.780195Z digest=sha256:9042de9ea194ba468b75158575028fd58445733f63ed869d72e046c4c48f49e0

Observation f9fbb4f0-4061-4a72-9d45-69a78e5d0c7e · inbound

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing cites this paper.

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:44:01.664060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-29T22:37:02.708284Z digest=sha256:2d3d6043393758b47f7d0a5c680aa8f27b1a0331381a72ca7f2510e4e41c556a

Observation 40c45581-7f6a-4603-a973-ad6231aaf9b3 · inbound

MRMMIA: Membership Inference Attacks on Memory in Chat Agents cites this paper.

MRMMIA: Membership Inference Attacks on Memory in Chat Agents Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:27.072292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T12:04:53.511344Z digest=sha256:09569e45a8eae01b786d122702b5baddac566a2f9302cab7d11d88fce4c62c15

Observation d64b9315-cf10-4cbb-b788-e426d461e0a9 · inbound

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing cites this paper.

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:56:23.234607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T13:49:47.542697Z digest=sha256:1f0870d89d60cf50aeb2e6bc9b7a9b5ccf001a8ed2f881f252022a8139e16342

Observation 1468ffd0-0d5b-4fbc-8846-251fc95b7cbf · inbound

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails cites this paper.

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails Extracting Training Data from Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.057543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T02:06:31.341519Z digest=sha256:ba0c973ce6db6aac11b6f13a0a5e0279ff6d9151e1fe6ea3a78657184bd51454

Observation 80ad5aa8-0703-40ed-8ccb-81d3bd7d3de1 · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Extracting Training Data from Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.468052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:4adcb8aac9b70450405a27bb4131f2ca9671ca8d692310dc56f737f001e226b6

Observation f30c694f-27ce-46d4-9fc0-2cd79110ae11 · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Extracting Training Data from Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.357949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:de923c435c4e03426e01155b45845bd1ad7cc6b8c2353acf886720d754b54a93

Observation b6dac66e-c769-457e-8edf-e64ccde2034f · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs Extracting Training Data from Large Language Models

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:07:30.231259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:f2e62808ce7716d5041c90a2aed2ec29c6878472729d247bc5553eed69fefa2f

Observation e2d36494-b0a8-4b0d-8507-d9c8ae0a36cb · inbound

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans cites this paper.

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans Extracting Training Data from Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:05:36.687299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T09:00:23.558245Z digest=sha256:a1b8640df826ea871f1af0afcebb222901dd37722c26b0c953ba0b4872dd8d5a

Observation 48cfd250-1281-4f89-a3b2-72bd4189fdcf · inbound

OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents cites this paper.

OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents Extracting Training Data from Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:58:06.926713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T09:11:42.204778Z digest=sha256:76899cd24fd35bdad893088c52b47f50ab312c9105586d1d8f59fb05c9a3035a

Observation fb4f126f-1954-4a98-80d7-4e31d24dede2 · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity Extracting Training Data from Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:58:47.182482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:2dd4c03324f4021c69e709716ab2e9ed2d61d980ac83dd39ce83ea5475f76a99

Observation 99fef884-5b0d-4d97-9e02-13ed9e2c4cf6 · inbound

Exposing the Illusion of Erasure in Knowledge Editing for LLMs cites this paper.

Exposing the Illusion of Erasure in Knowledge Editing for LLMs Extracting Training Data from Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.276219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T09:10:39.422141Z digest=sha256:4666ca2f8dced08dbf5c369b207ba156f932f93bf918789d012407f210e5e1b1

Observation ec1a872e-ec30-4beb-b1a3-9eb6d7dd83c7 · inbound

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents cites this paper.

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents Extracting Training Data from Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:09:53.281660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T04:29:16.386339Z digest=sha256:b8544444159f88b1f0d2d5a1555420bbf0997367d60aab06867717e3f7540623

Observation ffdbb1a8-433a-472a-900d-5d4097d0fa7d · inbound

AI Native Games: A Survey and Roadmap cites this paper.

AI Native Games: A Survey and Roadmap Extracting Training Data from Large Language Models

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:06:58.758903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-02T12:59:20.908659Z digest=sha256:fd3a44d775424bd8e280bb82ee290d619a0fc1641a95da8f5b249f396d875042

Observation 4d8c8d02-19f8-4a42-88a7-cde6e538eab4 · inbound

AI Native Games: A Survey and Roadmap cites this paper.

AI Native Games: A Survey and Roadmap Extracting Training Data from Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-12T09:30:07.729486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:30:07.729486Z digest=sha256:4a66fa92591a824e42952a91766f3982ec8ddd9c74b7bbc813bfe3b4af5a67d5

Observation 41d6405e-1aa9-43f1-8a6a-e244df868e0e · inbound

Auditing Forgetting in Limited Memory Language Models cites this paper.

Auditing Forgetting in Limited Memory Language Models Extracting Training Data from Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.312170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-02T13:09:34.491580Z digest=sha256:9846fc6ab95efb273204ace1b09ae9336986b1bddab13061affc8350b17ea6cf

Observation 724b7aaa-782d-4928-b56a-e74cdf712e32 · inbound

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification cites this paper.

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification Extracting Training Data from Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T14:11:43.482286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:11:43.482286Z digest=sha256:7c32ca34d581f7622d7edf406c48dd947e34c939667a220948d09236147a321d

Observation 9c1c3df4-b145-41bb-a08f-62af760ed21c · inbound

DECAF: De-Clustering for Adaptive Representational Unlearning cites this paper.

DECAF: De-Clustering for Adaptive Representational Unlearning Extracting Training Data from Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:33:49.808882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:33:49.808882Z digest=sha256:d8cd6aef1ea3602e38ca601ea180c09f3c2c48e2681b2f8e630860d3b24c838b

Observation c08c5f00-e512-4bec-91e8-c1bbee4adaf8 · inbound

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization cites this paper.

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization Extracting Training Data from Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T02:28:36.726398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:28:36.726398Z digest=sha256:2cc3177ff1889e1a4d80e4fc63c1018243082c52bd25612c95bacbe16b37407f