Pith. sign in

Paper Citation Record · LEDGER

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2411.10242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10242 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:54:19.224601Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:30.924737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T02:07:33.490588Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 21709886-3c2e-409c-bd6c-67600f594df9 · outbound

This paper cites an unresolved cited work.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:54:19.940546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.138509Z digest=sha256:dd1bbe55de2f3bb9a04d803d77c570b265e20d1fbb890d77976f27aa804664bf

Observation df1b763a-cafb-46bb-8611-068f4f2266d6 · outbound

This paper cites black hole.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models black hole

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.845550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.177600Z digest=sha256:744aa47c3844543de7e988aa94a6d813406075b3d6a9588081cad5da1d79d24b

Observation 1f4c4e58-dbc3-4fee-966d-21d868eb3749 · outbound

This paper cites Table 3: Model-specific instantiation of the assistant prompt.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Table 3: Model-specific instantiation of the assistant prompt

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.918726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.148620Z digest=sha256:2e7b0f75cb6436ec87b28de714e86893e747ddb3443f043e510e29bfc3486871

Observation 45043830-e10c-419f-b9cc-39d2afefb6a7 · outbound

This paper cites The Llama 3 Herd of Models.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.039863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.039863Z digest=sha256:5f6690c618852a40913896dd04c5e346f159528e54f532212cbe21cd955b2b0c

Observation 719ccea3-9be8-4c00-8a33-b6b979c14df3 · outbound

This paper cites AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.056145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.056145Z digest=sha256:46cf1bb76529d37295a32446589def8db00cbf919889d01097fb3dd2a6145e5a

Observation 879f1fd6-c7ff-43ed-9aa6-03c03bb0821a · outbound

This paper cites Scalable Extraction of Training Data from (Production) Language Models.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Scalable Extraction of Training Data from (Production) Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.063840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.063840Z digest=sha256:5d38860115d3203c10ca45200da87da607e9943b0e905b1fcf1d151a66d029a1

Observation ab1e3ec1-9963-45f9-ba81-a7f4da1e23a1 · outbound

This paper cites Understanding Transformers via N-gram Statistics.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Understanding Transformers via N-gram Statistics

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.071209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.071209Z digest=sha256:54e1701142a879f9cccac23403171bdd43ecac6d73285e1a82bc07bd9bcf3e31

Observation 46931ed1-aedc-402f-a259-8367be1cedc8 · outbound

This paper cites Does Writing with Language Models Reduce Content Diversity?.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Does Writing with Language Models Reduce Content Diversity?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.077929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.077929Z digest=sha256:6a72ab929d0d9b5a7911716231a4424b1b93c139302830fd59c5070ec50a5bcf

Observation cab7f8e7-8327-48b5-841a-24cea1494a33 · outbound

This paper cites Privacy risks of general-purpose language models.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Privacy risks of general-purpose language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:20.037587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.083846Z digest=sha256:41ba58d3cb24ff989e27c66eb2854d0c473e9cac6db6adf0414358b37e2d75d6

Observation 5686168c-ff99-4abf-903b-be182c3c0e6a · outbound

This paper cites Membership inference attacks against machine learning models.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Membership inference attacks against machine learning models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.088940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.088940Z digest=sha256:3aa80d2cd5b3bba2fdae055e61ef073c7b01882535858e2cc6a3f791d8bb369a

Observation e45a8618-8020-4f42-badc-8788c80c57c1 · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.100984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.100984Z digest=sha256:6c60122513331e2d384d437e8411024d02980e399cd1ba87c384f230d40d9c25

Observation 15a014cc-466d-4ad5-a6c6-1205169705b3 · outbound

This paper cites Privacy risk in machine learning: Analyzing the connection to overfitting.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Privacy risk in machine learning: Analyzing the connection to overfitting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.106680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.106680Z digest=sha256:d4515f2847063a2a62f918bb81957a49facf1bda929efbf6d50371f76fabf3c4

Observation 3a237033-3545-44c2-a880-502820b76ff3 · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.118870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.118870Z digest=sha256:4b16f8e497fc8f839a1a0e1b4c475e0c2dcc500f11b2cb6810079c67c5a6ba8f

Observation 529fbcb9-265c-43c0-9e7f-ca19810b8aef · outbound

This paper cites Under review.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Under review

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.983866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.124524Z digest=sha256:1547edd95cf82a4d122896563a51538347ff829cac9ed19886b6b3f396403959

Observation 2c1f8d55-e481-455c-9f1f-667dbf250cf6 · outbound

This paper cites Copy this text:.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Copy this text:

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.893700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.159515Z digest=sha256:3d05cd774e7c6776a64ef999aa2e48b59d87569095ec9ada772d4ce2d723ab1f

Observation c7d73801-61d1-4ab4-b101-f365967ca516 · outbound

This paper cites Under review.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Under review

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.870065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.169837Z digest=sha256:b022fcb5f998191cb00377584252cda7d355c41a3f4bddcb44d37e2ea2f3a9e9

Observation 06a26a10-f2e8-4482-a1c0-33fb45af7edb · outbound

This paper cites yolov3.weights.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models yolov3.weights

Reference 24

Resolution
verified exact
raw_fallback, observed 2026-08-12T19:54:19.393617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.188411Z digest=sha256:26cd71948a0dae51ebab102caa54884d607da27a87d726227f48f59ca1c225ac

Observation 0b3aa300-1b00-45ea-9e49-13a499cdd9f4 · outbound

This paper cites Claude 3.5 Sonnet generated the following text for the prompt “Write a news article about The Catalan declaration of independence.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Claude 3.5 Sonnet generated the following text for the prompt “Write a news article about The Catalan declaration of independence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.809539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.206186Z digest=sha256:41bd2ee76a7c9cd29e5875abec87c924ededcfafeb5ae400b63d38caf0999148

Observation 86472ce6-af98-45d8-b64f-473465a07804 · outbound

This paper cites Spain is living through a sad day,.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Spain is living through a sad day,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.787332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.217979Z digest=sha256:213b19e454238787007d6bfe30c8898f62bd62850c3d5fb80c47ad57bbdd5e42

Observation 87e7825d-a9e5-4939-831d-2c7126129d41 · outbound

This paper cites Schindler’s List.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Schindler’s List

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.766889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.224601Z digest=sha256:05f96b34589e1a83a15f6c82a319d08c09b879e20dbfbbb572d05bb85fdef6fc

Observation e7d930df-587d-4a9d-8113-997c31af06d0 · outbound

This paper cites independent writing.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models independent writing

Reference 500

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.960530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.131145Z digest=sha256:4eba81524dae7731d2ebc609ed303c7423e2bfd840b5a63e3b8d3ebd3e0579fb

Observation 024b0ba4-f64d-4789-b71c-86678b6bde7d · outbound

This paper cites This detection is the beginning of a new era: The field of gravitational wave astronomy is now a reality,.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models This detection is the beginning of a new era: The field of gravitational wave astronomy is now a reality,

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:19.827471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.199566Z digest=sha256:b2933f5f9bc4e6711b65723647ca36c1959482b46724a59b469fbb550ac67a10

Observation 27fa27ef-cd25-4ad3-8fca-2e4c18e65cd9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.095661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.095661Z digest=sha256:5a544887eb148a69727eaac85e482fe5ad4498c3073b9dc44c6d94a044a51d7a

Observation 5253d434-45fd-4aa2-b207-8f84c77fc6cc · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.113181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.113181Z digest=sha256:dc054a3ef47ecbe29eba8376e77f11e504d2788e9f0e568839769a1d03d61d18

Observation 6f5bd5f6-c68b-4d34-893d-f6464c6be5ef · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Quantifying Memorization Across Neural Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.032446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.032446Z digest=sha256:542e5ce97d94478102b006e06c6f4ee572ec224fcf415976587798dd6afc4e73

Observation 529f0d03-4e79-4f12-ba7d-604136337926 · outbound

This paper cites Reconstructing training data with informed adversaries.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Reconstructing training data with informed adversaries

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:54:20.065820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:54:19.025247Z digest=sha256:3f610c268f22d45816e2a12477cbed34314c0f76e693407ca878c2adbe6f6417

Observation 55fa256b-e306-464e-945b-b949beb1dc94 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Measuring Massive Multitask Language Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.049122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.049122Z digest=sha256:a3315824f7270b9f1b069662046e3628571203d6c0f4af385514a71b51c4258b

Observation 3dbb4d42-5518-4b8a-81e4-ff664c9dd5bd · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Measuring Non-Adversarial Reproduction of Training Data in Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T19:54:19.019137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:54:19.019137Z digest=sha256:bc023f2ce1d6cd967681209d3141667746a317c5e624f8198d630e25df3e2cfb

Pith citing papers

Observation 554a1aab-d4a9-4732-b58f-c68ceb5c21c6 · inbound

Membership Inference Attacks on Sequence Models cites this paper.

Membership Inference Attacks on Sequence Models Measuring Non-Adversarial Reproduction of Training Data in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:30.924737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:30.924737Z digest=sha256:8cdb160ac769dda078fdc75102413c14aec128a8536c55047fbb640667d533aa

Observation 125ac473-0634-420a-81d4-4ca8f561fee9 · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Measuring Non-Adversarial Reproduction of Training Data in Large Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.472921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:433f374c2a0a5c187e34eccf3bbfba53ed65b767791473aa75bebc98510a5181

Observation 403014c9-b9b9-47b7-a38c-6620ed20f3eb · inbound

SoK: Colluding Adversaries in Machine Learning Pipelines cites this paper.

SoK: Colluding Adversaries in Machine Learning Pipelines Measuring Non-Adversarial Reproduction of Training Data in Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.492459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T16:10:51.471822Z digest=sha256:3ef7a8620eca7a3c236b12b42c43ceb4a9a66f3e505ad2ad873f77364f60c580