Pith. sign in

Paper Citation Record · LEDGER

Sequence-Level Leakage Risk of Training Data in Large Language Models

As of 20 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 4 inbound Pith citation observations for arXiv:2412.11302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11302 v3

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:09:50.778525Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:33.766535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T14:15:55.154579Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29208af9-e6e5-4574-b99e-dc4c3f738315 · outbound

This paper cites write newline.

Sequence-Level Leakage Risk of Training Data in Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.628972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.628972Z digest=sha256:674de5c4da0bdd9b938c74a012a82bb36c25e4e528266727e25fcf9eaa899fdc

Observation 32ce6619-ad7d-4c6c-8e5a-09ba5d0d2aef · outbound

This paper cites C hat M istral.

Sequence-Level Leakage Risk of Training Data in Large Language Models C hat M istral

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:51.229009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.636028Z digest=sha256:cb0b37b6c00526cb911f4fe733b06a37d62eac6d03be05876b4e338aad380dd1

Observation 5730a92f-fd00-4696-a31c-700fc446bc91 · outbound

This paper cites an unresolved cited work.

Sequence-Level Leakage Risk of Training Data in Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:09:51.210913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.645721Z digest=sha256:f82b617d76b76811a014ec5dbd2c69d3955a7a8cc93d48020993ce866158ee1d

Observation a7c518e0-fe5a-472e-967b-bf9fd6590712 · outbound

This paper cites Emergent and predictable memorization in large language models.

Sequence-Level Leakage Risk of Training Data in Large Language Models Emergent and predictable memorization in large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:51.192501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.651308Z digest=sha256:df2a2f4d86e95df29270495470634b5e3ae4768a2975162c8ab946a3971d8043

Observation d1dd711d-cdaf-457f-a00f-c2b6314803ac · outbound

This paper cites The secret sharer: Evaluating and testing unintended memorization in neural networks.

Sequence-Level Leakage Risk of Training Data in Large Language Models The secret sharer: Evaluating and testing unintended memorization in neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:51.174257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.659687Z digest=sha256:725316c2eecd73a63afc657efe9780b8f1edfacedd485aadf114e52b244f3dc8

Observation 5efedde3-1c84-46c0-9852-0adf52c583c8 · outbound

This paper cites Extracting training data from large language models.

Sequence-Level Leakage Risk of Training Data in Large Language Models Extracting training data from large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.667920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.667920Z digest=sha256:8c7e161fa147921fd49b7556a1e49321d9bc6ba9d69dfe1090e3615829184fdf

Observation 15512697-6a4c-4017-8c1d-4a218d5fba9f · outbound

This paper cites Quantifying memorization across neural language models.

Sequence-Level Leakage Risk of Training Data in Large Language Models Quantifying memorization across neural language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:51.142214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.673629Z digest=sha256:db88df567e05d5f10f0d23b70ffcc82445197ebdffc09f0ca7ba67e1090b80d1

Observation a6143e59-c360-4aa8-bc25-f967fad68b7e · outbound

This paper cites an unresolved cited work.

Sequence-Level Leakage Risk of Training Data in Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:09:51.124399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.678274Z digest=sha256:d70e21307bda4764b36d607c34ba513a8ac30127d596d7cb285341cc20098c9b

Observation 0ef7962d-0d84-4b6b-81ee-79404630f759 · outbound

This paper cites Do Membership Inference Attacks Work on Large Language Models?.

Sequence-Level Leakage Risk of Training Data in Large Language Models Do Membership Inference Attacks Work on Large Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.684310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.684310Z digest=sha256:e6091d9b9bc5e5e53e57ec9ed4365a36bdce6f5791b9df869e436480039735cd

Observation 19dffadc-0e1b-4dd5-aedb-ca58187dda87 · outbound

This paper cites Does learning require memorization? a short tale about a long tail.

Sequence-Level Leakage Risk of Training Data in Large Language Models Does learning require memorization? a short tale about a long tail

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.689646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.689646Z digest=sha256:192dbc90fe321178596a9992cce2b7f5fc7fba7c7afd76a64bbbd6a546999d7e

Observation 456f9785-a1c2-4e51-8637-679c58bb63e2 · outbound

This paper cites SoK: Memorization in General-Purpose Large Language Models.

Sequence-Level Leakage Risk of Training Data in Large Language Models SoK: Memorization in General-Purpose Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.694182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.694182Z digest=sha256:2ca77304f0817b177776fb7bc229886c74db91cb2ce15f67eaa9d7d7110a3465

Observation bbfb7957-6198-4f91-ba7f-3126fd49aa71 · outbound

This paper cites Measuring memorization in language models via probabilistic extraction.

Sequence-Level Leakage Risk of Training Data in Large Language Models Measuring memorization in language models via probabilistic extraction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.701547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.701547Z digest=sha256:0dc161b37424d1d93bd6711e1b52f15799410b107ebe955311760d3f4522339a

Observation 103ef7f8-7d4b-4583-a9cd-b654093eaf37 · outbound

This paper cites Are Large Pre-Trained Language Models Leaking Your Personal Information?.

Sequence-Level Leakage Risk of Training Data in Large Language Models Are Large Pre-Trained Language Models Leaking Your Personal Information?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.706508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.706508Z digest=sha256:0ec86a28a3bf2d3f67251c96d7c66fa97302151c17259751afe948fc35825eda

Observation 04733d41-86de-4b99-83d2-eb337497d844 · outbound

This paper cites Preventing generation of verbatim memorization in language models gives a false sense of privacy.

Sequence-Level Leakage Risk of Training Data in Large Language Models Preventing generation of verbatim memorization in language models gives a false sense of privacy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.712445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.712445Z digest=sha256:480c61c089b843269d32de1a08303af920b65e4401346e3c7ea1f596de7f6e9d

Observation 9d85a289-75ec-4ecb-9055-5b354f654f5f · outbound

This paper cites Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs.

Sequence-Level Leakage Risk of Training Data in Large Language Models Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.716503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.716503Z digest=sha256:25124c79fed22a08aafad7b9e2d87eef1bb38ee25acc269169cedd11ca7f3166

Observation a29ef7da-327e-4151-bb9a-c1f292fdf8d5 · outbound

This paper cites an unresolved cited work.

Sequence-Level Leakage Risk of Training Data in Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:09:51.093483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.722385Z digest=sha256:590f86c9b17e889719970c2f2fb8eeac1711052b0466f39a95d2b5d10c7db964

Observation 0da8c77d-9d8c-405f-ab62-41e40e91e513 · outbound

This paper cites Scaling Laws for Fact Memorization of Large Language Models.

Sequence-Level Leakage Risk of Training Data in Large Language Models Scaling Laws for Fact Memorization of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.727089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.727089Z digest=sha256:dea953b6541c0ce003ac8734d4e67c6b71958f5c2f6195596ed0c9647b6a4f19

Observation 7f3b5048-3899-46ac-bfcc-fd7a77b5d186 · outbound

This paper cites Analyzing leakage of personally identifiable information in language models.

Sequence-Level Leakage Risk of Training Data in Large Language Models Analyzing leakage of personally identifiable information in language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:51.065094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.733185Z digest=sha256:a542a9de1b75813b2825277bf21d78eaf21527039e2838c074a5c19e45d05714

Observation d35d609d-137e-4575-82f1-95b787f2ed40 · outbound

This paper cites C., Sedghi, H., Lipton, Z.

Sequence-Level Leakage Risk of Training Data in Large Language Models C., Sedghi, H., Lipton, Z

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:51.044310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.738595Z digest=sha256:d7b9c329a65d6309debb8395f874cff3120c708eb38ea28d035b66656f8af467

Observation 3ef6e066-11d8-48ff-94d2-1890b7955b73 · outbound

This paper cites C hat G P T.

Sequence-Level Leakage Risk of Training Data in Large Language Models C hat G P T

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:51.027772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.742966Z digest=sha256:b6eb3f56186d20ae75eddc33bb101e63ba2230b7997f4c71ce2b6b8ae327f5e6

Observation 9d60a4bf-05fb-47e4-9039-eed3acce549e · outbound

This paper cites Detecting pretraining data from large language models.

Sequence-Level Leakage Risk of Training Data in Large Language Models Detecting pretraining data from large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:51.007931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.747433Z digest=sha256:218e000183e31ed61c23d8973b9d9de13b605fc7acae03d46011ceb2a906555c

Observation 64e3bba5-c49b-4144-a6f6-2dc091111e60 · outbound

This paper cites Memorization without overfitting: Analyzing the training dynamics of large language models.

Sequence-Level Leakage Risk of Training Data in Large Language Models Memorization without overfitting: Analyzing the training dynamics of large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.752799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.752799Z digest=sha256:5001e0aff53589cfb0416e46b998b8961a5936ccc76bf8a3d32d5923f6d4231d

Observation ef3377b9-db58-44d2-8090-b6094a1efae0 · outbound

This paper cites Bag of tricks for training data extraction from language models.

Sequence-Level Leakage Risk of Training Data in Large Language Models Bag of tricks for training data extraction from language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:50.975569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.758964Z digest=sha256:0658d354fcf6a8daa7323c1390c2570330428b150f5afd47bbf64991baa2a188

Observation d474e47b-2bdc-41a0-b8f7-cd1b399c2275 · outbound

This paper cites @esa (Ref.

Sequence-Level Leakage Risk of Training Data in Large Language Models @esa (Ref

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.765890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.765890Z digest=sha256:7c9e7ca5d18667c35c6aacf105916934d8dbef26d6078a900386c2b535591e8c

Observation 659ac40c-5641-4597-8fa8-cc197dce25ec · outbound

This paper cites an unresolved cited work.

Sequence-Level Leakage Risk of Training Data in Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:09:50.771333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:09:50.771333Z digest=sha256:6f5a0f97adf5f632a2fcaba0bc4781dc1aede546e88f5b635580146d38b5247c

Observation ab1648f5-ca93-4ec8-9a80-05d74c8ea4bd · outbound

This paper cites More recently, duan2024membership,shi2023detecting also explore such Membership Inference Attacks on LLMs.

Sequence-Level Leakage Risk of Training Data in Large Language Models More recently, duan2024membership,shi2023detecting also explore such Membership Inference Attacks on LLMs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:09:50.931597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T15:09:50.778525Z digest=sha256:ffe6e4c760ca9af1d64cba3331bd4b5a65227e99636fec0c7a9b5496fb497f9f

Pith citing papers

Observation 1bca4c09-851c-4338-ba73-ae22ed7760f4 · inbound

What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests cites this paper.

What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests Sequence-Level Leakage Risk of Training Data in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:33.766535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:33.766535Z digest=sha256:e8d4524ce809ab02abab68b3123eb763d881ff3e84c33b378d6134020ab36dc9

Observation 3451d5e0-9f5a-4f23-87fe-29db27f6e25f · inbound

Security Considerations for Multi-agent Systems cites this paper.

Security Considerations for Multi-agent Systems Sequence-Level Leakage Risk of Training Data in Large Language Models

Reference 241

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:15:55.156674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T14:12:14.160789Z digest=sha256:28c924f5397735d407ac18dd8ca1849edd44236fe33b6948d2d7a1ff7fd7fb73

Observation 300dba33-ed9d-4f03-b3b9-860171231052 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Sequence-Level Leakage Risk of Training Data in Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:26.834349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:26.834349Z digest=sha256:2e936849ba05ee126f67cb4ee4d36d1176346d883acdce17f460251ba6d76035

Observation d2b4187d-9b88-4266-ab9e-01d5497c9ee6 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Sequence-Level Leakage Risk of Training Data in Large Language Models

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:35.811486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:35.811486Z digest=sha256:1b66d36c30d8ec68f2a25b82011ed23bbe20cc565f4983bda441b3997b5e14b4