Pith. sign in

Paper Citation Record · LEDGER

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2506.19262.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19262 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:19.055117Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:53:39.113934Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy6
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 645864b0-6078-460d-894c-0f73a24b1c66 · outbound

This paper cites Phi-4 Technical Report.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.851748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.851748Z digest=sha256:cfa25c20c20e80af129fd1654b3735a2436c6c8007b9781f8a102d70c2f9d98a

Observation bba064c8-de37-4c7e-84c8-f3930b16b45b · outbound

This paper cites GPT-4 Technical Report.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.857973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.857973Z digest=sha256:c9035d9407d9ef25fb5c8ce6393c0437770803ace5643f1ec1749dfb183e03f4

Observation 7043876b-332c-475d-a4fb-22d500c877cd · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.863690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.863690Z digest=sha256:4622fa5f06ce4e750f5e51dcb0ad52ad714367b1c1e1c152f7c909af5a9e020a

Observation 6492b30d-3ec5-4f39-a21b-099f9cb35650 · outbound

This paper cites Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.868972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.868972Z digest=sha256:c501906195d02c284e5c2703d5e545dcc4897cdb456b9f79efaec394d101bcc5

Observation b6c58840-c707-4615-bd4a-26565c4b61e0 · outbound

This paper cites Language GANs Falling Short.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Language GANs Falling Short

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.875688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.875688Z digest=sha256:afc0bcf41128d0e8bd806a4fa5b60036cc5c8b2e514444b6b8ac53bffc660ea6

Observation 89491718-442f-463f-aaa7-146c097564c0 · outbound

This paper cites On the Diversity of Synthetic Data and its Impact on Training Large Language Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning On the Diversity of Synthetic Data and its Impact on Training Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.881689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.881689Z digest=sha256:180981759bec0541af1c0dbb8271b4a4a2d87be1cee46db9ecd7fb8736510086

Observation 5567f7ba-0961-48e8-9080-1a02a7cf09ed · outbound

This paper cites Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.888836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.888836Z digest=sha256:412fed5f7f03594a326b99d1925cde803b1575fbd3b03cf1188d11a17e7a2235

Observation 505260d3-412a-48a7-a24d-6128c7fccdf9 · outbound

This paper cites AugGPT: Leveraging ChatGPT for Text Data Augmentation.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning AugGPT: Leveraging ChatGPT for Text Data Augmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.893945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.893945Z digest=sha256:78fda0406d3be8019f0d00df7fb074e49d539a25cc4f86779cce526eb12408bb

Observation 0df68e44-eaa5-4ce3-a3f6-002123cd2fe3 · outbound

This paper cites Universality of the π2/6 pathway in avoiding model collapse, 2024.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Universality of the π2/6 pathway in avoiding model collapse, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.787288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:18.899052Z digest=sha256:ad87c0af74f071f98ccea380f24c77b95ac8d0a17d58936fbc016e80fed244f6

Observation 42019cdb-260c-4c0f-a28c-833ec6c747c1 · outbound

This paper cites Is GPT-3 a Good Data Annotator?.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Is GPT-3 a Good Data Annotator?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.904149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.904149Z digest=sha256:cf3d7043dccf5a2e660e5cb97c24b2896c6d6a743d71442b70dd15dfa3a65095

Observation 8c071fe3-c614-4553-ba3d-e2f9dc40d8b9 · outbound

This paper cites Data augmentation using llms: Data perspectives, learning paradigms and challenges.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Data augmentation using llms: Data perspectives, learning paradigms and challenges

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.769259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:18.909378Z digest=sha256:94f265922bbb5a27463d80c29afb7339b179a8e918ff06df24af1d26f16b2da2

Observation e53f2ab7-2b3f-4ea7-98ae-47bb7a5f405d · outbound

This paper cites Model Collapse Demystified: The Case of Regression.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Model Collapse Demystified: The Case of Regression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.914754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.914754Z digest=sha256:7556207f31ff1dbfde412f3ff96fbc19d65cc5707e1532075c685a63351e3682

Observation 5bcb2006-a0a3-4586-9a10-828463c2288e · outbound

This paper cites Strong Model Collapse.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Strong Model Collapse

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.920822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.920822Z digest=sha256:43e036e85a603bc780a2f19bb0758d693d46dcc9a8c6694b68be9a75567c331a

Observation fd1d13bd-5531-4d84-ac15-50c6891f3cbe · outbound

This paper cites A Tale of Tails: Model Collapse as a Change of Scaling Laws.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning A Tale of Tails: Model Collapse as a Change of Scaling Laws

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.925506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.925506Z digest=sha256:aabb1461faa362b985b60c129dcb96e3abb444398639850ebd7b154ab47e6c34

Observation c0e572d1-5b86-4daa-95cb-0c81f392a9c2 · outbound

This paper cites The Llama 3 Herd of Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.930500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.930500Z digest=sha256:14b3ceebb3d18ec705a46432d1088b67db8c43320ee8b99396cb19d8529f16ea

Observation 4a98b15e-7d2e-4c28-b846-013385ff65ff · outbound

This paper cites Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.934682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.934682Z digest=sha256:d355cf343aa0b2c1d5c2dc6739b8a365998092bc7e50349ebcbefc6485bb9627

Observation a81f5911-b703-410c-8c4a-ea02c64c4425 · outbound

This paper cites Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.939264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.939264Z digest=sha256:7e6bc079da7be507fd5b397f78e015b20a10a867a8702e979b40c65571ce98ac

Observation 1991e3be-c794-4a86-925d-05e592661c07 · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.943859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.943859Z digest=sha256:a9dbbb9e4624dcb656540b5c50d112e4b0c3c705492c55291321dd4419ec0a3b

Observation d43b8f4d-bb3b-462c-8989-9ac16ce16723 · outbound

This paper cites The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.948737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.948737Z digest=sha256:3c620872e0a054a2fb464a6e6262ff63c1c14a15f7e6f956f0e0ad206d231d61

Observation e7322882-a784-449e-93f6-5e77106b91ab · outbound

This paper cites TarGEN: Targeted Data Generation with Large Language Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning TarGEN: Targeted Data Generation with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.953175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.953175Z digest=sha256:69bc511e1a8d237858837858af5e5e4f3aa5ccee056b2a97c09b75b5007c0e7b

Observation 7de1fa9a-a33c-4f24-8c98-16f7edaedbb3 · outbound

This paper cites Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.957552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.957552Z digest=sha256:32359f85d68ab62209a93bbd5d08a7577897024dbb70108f67500d3471262b70

Observation 551809de-a8ab-45bf-a7ed-46802623ffbc · outbound

This paper cites The narrativeqa reading comprehension challenge, 2017.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning The narrativeqa reading comprehension challenge, 2017

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.962396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.962396Z digest=sha256:fa4178159f58ed26bd16a751cb55c80f29d857a9177be780a11e87a784b5e2e2

Observation 00dbaa3d-0fe8-48ec-ad14-e4843330d71f · outbound

This paper cites Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:19.334959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:18.967285Z digest=sha256:61e7c1f2625c042011fe94e47c4bf2d1a606e17f7557975771ac3364f2397947

Observation 7c60ce21-922e-463b-afa8-1cbbd6de9982 · outbound

This paper cites A diversity- promoting objective function for neural conversation models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning A diversity- promoting objective function for neural conversation models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.732046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:18.972529Z digest=sha256:df455e2e2edf78a08c132b98a263bf3da471ee9c54ae35105ea89073b58d9b95

Observation 5c87fad9-1675-4c54-a5c8-50224cb584aa · outbound

This paper cites Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.977169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.977169Z digest=sha256:f4f672b94be45ae37f6a093456043d862d41dafa77abc0ebc1ce2169b219c0b2

Observation a74c14c8-a7b2-4817-a577-3595866ff611 · outbound

This paper cites API-guided Dataset Synthesis to Finetune Large Code Models.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning API-guided Dataset Synthesis to Finetune Large Code Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:19.293347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:18.981774Z digest=sha256:1ae988098a95af7b6c25c8552adc87515253f839724a064b28961c31738d29cf

Observation ec65d921-5081-405f-910d-cd222a216ae7 · outbound

This paper cites Generating training data with language models: Towards zero-shot language understanding.Advances in Neural Information Processing Systems, 35:462–477, 2022.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Generating training data with language models: Towards zero-shot language understanding.Advances in Neural Information Processing Systems, 35:462–477, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.716531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:18.986716Z digest=sha256:c7cd63df02d7038cb1933cf73e8a67bf4149f5deb2e32103cf9a880ace6d2f86

Observation 04faf82d-35c8-4fb2-95f5-f385cf7333e1 · outbound

This paper cites A corpus and cloze evaluation for deeper understanding of commonsense stories.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning A corpus and cloze evaluation for deeper understanding of commonsense stories

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.990942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.990942Z digest=sha256:a838182a022f6234e3d9906335f5778a9f4bafddefd7181be974524f4800469d

Observation e8b62230-41bd-4f64-ad14-816ffc6addee · outbound

This paper cites I learn better if you speak my language: Understanding the superior performance of fine-tuning large language models with llm-generated responses.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning I learn better if you speak my language: Understanding the superior performance of fine-tuning large language models with llm-generated responses

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.685719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:18.995370Z digest=sha256:b14213bd38fd0584a5df50010d2dce5c92c032532c770909076ca56abdbf0b8f

Observation 9e77a416-a27a-4afa-88bf-749bc2797c13 · outbound

This paper cites How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:18.999831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:18.999831Z digest=sha256:f956961a0dbf94a58d9f0a57fcc90d6e88ef609289756d7ac51f7de7fc186220

Observation 7562b1ae-da84-49a0-98c6-5009f362cd28 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.004646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.004646Z digest=sha256:f75ff1742e605a501969f0882014cb0c26b371a427f70241948377548a3e3737

Observation b8ed6bdd-5545-402b-8ab4-7adc93da7581 · outbound

This paper cites Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Ai models collapse when trained on recursively generated data.Nature, 631(8022):755– 759, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.009713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.009713Z digest=sha256:8de3384bfa38973167b94cca700e6d39609f329869d2085d1bdc3b8cc6e6118e

Observation 9b67ba55-54a6-4052-b681-92f2c2ea78dd · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Large Language Models for Data Annotation and Synthesis: A Survey

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.014617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.014617Z digest=sha256:2660f3862410412d61530064fd4d3915d97bd223a3881c7139ae4fbe103e04ed

Observation ecf7367b-4479-42d4-a809-59f21da280f5 · outbound

This paper cites Evaluating the Evaluation of Diversity in Natural Language Generation.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Evaluating the Evaluation of Diversity in Natural Language Generation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:19.211776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:19.020253Z digest=sha256:8cd6b7496b0fe1ae113a1dd119b6ee96b5c1b2a86dc0e44403d73b9d36126f30

Observation 1b8f8fe9-b838-442d-9eca-3a0233a5f284 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.026402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.026402Z digest=sha256:8cb9ab81dd9cb589d47f8fbe4324926d4fb886f606512c5a477d01d8764f6f08

Observation e757d0cb-d62c-4142-8bc9-97b74cad6e9e · outbound

This paper cites CodecLM: Aligning Language Models with Tailored Synthetic Data.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning CodecLM: Aligning Language Models with Tailored Synthetic Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.031822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.031822Z digest=sha256:95ccc3251981a631082fd05de5087e7b49fd7286f2948d7848fcc9dc9697e02c

Observation 67fb0533-6b1d-479a-83cd-02341d3088a7 · outbound

This paper cites ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.036726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.036726Z digest=sha256:85d9ae2b832f2bd85bc98031bcce9479daad6546f73683676d4e8f9cd093d6aa

Observation 7d3c46d2-027f-47ac-8757-7cd61c6761b9 · outbound

This paper cites ZeroGen: Efficient Zero-shot Learning via Dataset Generation.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning ZeroGen: Efficient Zero-shot Learning via Dataset Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.041283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.041283Z digest=sha256:0c6c5a86d1d2c4d4c4d50e4c5cf7501390da3c0aededcb03fe0bccf36c71517a

Observation 19c461a0-09a6-4266-98e1-70345fc5d622 · outbound

This paper cites GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.045722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.045722Z digest=sha256:994d8d3250ab87beef9a4028ad4f4ac5afa0db4d773e8ec833f08ded990b8b65

Observation 8804a4a2-3eb3-4378-9620-e07eea7c5218 · outbound

This paper cites Large language model as attributed training data generator: A tale of diversity and bias.Advances in Neural Information Processing Systems, 36, 2024.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning Large language model as attributed training data generator: A tale of diversity and bias.Advances in Neural Information Processing Systems, 36, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:12:19.654995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:12:19.050725Z digest=sha256:5a8cb87b124cb212cfb6366714cf64f3c27dc3df8cf0221b4607bc1a9b25db53

Observation 82b574e1-7c95-4cdf-9616-d733186d0f80 · outbound

This paper cites How to Synthesize Text Data without Model Collapse?.

What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning How to Synthesize Text Data without Model Collapse?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:19.055117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:19.055117Z digest=sha256:ae9b81e8fb1990da6c0219ecf935923ec7b11d647211081543f5c335694a7c48

Pith citing papers

Observation f400b560-54ac-4a6e-a286-d40eb55b50fc · inbound

One Joke to Rule them All? On the (Im)possibility of Generalizing Humor cites this paper.

One Joke to Rule them All? On the (Im)possibility of Generalizing Humor What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:39.113934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:53:39.113934Z digest=sha256:9510bd52281f82ef6338023380fa0d3a4bf0aea0dc02dd66cdae2d0663390285

Observation f76b29ef-44c0-4475-9f5d-890516a762cf · inbound

Epistemic diversity across language models mitigates knowledge collapse cites this paper.

Epistemic diversity across language models mitigates knowledge collapse What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-03T15:59:04.798474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-03T15:57:18.174038Z digest=sha256:9881bfb13b559af0e69def06419a78c68112c0182bfcf90486f81e6062095ef2

Observation 41c68f9a-9d4c-4b38-ae8d-500c2ce2c6f5 · inbound

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders cites this paper.

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:11.723762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:11.723762Z digest=sha256:b4c46e24d54e7cf140c2532baa0442ede7e271fe247fd5736f3c7d4d9b3153ea