Pith. sign in

Paper Citation Record · LEDGER

When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2309.04564.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.04564 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:31:49.697043Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c894b174-2f1a-4007-92fb-7003f519a3fe · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:46:40.287347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:161aa9f23faca0e87e4510d81b5590c5a1093c112839304a05da33801ddbefd0

Observation 2cabdda0-928c-42f8-8775-028b8a8be8ec · inbound

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models cites this paper.

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:07:53.863501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T10:07:53.748795Z digest=sha256:8c455b577cb4dab9532ee5cf20b6d882892e47c63d9e3fd4c004a0fcaad58c9c

Observation 80cb290a-d72a-4454-980d-5e17b1e77046 · inbound

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP cites this paper.

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:38:15.898675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T17:36:18.451771Z digest=sha256:53ef19dbdd98a0b91bb7a71a77248ee71c18883984e7eb42142a97a6cfeea3a9

Observation dc25a264-3f1c-4d21-b0ac-b1506e6ed1bb · inbound

From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation cites this paper.

From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T19:31:49.697043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:31:49.697043Z digest=sha256:d5829855593f0c1cb56c906a2db7c6fcfa1591728dbe0e010acaea4998d60942

Observation 9fa03547-68b8-47b7-b02f-0e39ab57bd3c · inbound

Visual Compositional Tuning cites this paper.

Visual Compositional Tuning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.142850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T17:39:09.890605Z digest=sha256:3682fdf4c409b3d49d142ac04245aa6961a90a416cca12e21b8128a60004e1a8

Observation 58768ad4-65f5-4480-8f5d-e1fd7c9863c9 · inbound

Enhancing LLMs via High-Knowledge Data Selection cites this paper.

Enhancing LLMs via High-Knowledge Data Selection When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:02.614781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:02.614781Z digest=sha256:455d8116781799b4e9848eabec4c57d439a6a6956b65438106255559c142a13f

Observation d1e39d41-ff38-45d0-ba22-d7320806b40d · inbound

Small-to-Large Generalization: Data Influences Models Consistently Across Scale cites this paper.

Small-to-Large Generalization: Data Influences Models Consistently Across Scale When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:26.725536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:26.725536Z digest=sha256:8981c0ac5afe1c30772f67832dd24ab81eb9f23897088e59e11c31ae5487bb1c

Observation 9d0b6135-a217-482e-9dbc-88e0d434b1f0 · inbound

Improving Chemical Understanding of LLMs via SMILES Parsing cites this paper.

Improving Chemical Understanding of LLMs via SMILES Parsing When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:45.356862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:06:45.356862Z digest=sha256:19a64566a030030336a42f80906abd6d3265a2d07d9b1ceb60323d906d93af56

Observation 7de29b8e-dede-49df-828d-71792f8cc102 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 285

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.502620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.502620Z digest=sha256:2e875fff907d7d810ca7452f39c3814461f3808b1b27657c3160ac98a165fa66

Observation d619e2ab-6886-44f8-9a45-7f500f198bcc · inbound

Efficient Data Selection at Scale via Influence Distillation cites this paper.

Efficient Data Selection at Scale via Influence Distillation When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:18.051571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:18.051571Z digest=sha256:2c918a9d674c21c5bfbb7a80917d0489ed8a70008543cdfe9b28b85a5b1c7373

Observation ea4f463e-8f5a-43cf-a59f-4cf597277efc · inbound

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining cites this paper.

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:48.284205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:09:48.284205Z digest=sha256:90ce270949d82fe9145a15a721b92fa578519990af8fdaa3228a5452967074d9

Observation b36ccb73-9dcf-40a7-b18e-44f33a2c34eb · inbound

GORACS: Group-level Optimal Transport-guided Coreset Selection for LLM-based Recommender Systems cites this paper.

GORACS: Group-level Optimal Transport-guided Coreset Selection for LLM-based Recommender Systems When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:21.826147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:56:21.826147Z digest=sha256:2d9dd150e025ae48e512905543492c7700eb01e57684dee029567235de6547c9

Observation 64bb2aec-abbc-4d16-aafe-645de5f60d18 · inbound

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning cites this paper.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.455874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.455874Z digest=sha256:4906c0e2bd74fbf2de5038b0fc31cfaf2b33f44da37cf47978ca0d99b442a71f

Observation 59f6abf4-afe3-47c3-a6ba-d8b95481c1ed · inbound

Efficient dataset generation for machine learning perovskite alloys cites this paper.

Efficient dataset generation for machine learning perovskite alloys When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:07.465093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:07.465093Z digest=sha256:267aed68bd5d3088daa6fefc5ccc008f12efc02de2fee87ebd760a3adfccbe29

Observation ca62f42a-f897-4865-96b0-edae64c870fa · inbound

Disentangling the Roles of Representation and Selection in Data Pruning cites this paper.

Disentangling the Roles of Representation and Selection in Data Pruning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:24.011287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:15:24.011287Z digest=sha256:41e1421967a0af6e375500488092dca36ab66013e183de284f5af9bd62437f30

Observation 6d07eb10-06d0-43e4-8429-1cb456aa0180 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:11.656872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:11.656872Z digest=sha256:1eb577c7bb036d55a7806f3cf0e419e2e6dacff90df1d7aa1c7c0fa1500485e2

Observation f1ed7484-fbb0-434d-9557-1e0cf813a131 · inbound

ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization cites this paper.

ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:00:05.194995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:00:05.194995Z digest=sha256:01645a93f33756e5cfd0b16c9052267f54dba7879c767d6d6b00f1be9aa72507

Observation 2e348e6e-7e5a-4f43-9d9d-a883ad6a9fef · inbound

GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization Geometry cites this paper.

GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization Geometry When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:30:07.825970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T12:26:14.261351Z digest=sha256:eefd788903be247220ff3e0bb6355c9421b7df3ac245aa596df9653dfcd40cb8

Observation 02f8e2ee-dc95-42cb-a6ab-97b1549804ad · inbound

OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation cites this paper.

OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:19.737306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:19.737306Z digest=sha256:f34b06bd9e354754c291565e9f832abf47983d5c7e657df385816b7d95308d6a

Observation f79798d1-047b-49c3-a62f-082c3e17873d · inbound

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics cites this paper.

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:00:31.597557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T02:59:27.807789Z digest=sha256:875430b8c965dd6c9ff8e7142136f14c3190f6c440e63fc64d6f0d68d49d8298

Observation bec684eb-a500-4351-a15f-776c0af3754b · inbound

A Systematic Framework for Tabular Data Disentanglement cites this paper.

A Systematic Framework for Tabular Data Disentanglement When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:55.343421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:27:39.046005Z digest=sha256:cdec548bf6c0e12ba907155d529a7d6d60eb6fdcc356b0ea0b69b3a07735414a

Observation a3ddf1a1-2cb4-416e-b88d-752280d835eb · inbound

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization cites this paper.

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:30:49.021001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:06:46.131725Z digest=sha256:a2a18de8e7eaf4274a28392e28169b28a175e1a8a7f784ccdbc479f1c2418f40

Observation 3a677dd1-e712-45c9-805e-2f19681436fb · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.197284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:82580b8b401c0d7cc47219e691f3a5490654c014053b61fab876bd791b607f1a

Observation 42d6e4f0-05f0-4d7e-b871-71d586e13301 · inbound

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees cites this paper.

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:07.878353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:50:39.734124Z digest=sha256:63a98aa0ec415e3347163d0ddf973db018a6ea37e25e91ffaef3479c20ef2d61

Observation 06f0d75f-3540-4396-bc2f-322592cb7c9d · inbound

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees cites this paper.

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:41:17.710831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T02:38:45.322351Z digest=sha256:722a43f5a0491a8c179760bdb8b48cb26dfe1f7ce369c89f1c19f8b592682717

Observation 9b194a4b-747d-408c-854b-b7680421d8ca · inbound

Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods cites this paper.

Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:11:53.392676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:09:21.652035Z digest=sha256:e12bbacba28e638aa2236817a95561e7b6302fbfdc06184b3b91dea39960f90d

Observation b00f21c8-e6b6-4814-ac7b-59860a533759 · inbound

Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning cites this paper.

Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:07:54.465343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T20:02:58.318276Z digest=sha256:411cf4a1deebb0ea1f2e355e9e6c41dd672c10949ac3155e160b0dca19dcac4a

Observation d5555e70-5a61-4bd6-98b2-0169e925a5cf · inbound

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code cites this paper.

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:08:05.057764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T05:06:46.360174Z digest=sha256:a96a12cebac983e8569ba838b3024da5f1be20a950f68e9c23522705dca575ba

Observation db30ffd2-cf6e-4225-839f-d856a63944c6 · inbound

Unified Data Selection for LLM Reasoning cites this paper.

Unified Data Selection for LLM Reasoning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.137379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-22T05:33:20.930156Z digest=sha256:66f376a9f642ddadb97d81bc91052afc78390a55578e1b5cfcb170d5b5a59af8

Observation ab1d612f-3fa9-44d0-923c-f018258698cf · inbound

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection cites this paper.

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:13:30.110956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T14:08:40.968105Z digest=sha256:11d4c2c96a383e3fe2801edd6aacc9efb1e078d958532ecb0a6db642c97b4e06

Observation 9977346f-1710-4a74-a843-032e2e0f72e6 · inbound

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework cites this paper.

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:27.218492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T18:26:32.883834Z digest=sha256:0631b04a3d080643736894f3695fe60a97e1873a865084393ffc834ad4a0eab7

Observation 8b75c3a5-0934-4ce1-a63d-ccd01945265c · inbound

Data Selection Through Iterative Self-Filtering for Vision-Language Settings cites this paper.

Data Selection Through Iterative Self-Filtering for Vision-Language Settings When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:49:44.946514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T09:22:47.537137Z digest=sha256:e3430694700bc764352c8b16996fc3ccfaca33d30a3a889359e39c7bd8d14559

Observation ddfb4caa-9c82-4503-86fc-8301fe8837a6 · inbound

Internal Data Repetition Destroys Language Models cites this paper.

Internal Data Repetition Destroys Language Models When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.884375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T00:12:56.745617Z digest=sha256:a9f014e5ef56949e72ae192486e7142e9fe0ad2c68fbb39c98ef75a17fb3bf2a

Observation 870d7ae7-8728-4ea3-8419-a661c6fd1a08 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 216

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.707543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:508c9b4c94f57cbbb075a11835a9b52e9fc92e0e3709aa6df4834ddeec7baf07

Observation 94a9d6c8-a587-4be7-8fed-9f93f61025bc · inbound

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference cites this paper.

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T11:19:34.319148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:19:34.319148Z digest=sha256:abbfa877e99e5416d9f35d7718566a0e5fb16bbcc18764965a49ea977025a4cd