Pith. sign in

Paper Citation Record · LEDGER

When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2309.04564.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.04564 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:15:24.011287Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c894b174-2f1a-4007-92fb-7003f519a3fe · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:46:40.287347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:537a8f057ed86da1dee8f4412a045a826c40b4341186f3e9e0a5141b69520e36

Observation 2cabdda0-928c-42f8-8775-028b8a8be8ec · inbound

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models cites this paper.

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:07:53.863501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T10:07:53.748795Z digest=sha256:984f56d84eb9c2d9cd8094f1772d5cedecbba146a0510b58d0ca460235451b8e

Observation 80cb290a-d72a-4454-980d-5e17b1e77046 · inbound

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP cites this paper.

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:38:15.898675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T17:36:18.451771Z digest=sha256:26ec752f4e132b438f19c469a4e367d0a709e20f654188633cbc80797dd8431d

Observation 9fa03547-68b8-47b7-b02f-0e39ab57bd3c · inbound

Visual Compositional Tuning cites this paper.

Visual Compositional Tuning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.142850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T17:39:09.890605Z digest=sha256:4a019a7cb28817782bef603d9f0dd7d61c5d5f16e14b3b91ddf2e5fa3809589f

Observation ca62f42a-f897-4865-96b0-edae64c870fa · inbound

Disentangling the Roles of Representation and Selection in Data Pruning cites this paper.

Disentangling the Roles of Representation and Selection in Data Pruning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:24.011287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:15:24.011287Z digest=sha256:77b58cf48e01701121c5ed936eb73d98853ba97b3cabb2521deeed3d1f9d42f3

Observation 6d07eb10-06d0-43e4-8429-1cb456aa0180 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:11.656872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:11.656872Z digest=sha256:3b53c789729f00f204efcf61b77d6aa9cedd4e46b8d09c51b0f34c124fe7683a

Observation f1ed7484-fbb0-434d-9557-1e0cf813a131 · inbound

ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization cites this paper.

ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:00:05.194995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:00:05.194995Z digest=sha256:3aac04e3015e5be49e862cbc386471f7699e742bb293871883a3ba6bbf57a64e

Observation 2e348e6e-7e5a-4f43-9d9d-a883ad6a9fef · inbound

GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization Geometry cites this paper.

GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization Geometry When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:30:07.825970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T12:26:14.261351Z digest=sha256:c9720cfeac4168d03f4e5a689807f16cc574d6353b9a9a6bd32677986a80aa57

Observation 02f8e2ee-dc95-42cb-a6ab-97b1549804ad · inbound

OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation cites this paper.

OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:19.737306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:19.737306Z digest=sha256:0fd109d1a138bf7b68b6ad7a6cc94957322eb52561099cfa629d8190eb309bfb

Observation f79798d1-047b-49c3-a62f-082c3e17873d · inbound

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics cites this paper.

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:00:31.597557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T02:59:27.807789Z digest=sha256:5dfcfcf40f37ef10f49667d75016ac3d01f7745a91d4deeb3a2c20027b47bcf2

Observation bec684eb-a500-4351-a15f-776c0af3754b · inbound

A Systematic Framework for Tabular Data Disentanglement cites this paper.

A Systematic Framework for Tabular Data Disentanglement When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:55.343421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:27:39.046005Z digest=sha256:2ab42e66c6a75c12cdf26a74716ca17ca2da599776c578f60c737bc0ddb421b5

Observation a3ddf1a1-2cb4-416e-b88d-752280d835eb · inbound

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization cites this paper.

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:30:49.021001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:06:46.131725Z digest=sha256:26633efddfb98be62c5969a78fd40b33c6f834c20d66421336c6cd397c2d9b2a

Observation 3a677dd1-e712-45c9-805e-2f19681436fb · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.197284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:d59dab2c2dd24545d878a595b5e62257628e0e83db97b53ae42e58fa28cdf410

Observation 42d6e4f0-05f0-4d7e-b871-71d586e13301 · inbound

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees cites this paper.

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:07.878353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:50:39.734124Z digest=sha256:4a9c7e060b69db10ec38c413586de6cac8bc3f3eb2132172363c71fcc9344455

Observation 06f0d75f-3540-4396-bc2f-322592cb7c9d · inbound

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees cites this paper.

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:41:17.710831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T02:38:45.322351Z digest=sha256:f9f36883fe61f16d813c5e37d905f3660021d0911f4aefe36be4ef148f926a93

Observation 9b194a4b-747d-408c-854b-b7680421d8ca · inbound

Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods cites this paper.

Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:11:53.392676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:09:21.652035Z digest=sha256:7e8617c4b438983fd6a09dbe2a80713581200a2580e6f4d46179fafc55726d05

Observation b00f21c8-e6b6-4814-ac7b-59860a533759 · inbound

Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning cites this paper.

Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:07:54.465343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T20:02:58.318276Z digest=sha256:755a2f6f654189e9d3a01356bc9068bc33e085b1f91e4877c244800526db5c2e

Observation d5555e70-5a61-4bd6-98b2-0169e925a5cf · inbound

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code cites this paper.

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:08:05.057764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T05:06:46.360174Z digest=sha256:9f30e7de24a4bbd1671e7ed7077bb7e766223ce1d74fd21e2f2fad167bec3ebe

Observation db30ffd2-cf6e-4225-839f-d856a63944c6 · inbound

Unified Data Selection for LLM Reasoning cites this paper.

Unified Data Selection for LLM Reasoning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.137379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T05:33:20.930156Z digest=sha256:04ef399cb356590c4655d9c7274e3efc7212e267d0e31b20c5def3f8ba4a2030

Observation ab1d612f-3fa9-44d0-923c-f018258698cf · inbound

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection cites this paper.

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:13:30.110956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:08:40.968105Z digest=sha256:0d5bde2f6b9a33b3e8224ac27dd8f3595cc0bca8280761166ee4379c2f1b5893

Observation 9977346f-1710-4a74-a843-032e2e0f72e6 · inbound

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework cites this paper.

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:27.218492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T18:26:32.883834Z digest=sha256:74bd648eb302ea8f54bd4001db01a9a20de5c5132b6ed05ef83f1b86d68c8278

Observation 8b75c3a5-0934-4ce1-a63d-ccd01945265c · inbound

Data Selection Through Iterative Self-Filtering for Vision-Language Settings cites this paper.

Data Selection Through Iterative Self-Filtering for Vision-Language Settings When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:49:44.946514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T09:22:47.537137Z digest=sha256:7abf9682a37ebe350dbe859a53e9b853219a2a2014ceaa78e75c9eebf91c6d8a

Observation ddfb4caa-9c82-4503-86fc-8301fe8837a6 · inbound

Internal Data Repetition Destroys Language Models cites this paper.

Internal Data Repetition Destroys Language Models When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.884375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:12:56.745617Z digest=sha256:634dd4a02ada4323adfa4643b86c2d35f896c76dcdc1e0553e1ed7f2e7f4b7f0

Observation 870d7ae7-8728-4ea3-8419-a661c6fd1a08 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 216

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.707543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:3efbd46a8bdfcb1f2c139a4a5183ab496a0c0d9ab91b1191858b27cf0a9a78c5

Observation 94a9d6c8-a587-4be7-8fed-9f93f61025bc · inbound

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference cites this paper.

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T11:19:34.319148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:19:34.319148Z digest=sha256:c48fe262e4b4b506e1f99cac2ee48168fa9f379909512d0c29d888575d9a70ac