Pith. sign in

Paper Citation Record · LEDGER

How to Train Data-Efficient LLMs

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2402.09668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09668 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:03.600455Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3d59eb21-f235-4807-978f-cb47c3275d11 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models How to Train Data-Efficient LLMs

Reference 237

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.315894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:115edaa327ae688834a8a407af081c1519b3684384e700bc8104d0a8bb1f3209

Observation 792957dc-e756-4ca0-bad0-21339c6cd949 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:44:38.337372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:e095a9e3a3d84b5055ce3db2900a3f9269ee33e3cee5c30b90e0e6b58312b305

Observation 149a544f-2fdb-4df2-bf3a-ca4639ca9b4b · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models How to Train Data-Efficient LLMs

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.197578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:6c94e2d432c00fc49df2100105f1ad9333a71837cc2f693e3592f117ab3e6a27

Observation 2e7b6d1c-21d2-49e5-8281-2fa89251ff4d · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report How to Train Data-Efficient LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:22:12.189477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:7275b00aacb6a3d417d56403ef23573fc3f3dafd14ee2720b4bdbad5187c371c

Observation 008d596f-f874-4fec-ae02-f2b8d97f73f7 · inbound

Enhancing LLMs via High-Knowledge Data Selection cites this paper.

Enhancing LLMs via High-Knowledge Data Selection How to Train Data-Efficient LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.600455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:03.600455Z digest=sha256:de13ba4c939e17ad4c2272d73b5b8944b97c7c82ec285bfb8e55bb7af4196848

Observation 75691a37-0af0-4eb6-8cbb-6578d75e4297 · inbound

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain cites this paper.

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain How to Train Data-Efficient LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:54.356737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:33:54.356737Z digest=sha256:2270abb2f0ef37568b64ecf3d5991086e111f0e34b6bce5a4f1d724538649556

Observation 7be11a93-fc4f-456a-b1ce-c09431db1d5d · inbound

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining cites this paper.

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining How to Train Data-Efficient LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:49.740308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:09:49.740308Z digest=sha256:ff4d9eba259b7fd029689767fdc63123dd83fb362157b98aee0a1c05b44c3560

Observation 8e5f661e-1df7-4c21-b0f1-6e0f8d337875 · inbound

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets cites this paper.

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets How to Train Data-Efficient LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:21.726501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:21.726501Z digest=sha256:3a7fb660fdcda570280b28791698c8481f7516e20b9d59cb169328e040772df7

Observation 45a2bb00-27e7-413c-aad1-04bfb0365e04 · inbound

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models cites this paper.

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models How to Train Data-Efficient LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.769180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:19:31.769180Z digest=sha256:b196e0c4b72bc13a9cbd049ad0e23a057edf78abf0b9d9b4c1af81b7ef857e56

Observation baf9518d-8283-4d8e-9693-627a73a6bf4c · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning How to Train Data-Efficient LLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:22.198008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:22.198008Z digest=sha256:93686d007ac116258a0f1197fbc2aba5c0bc2a6a17b7bdffa2a1b96e21a75341

Observation 9d290ed7-dc77-45a3-b4e6-29d20b1418fa · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation How to Train Data-Efficient LLMs

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.750924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.750924Z digest=sha256:1322ab7c8f10b57218e5ac6736132076cfa70185f110afa1d0d878757977ab93

Observation 2a1f4957-80ec-4cd3-a2fd-c1a72c8777ca · inbound

Assessing the Role of Data Quality in Training Bilingual Language Models cites this paper.

Assessing the Role of Data Quality in Training Bilingual Language Models How to Train Data-Efficient LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:57.821522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:57.821522Z digest=sha256:33cdc2d1330644a2e7d925bd68211c934b8752a79304176265210aa6623f5dff

Observation 98bac85f-4f08-4101-acf8-8a2b1e0ef229 · inbound

Disentangling the Roles of Representation and Selection in Data Pruning cites this paper.

Disentangling the Roles of Representation and Selection in Data Pruning How to Train Data-Efficient LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:24.716339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:15:24.716339Z digest=sha256:6533db6d13d7b214bc89a7c4d45f19aaa4b202ca06656f75c4f472cefc07a080

Observation f1040ee3-a77b-4e43-8c1c-ea8522bbea7f · inbound

Efficient Training of Deep Networks using Guided Spectral Data Selection: A Step Toward Learning What You Need cites this paper.

Efficient Training of Deep Networks using Guided Spectral Data Selection: A Step Toward Learning What You Need How to Train Data-Efficient LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:46.325370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:46.325370Z digest=sha256:61f139db52beecfa183ae286e0c3e22452d93ece107a48ed80f5225f32cf0458

Observation 6d867eea-4249-4856-8592-b3db67339c98 · inbound

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs cites this paper.

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs How to Train Data-Efficient LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:44.081602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:56:44.081602Z digest=sha256:2aa3f4a5bd628c0c3273dbf3c18bbd7363ef608613311368c2fc8ea1b629f0f4

Observation d6c6eb5c-08c4-4a4b-a6c8-30527e315b02 · inbound

Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification cites this paper.

Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification How to Train Data-Efficient LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:32.871288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:21:32.871288Z digest=sha256:8da471cec29c91c63af8a2a184cf8ef026090178c192b815b3023db8f53ef9d0

Observation 8490cd5d-1c4f-4bab-95f1-f8eca72731b9 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks How to Train Data-Efficient LLMs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:13.006591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:13.006591Z digest=sha256:81cead57c816cb873cf6d5030943b65cb19d5a1bf7c6c8f24d3cb6d75f355cdc

Observation 3a86b5a3-61a3-4750-a132-2cb73815b823 · inbound

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection cites this paper.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection How to Train Data-Efficient LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:42.206386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:42.206386Z digest=sha256:358ac6a35d54a90af44761fab20353f561bf19971f3b971378663758dc1a25e6

Observation 7860f5b1-4c1e-4780-86e1-89b7fa135eb3 · inbound

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models cites this paper.

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models How to Train Data-Efficient LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:04.764330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:13:24.750244Z digest=sha256:c6df1aabb711fd2ed246699ce71ca7ba7b0d6105510eb46c5218433b271c9a9f

Observation 6b912b53-402d-4d00-bb74-0e15c36c04ac · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:59.080239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:50306b533da58447b453dcc08b91b1cbecf4eca2e89d7d80e72219bdf609b38e

Observation 2b4a269f-1cbb-4eaf-b0d9-3a16f4405fe2 · inbound

KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates cites this paper.

KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates How to Train Data-Efficient LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:41:04.697353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:22:59.883307Z digest=sha256:042d0f09686c2c02e5614ffacde4e9fff3801ed6537e23ca277ef093bf0c3d67

Observation 6601df6f-9c38-41f9-ab46-d3096b43ba4e · inbound

DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models cites this paper.

DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models How to Train Data-Efficient LLMs

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:21:26.951063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:19:11.231416Z digest=sha256:4dbb1bff86aef3722f204155de1eda90f23543ae5d8f45d17efa64217765eb75

Observation 98535e94-9a8d-4b8a-8662-34780db3a8c3 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:11:08.529589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:7f0c59068becd31ad627cda401d1653d9247faf96ec263a941afc0fddd862bb1

Observation af26c0cc-ad3e-43fe-9683-2353c9ae1b0d · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:54.594603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:f3c538f003cf935c38585238d0958e7601c346f00a00ed7ea561d352768d3706

Observation 9fe4cc44-6781-46ce-ab2f-60a76e197dfd · inbound

Accelerated Relax-and-Round for Concave Coverage Problems cites this paper.

Accelerated Relax-and-Round for Concave Coverage Problems How to Train Data-Efficient LLMs

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:05:57.094895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:52:41.972226Z digest=sha256:1b96922df979b96b0c222f7816c7326df27eb47d65ccd2d28990bab9414e6eae

Observation ad0f918b-ca1e-4b43-a0a9-df8466e9b472 · inbound

Reflections and New Directions for Human-Centered Large Language Models cites this paper.

Reflections and New Directions for Human-Centered Large Language Models How to Train Data-Efficient LLMs

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:25:56.843634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:25:46.378350Z digest=sha256:4adaad8b4da06d0bfb6c94c26a10cf7a430adb59daff11fa2b6694f53d74383f

Observation 10576518-8575-43d8-9ea9-7263c128c47b · inbound

Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching cites this paper.

Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching How to Train Data-Efficient LLMs

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:03:15.950487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:58:52.511363Z digest=sha256:bc94c1d5700250e78c43b671c7b4ddcda455a98f41d1a8ddc500a275e5c29f89

Observation 741108f7-c40c-483a-9b45-dd06febbe00e · inbound

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them cites this paper.

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them How to Train Data-Efficient LLMs

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:22:46.237668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T23:19:22.355753Z digest=sha256:e97c03c8fef2a4b7e7940a9e946f8ece354c5d80996ebbf46b525a5e57c7fa9e