Pith. sign in

Paper Citation Record · LEDGER

How to Train Data-Efficient LLMs

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2402.09668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09668 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:40:15.672658Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3d59eb21-f235-4807-978f-cb47c3275d11 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models How to Train Data-Efficient LLMs

Reference 237

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.315894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:b4f657856301c7ced21a158433da8be3b1a9abc00b6262eff18351c09a91c3c9

Observation 792957dc-e756-4ca0-bad0-21339c6cd949 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:44:38.337372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:ee04592b017622636bd600a4b7bcac49622e9e4eda7eb4b1e2b6654e51dce1ad

Observation 149a544f-2fdb-4df2-bf3a-ca4639ca9b4b · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models How to Train Data-Efficient LLMs

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.197578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:82324be0ac4d5e0da5055446066963994c09e82292a147929b642189862f06d4

Observation 61de4bd6-3171-4400-8060-c1e13931fd20 · inbound

Efficient Alignment of Large Language Models via Data Sampling cites this paper.

Efficient Alignment of Large Language Models via Data Sampling How to Train Data-Efficient LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:40:15.672658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:40:15.672658Z digest=sha256:6f3b79215221b1dfa1e17dedfa68907a5c7e81f7d7daf3dc97b14fb0788dd590

Observation c3d4efd7-58b9-4bca-be02-88d5c24569b8 · inbound

LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models cites this paper.

LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models How to Train Data-Efficient LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:45:47.427739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:45:47.427739Z digest=sha256:dadc81c33e1fe6c6f836d31f813bf56c5d816bc255677e314f3e668e8855f805

Observation 34c8fa16-7965-4daa-aa31-b2f4b70e9ac3 · inbound

Training Bilingual LMs with Data Constraints in the Targeted Language cites this paper.

Training Bilingual LMs with Data Constraints in the Targeted Language How to Train Data-Efficient LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T17:04:34.463318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:04:34.463318Z digest=sha256:c729d3e3f8f95babdefca45c19512dc106a5b93193798aaae6601290991aa48f

Observation ec3b47f7-da73-41fe-a6ce-030286bc9153 · inbound

7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement cites this paper.

7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement How to Train Data-Efficient LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T20:27:18.187919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:27:18.187919Z digest=sha256:6a5bed395204ea8ea0c5be695a5db812837278595dcf52c3d37567cff691dfd9

Observation d1c50baa-eef8-43cc-a5b4-0be00d64494a · inbound

Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data cites this paper.

Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data How to Train Data-Efficient LLMs

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-10T17:34:38.911261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:34:38.911261Z digest=sha256:32a032af52b82ff1c0892f61168c89805186db21f564741d5054d9b022aa009f

Observation 102ae0a7-155c-48e5-80d8-c6c9c7ab32da · inbound

FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training cites this paper.

FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training How to Train Data-Efficient LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T17:54:00.454980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:54:00.454980Z digest=sha256:cb203e505eebce91033fcd6a24530e8bd3eaf765f8d1637fe85b0768f8076b56

Observation fd3cbe97-8165-41a3-9aa7-4928f63c3f7a · inbound

FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy cites this paper.

FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy How to Train Data-Efficient LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T18:58:34.553634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:58:34.553634Z digest=sha256:2ae8506807c8750d8a5181bd35764826f39b7663186d1af5323cadc528f795a6

Observation 1d43c75f-11a3-48e7-8928-a35bf4795c4b · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts How to Train Data-Efficient LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.347017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.347017Z digest=sha256:0dc9ac29a96c91e5d9f37d3b19cbace1429588d6fa57ea593ad1877db2aeb115

Observation 2e7b6d1c-21d2-49e5-8281-2fa89251ff4d · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report How to Train Data-Efficient LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:22:12.189477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:56891698d33687bb7ab21841f0c788a1e0cb6dec042afa8eafdc6046758726ab

Observation 008d596f-f874-4fec-ae02-f2b8d97f73f7 · inbound

Enhancing LLMs via High-Knowledge Data Selection cites this paper.

Enhancing LLMs via High-Knowledge Data Selection How to Train Data-Efficient LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.600455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:03.600455Z digest=sha256:bc3d93c502cb4665e4e293cab68034a8c79193280d750bc3ca27fb8b7bcd036b

Observation 75691a37-0af0-4eb6-8cbb-6578d75e4297 · inbound

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain cites this paper.

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain How to Train Data-Efficient LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:54.356737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:33:54.356737Z digest=sha256:2bc835f0f2b9e036449712d0b57d0e1e1218d3120773434e4c030df2e2e5542b

Observation 7be11a93-fc4f-456a-b1ce-c09431db1d5d · inbound

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining cites this paper.

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining How to Train Data-Efficient LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:49.740308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:09:49.740308Z digest=sha256:9171384abd52842eb58314bc680a58dd9ac832f7b82968b5353b29e1a0ecc711

Observation 8e5f661e-1df7-4c21-b0f1-6e0f8d337875 · inbound

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets cites this paper.

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets How to Train Data-Efficient LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:21.726501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:21.726501Z digest=sha256:da2d6e1216adecfcdf661cba66df86dc7fd9eab3391da326c0b0c4cb156e84a6

Observation 45a2bb00-27e7-413c-aad1-04bfb0365e04 · inbound

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models cites this paper.

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models How to Train Data-Efficient LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.769180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:19:31.769180Z digest=sha256:bb5ae30aedadf555594b10db00f46a6355f334aa7a8c0713ed171545ec8bcc9a

Observation baf9518d-8283-4d8e-9693-627a73a6bf4c · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning How to Train Data-Efficient LLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:22.198008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:22.198008Z digest=sha256:12719bac57443a64032add99830aaccfa46dd73ad39458ff87866795cc430f00

Observation 9d290ed7-dc77-45a3-b4e6-29d20b1418fa · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation How to Train Data-Efficient LLMs

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:09.750924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:09.750924Z digest=sha256:8d9fb771d2eaafadf4ca7c6748c532f9e01fb01e68470b1d8f1ae052f30db305

Observation 2a1f4957-80ec-4cd3-a2fd-c1a72c8777ca · inbound

Assessing the Role of Data Quality in Training Bilingual Language Models cites this paper.

Assessing the Role of Data Quality in Training Bilingual Language Models How to Train Data-Efficient LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:57.821522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:57.821522Z digest=sha256:18f9a4f1a78651d5a66e6a929ae538ea760b23d3968764fcd22e33348be8c18e

Observation 98bac85f-4f08-4101-acf8-8a2b1e0ef229 · inbound

Disentangling the Roles of Representation and Selection in Data Pruning cites this paper.

Disentangling the Roles of Representation and Selection in Data Pruning How to Train Data-Efficient LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:24.716339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:15:24.716339Z digest=sha256:a26584f7a3a724ef89b699bfa6e70e9dc1cc7d622c274e0f8708db7a486024eb

Observation f1040ee3-a77b-4e43-8c1c-ea8522bbea7f · inbound

Efficient Training of Deep Networks using Guided Spectral Data Selection: A Step Toward Learning What You Need cites this paper.

Efficient Training of Deep Networks using Guided Spectral Data Selection: A Step Toward Learning What You Need How to Train Data-Efficient LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:46.325370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:46.325370Z digest=sha256:277d4865be5f27ea2942fad4a58ef3df6d14d399c50f403a6963d2847280482d

Observation 6d867eea-4249-4856-8592-b3db67339c98 · inbound

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs cites this paper.

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs How to Train Data-Efficient LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:44.081602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:56:44.081602Z digest=sha256:95def3b3a3926a82aa5bbfba9f6a773d9c3a44e751d573969f3e7fc19e0fd46e

Observation d6c6eb5c-08c4-4a4b-a6c8-30527e315b02 · inbound

Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification cites this paper.

Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification How to Train Data-Efficient LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:32.871288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:21:32.871288Z digest=sha256:2ee9e0ae22e1a237c0123162f278dfb98381b7e5709e0235f3164e81f23544e3

Observation 8490cd5d-1c4f-4bab-95f1-f8eca72731b9 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks How to Train Data-Efficient LLMs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:13.006591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:13.006591Z digest=sha256:dbde21eb71d071dce0756257ed291c1440e07defc4984b7d9d5aef9138ac0976

Observation 3a86b5a3-61a3-4750-a132-2cb73815b823 · inbound

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection cites this paper.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection How to Train Data-Efficient LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:42.206386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:42.206386Z digest=sha256:5de82298435789a67c8ccafb74d449bb5cc8e01cf540c10956ac2c473cf6b5f8

Observation 7860f5b1-4c1e-4780-86e1-89b7fa135eb3 · inbound

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models cites this paper.

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models How to Train Data-Efficient LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:04.764330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:13:24.750244Z digest=sha256:b68969fdc0dc5224f5a5ce951b306a0dcb260d1ef3b54a6aea42419ac64a56a0

Observation 6b912b53-402d-4d00-bb74-0e15c36c04ac · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:59.080239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:14506df99e4ef3dd2a39d4e76bf9bcd003eee537fd707ec0ed2a89fec7214be2

Observation 2b4a269f-1cbb-4eaf-b0d9-3a16f4405fe2 · inbound

KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates cites this paper.

KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates How to Train Data-Efficient LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:41:04.697353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:22:59.883307Z digest=sha256:b3df40371ec1483f686d397d281cdc59399f3477e563c1a152634ac3ef837687

Observation 6601df6f-9c38-41f9-ab46-d3096b43ba4e · inbound

DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models cites this paper.

DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models How to Train Data-Efficient LLMs

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:21:26.951063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T06:19:11.231416Z digest=sha256:a5af5134a154c1c037d1b9f7b1c66469e4b142e49a77d726332f73227772d72a

Observation 98535e94-9a8d-4b8a-8662-34780db3a8c3 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:11:08.529589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:0abbf36b2ccf0104bb5a4e0721d23e28803ee496896672b2c26c2eeea144e57c

Observation af26c0cc-ad3e-43fe-9683-2353c9ae1b0d · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence How to Train Data-Efficient LLMs

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:54.594603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:663edd442d916f803569e17ef9c8bc5b629ddf106d91de095c0810863573ac2d

Observation 9fe4cc44-6781-46ce-ab2f-60a76e197dfd · inbound

Accelerated Relax-and-Round for Concave Coverage Problems cites this paper.

Accelerated Relax-and-Round for Concave Coverage Problems How to Train Data-Efficient LLMs

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:05:57.094895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T00:52:41.972226Z digest=sha256:d6abb911c211c437eda338b50b6d2e7e0ad7f75a39f2647dc61933f042322408

Observation ad0f918b-ca1e-4b43-a0a9-df8466e9b472 · inbound

Reflections and New Directions for Human-Centered Large Language Models cites this paper.

Reflections and New Directions for Human-Centered Large Language Models How to Train Data-Efficient LLMs

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:25:56.843634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:25:46.378350Z digest=sha256:4c0db3b4d69607eda58d63e0a38a8608c37d9a97dcd6a3b10e52a0fca3a87f3b

Observation 10576518-8575-43d8-9ea9-7263c128c47b · inbound

Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching cites this paper.

Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching How to Train Data-Efficient LLMs

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:03:15.950487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T08:58:52.511363Z digest=sha256:d3f0562c92502990dc3dde09e437cf4f36f36f151a5459d755d2703041f45f7c

Observation 741108f7-c40c-483a-9b45-dd06febbe00e · inbound

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them cites this paper.

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them How to Train Data-Efficient LLMs

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:22:46.237668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T23:19:22.355753Z digest=sha256:df38e32eb450be63d21d5f318fb53899ad13409d1e1a144890014bc550f145a5