Pith. sign in

Paper Citation Record · LEDGER

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives

As of 10 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2505.21598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21598 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:33:55.004931Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36bc18a6-cb30-40c1-a97f-4897f475f9cf · outbound

This paper cites A Survey on Data Selection for Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Survey on Data Selection for Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.370150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.370150Z digest=sha256:314383d6fcd8d17fc88c3e264bd1a222a807101c3775f4af6e21cbd1ae0680cf

Observation 62f76901-e1cd-4dd5-b557-a4ee9abbe3a4 · outbound

This paper cites Efficient Online Data Mixing For Language Model Pre-Training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Efficient Online Data Mixing For Language Model Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.487199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.487199Z digest=sha256:85e4df5a72ff535c4eb7cba3e3cafacfac6a4c32952204464ef0beddcb3492e0

Observation 882719b5-0d86-4d79-951b-57fdb5e0f51b · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:58.307035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:48.603636Z digest=sha256:2bd685acbd52f154be2fc2941d246ea506624833f18f644ce046a986ef72a923

Observation b4addb51-6306-41eb-b8c1-294a5259ce6b · outbound

This paper cites Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:57.368547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:48.730422Z digest=sha256:0d9a1826cdb21d6e8f32e60079dbf16e2d1fb4a3d573bfaf05c4224f669c3c00

Observation ab01c965-32fd-4ebb-b0ae-849adfda1d2f · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.893722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.893722Z digest=sha256:6296f04fb25d0ab69e3cbff4b48287f843dc980830e16238fae3cf2c60aaf476

Observation 74c29320-b6e8-4c06-b69c-a39e6b5ba3d9 · outbound

This paper cites Aioli: A Unified Optimization Framework for Language Model Data Mixing.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Aioli: A Unified Optimization Framework for Language Model Data Mixing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.082128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.082128Z digest=sha256:aca0fffddb514e5f65d19aca35a148f2a5002bae87104f6f5b514de1cf27d55d

Observation bcc8f024-6663-461c-8527-49b1cecfbe75 · outbound

This paper cites Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.258242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.258242Z digest=sha256:52630b036eba5f2e8ff7ed9cec2ff4644ca55894a8807435d086a66d2d6bde4e

Observation e3ce5e90-f245-477f-a761-efa5cf0f381d · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Instruction-Finetuned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.385365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.385365Z digest=sha256:e351c63d7a451c2e4861fcb6dd14d37eab7eeec3ff1cf615eba2db4b1399505b

Observation 23546973-3ab1-4667-a1c8-176bd45dbb1e · outbound

This paper cites Training GANs with Optimism.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Training GANs with Optimism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.511514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.511514Z digest=sha256:886fa1c654d7fada723707d2fb430aa8187cc23b8e24c21127aa1d97b8329927

Observation f28c872d-09c2-4910-96ae-4e6d2f031454 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:33:57.103547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:49.614639Z digest=sha256:5771cf1e6cf922c9daafbe567c0d8c828ec6a9fef02d6d9871b80e420f6c58f3

Observation 963250a7-090d-4ab7-9b92-d715f9f2a4b9 · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.759942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.759942Z digest=sha256:d44c4c498ed383fa0f821064eb04805fdf1b881187f2dab200afc5441f10f7bb

Observation fd576d84-729d-4da9-a0a7-2ed46f8e60e4 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.915397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.915397Z digest=sha256:ed9a90a7a76fcd8c64df8dd1a2aa25386e1c94b23f4a97a669bf95624612b30e

Observation b5ee6c42-924a-406c-ba0f-feb001fb6ba3 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:58.127646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:50.039377Z digest=sha256:3bc03d58c082484e40d61130b1da42cda8d5731f1053536c1eee33c4722f11eb

Observation ccbd187d-b65b-45d9-b2fd-9eee1d0eddbc · outbound

This paper cites Forward and Reverse Gradient-Based Hyperparameter Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Forward and Reverse Gradient-Based Hyperparameter Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.797037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:50.178732Z digest=sha256:7382b2e9ae013142f784e36bb5768e01b8eccf86c82f975df2cad6772f4f6256

Observation 1e4ebb40-ae5c-4daf-960c-187a2e144079 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Language models scale reliably with over-training and on downstream tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.325132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.325132Z digest=sha256:b696c2a157e0597b697f585861877be4f0b2509db4e4cea34f2d94a869bb5de1

Observation df30824e-1b19-4511-92de-89ab5d0166ef · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.472683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.472683Z digest=sha256:78eaff6d58797ada5461c86b9b45dc9c1378deebac47711d602e02864fa4ca10

Observation 66437165-7de4-4550-accc-8ec38fb8c18a · outbound

This paper cites BiMix: A Bivariate Data Mixing Law for Language Model Pretraining.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.604941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.604941Z digest=sha256:1b02be71e701ff1221efc24eed9ed2fd7602d22d165ddca4eea5fe3450642c46

Observation dea4f49c-d8e5-4686-86dd-5cdaa69a665a · outbound

This paper cites Scaling Expert Language Models with Unsupervised Domain Discovery.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Expert Language Models with Unsupervised Domain Discovery

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.715368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.715368Z digest=sha256:d2710e718b5bb73838ddd332e5ae93e64dd26f17b1673530b185390d73e1359f

Observation 40a15f17-2fd7-4935-88c7-255670c03bb2 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Training Compute-Optimal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.850028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.850028Z digest=sha256:70eddd98894e7d6402d183eee2ac8bcec01b15bacb38035907ba3e5809ed8e36

Observation b311690c-0dce-4614-b305-f00efb81c32c · outbound

This paper cites Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.960916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.960916Z digest=sha256:43130479d768bc0f6a28bf21e36de1e964c808ee61dc2a3bb8a5b040b1a6284d

Observation 108619fd-ffb3-4a8e-a9bb-e1fe273c38d7 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.071911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.071911Z digest=sha256:37f079cb90798834954b277a1792096ddc0bc89d5d063f8281927da556b6ae8f

Observation c14ea77e-2f43-4978-9bd0-6ac2aab41cba · outbound

This paper cites Scaling Laws for Neural Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Laws for Neural Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.178477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.178477Z digest=sha256:796198ed683f3225b7f358d650fa3ce0fb1d98fbb8827c9d75a09c420569f322

Observation 7f11cfa6-090e-4ecf-a0a8-b8fda047eca8 · outbound

This paper cites A Fully First-Order Method for Stochastic Bilevel Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Fully First-Order Method for Stochastic Bilevel Optimization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.324802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.298763Z digest=sha256:468e79c0f0db7eaad75e6c5c8fb1d174d7e3cf1ba101a8954d1050dd8d0796a9

Observation e5192bdd-f44b-4a8a-9ca8-ee3e999773ef · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.971782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.414889Z digest=sha256:f59de3d53d48840293af041008307e02c613b51f2c71beeef95d500398dd1ad6

Observation 72f6d54c-d65e-473f-b53d-f57aa729b933 · outbound

This paper cites MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.548263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.548263Z digest=sha256:8ac29628c92224c4786d7d942edfd8270461f5c441ecdee5d897f66cae4130f5

Observation 1004f965-6e3d-4c85-8453-fc39bb0c4af4 · outbound

This paper cites FAMO: Fast Adaptive Multitask Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives FAMO: Fast Adaptive Multitask Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.661519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.661519Z digest=sha256:a371b1d5b084f1900f0b56c9e1ce0cce3bcbd1096ba92cb2a670b169a2bc84bb

Observation ae44435e-3711-4976-ac39-7797fd167ca0 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.813615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.772616Z digest=sha256:f16a0c84bf7d2e8c1aaab477c4a1669ee385ea1ca9412cc6a48beae29aa503e1

Observation ea6153a1-ef08-4c32-bce6-d6fbdf085acc · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.884698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.884698Z digest=sha256:7f00555e836d240612cb53a3f14c8e73c6189eaeda18c8afb542130264203a07

Observation 5f0f9683-56eb-49ac-9295-c930cdaa501e · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.644662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.984488Z digest=sha256:08a6c3ba7ee091e4a6449da01deabe14021a898c6ad272dd08dd30c60ad719e4

Observation 02ee085f-bef2-4e81-9cac-ed70b0ea99f1 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.124575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.124575Z digest=sha256:a71476118f61686f0bf0415aa513c3968c762739e450feaaeb617403f0331544

Observation 5b8c3ab1-ee23-4537-86aa-73cb9e83704b · outbound

This paper cites Optimizing Millions of Hyperparameters by Implicit Differentiation.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Optimizing Millions of Hyperparameters by Implicit Differentiation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.282275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.282275Z digest=sha256:16db6de6b41e4404239455e3b3d4a1228ef5e277589b5f70854aa72485e13bc5

Observation a8254859-93c3-45ee-ba9e-011c5d60762b · outbound

This paper cites Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.387698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.387698Z digest=sha256:384d2802ee176bdb495ccd6ac044fb497881e7a3e2cf639389cecf3eba12a646

Observation 583b578c-c295-45fa-840c-1ee032eaad82 · outbound

This paper cites OpenELM: An Efficient Language Model Family with Open Training and Inference Framework.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives OpenELM: An Efficient Language Model Family with Open Training and Inference Framework

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.514060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.514060Z digest=sha256:3f4197664ff596c06f8fbd1f1e5c9fab4b67535cb5cc66d8774f997009a3739c

Observation 9d15914f-4f31-498a-8937-b25454cee021 · outbound

This paper cites Hashimoto, and Percy Liang.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Hashimoto, and Percy Liang

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.622870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.622870Z digest=sha256:1c264e2184d69d48b5ff0484d173281b845a10f469d40b18cf8ae9a98e4196cc

Observation 47ac372b-45d2-4da2-843c-777b6e795982 · outbound

This paper cites LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.754333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.754333Z digest=sha256:77c7aa42d48b7b3ef7a66b6e4b642a231af9fef717b6bc0754db7d952af358d5

Observation 29e0cfed-d8b5-454a-b07d-51453c2f516d · outbound

This paper cites ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.895760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.895760Z digest=sha256:2e346a4d2f6a87db191a9adcf9085db5694e862337687d0e20cda45d50951830

Observation 16c2a519-496c-4d04-b4fa-96a085159c1a · outbound

This paper cites D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.994596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.994596Z digest=sha256:0800bf4f31b13d21c046cef1fd5456834fd663a756a8112afaa961b22f34b8e7

Observation 6914823b-34a4-443b-93df-1ad0ec8df797 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.139442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.139442Z digest=sha256:48ca52f664510e72063e68f3ce9fc7bc590a51dad4bc3d65c91e6cb145f3cb53

Observation af4c5844-478f-4d49-979d-0b99705c36cc · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 39

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:33:55.904518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:53.242787Z digest=sha256:1e0cc9dfd08c892d830da5d094b94469583b6aa7c450750b733095dad8a4af67

Observation 70d827fa-3469-459e-8020-9b16f8b81515 · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.325891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.325891Z digest=sha256:40f571e6a62b567340e3e79c2e7c9206c745ab2fd2194cd91f6c7fb793ff29c4

Observation 7fe70179-7f88-46f0-8088-731f97cbb0cf · outbound

This paper cites SlimPajama-DC: Understanding Data Combinations for LLM Training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives SlimPajama-DC: Understanding Data Combinations for LLM Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.419429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.419429Z digest=sha256:96879dff7d48848c6959c7c0135e85b7b3cc13f4f8237da5f0c63ad0612157fc

Observation b6821ead-e6e2-44de-843c-4a7f822e84a6 · outbound

This paper cites The Vizier Gaussian Process Bandit Algorithm.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives The Vizier Gaussian Process Bandit Algorithm

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.554390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.554390Z digest=sha256:d31e7a97581b41cabdf9c9204cea2c78dd86ae677c20bfcd53455c9db26075f9

Observation 66e1e1d6-b288-4248-a255-e95759f080ca · outbound

This paper cites Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.715579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.715579Z digest=sha256:8f61751ee3ba87d89dac21517ee38a009b25d351ff97e8436bdb1fce74dcc866

Observation 1c3b19b2-514e-4dc4-8c72-bee9fdeb6b64 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Finetuned Language Models Are Zero-Shot Learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.834351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.834351Z digest=sha256:05345caa299d733b7f96bca21ffcab0c870d356cff9928e8b7a9416491090afb

Observation 3597c35a-4d91-4682-82bc-7cb147cc2dab · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.990057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.990057Z digest=sha256:4342386c8a5cf560a5138822c839165601e1db0cf53010c08013bfd961b353b7

Observation 8d4d2737-a203-4e5d-a27c-6133fb09a45f · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.125833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.125833Z digest=sha256:93ffb6e85767be8eb65fdd781896067941b3071201b45333d68362f0ddf1c843

Observation 70accda3-7105-445a-8a53-0fa773379f6d · outbound

This paper cites A Unified Perspective on Multi-Domain and Multi-Task Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Unified Perspective on Multi-Domain and Multi-Task Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.253956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.253956Z digest=sha256:857816637320a21fc92c21763087b92bf24af0103e68e58046c095931c185f13

Observation eb873359-d146-41a6-a21c-3727d824a5d8 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.384371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.384371Z digest=sha256:5ef0d55c019962c4056f8fded07358102685ce2c563d5b807758fa60860fcd8f

Observation a18e95ba-7dc5-42a9-b983-6530db2b0d41 · outbound

This paper cites Gradient Surgery for Multi-Task Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Gradient Surgery for Multi-Task Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.509096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.509096Z digest=sha256:e8fd7efbce59fa54f5564a0e9e65703e5f5f96516d05c1cc918d4ad2554f6d50

Observation 34000aaf-f730-4f6e-9659-43a4e2ce50f5 · outbound

This paper cites A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:55.333064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T13:33:54.674767Z digest=sha256:d63a855c107001e9d0c05018d87a2ab1fba7426e661e2566bff9dfbf9eebba91

Observation 141ad7d3-e01b-4495-b403-50e15e80cb69 · outbound

This paper cites An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.796497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.796497Z digest=sha256:0230da624d78b61ef2079579fcf1835928547646178090b9335fd9734a93a66f

Observation 5bb3652e-81b6-4dc2-9b03-43b5ae33a944 · outbound

This paper cites online" 'onlinestring :=.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.889501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.889501Z digest=sha256:73c2609e438b1ebfbb93847d7542b9639886c4a4aef6334dcd69a074afb14b18

Observation b1e40183-72ef-4dd2-9328-0352283a0f8c · outbound

This paper cites write newline.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.004931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:55.004931Z digest=sha256:c3c40883ef80115c27188ed88b0bec59ecabf622883b8a00eca651d2131a7420

Pith citing papers

No inbound Pith citation observations are available.