Pith. sign in

Paper Citation Record · LEDGER

DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2305.10429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10429 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:49:32.945685Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.134017Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ddfd013d-5f9e-44ca-b6eb-4fb10d7b5e35 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.480274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:83d121ce1150ad2fd9768dee00a66b6f3adf84872161a2df4eec3f9427d192ef

Observation f908cae6-8cff-4bac-8b2c-effc1ca44a04 · inbound

Llemma: An Open Language Model For Mathematics cites this paper.

Llemma: An Open Language Model For Mathematics DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:17:46.419903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T08:17:46.055279Z digest=sha256:2b8d5356ef12005d4a30218259074d8a8aae74b6e6cc4dc8be397a3b8900ad6b

Observation aa434cf1-ebd4-4aef-9ead-0802752e0920 · inbound

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging cites this paper.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.557385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.557385Z digest=sha256:8fe3a7db8c9f520f98de1fd2d9dbc6927359ce95967ca9c09145357bbbefd8f3

Observation f0addaf2-1bbc-44bb-a1ac-08568136d1bd · inbound

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining cites this paper.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.566504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.566504Z digest=sha256:e36feb0c7b0df9e8db2f522aa14f4fc8bfbf157ab68813d1a6690c30a32ae2ed

Observation 3daf7f01-a17b-4f35-84cb-77d48c8e660b · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:34.937484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:34.937484Z digest=sha256:ff34e648234180d84c3214961218f836795fc858d7ba6edfeb11042530ed3c83

Observation 224d3fe8-7e53-490e-b635-2bc8bc3f5401 · inbound

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training cites this paper.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.945685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.945685Z digest=sha256:447a17281061a90cbd2321e13cda02ef241b85503b9484c036729bba379402e3

Observation 06eaf234-dbfa-4b6e-8877-ee3e34034ffb · inbound

ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects cites this paper.

ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:03.006603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:03.006603Z digest=sha256:f0bfbec3f0e84d6ffbcd2b2e7182bad52f8c627d271552c4ac33e8dcdf10e348

Observation b4303d09-6850-428d-9ed4-32fa963c15ef · inbound

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining cites this paper.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.550042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.550042Z digest=sha256:389383f8daf3a52255328e1c8f9cdd7bd37e88c2a90e0809275eee15ac0ab047

Observation 8d4d2737-a203-4e5d-a27c-6133fb09a45f · inbound

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives cites this paper.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.125833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.125833Z digest=sha256:d5242dbbcdebe514e00f40a00459e54f2348f3550d630462a1e894c42b433569

Observation dde6cae9-73d6-4065-836e-a0cc34f9af89 · inbound

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning cites this paper.

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:28.811231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:28.811231Z digest=sha256:1ba2b19a30822377081ee3ae09b31cabd2400b17c293d5a3f527713f1d0da05e

Observation 41496677-dd96-420c-a33b-7e258923b5e6 · inbound

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models cites this paper.

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:41:53.033134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T22:41:26.047957Z digest=sha256:decaed76a0233d2289d0f12ec6e76f11f3d67520031094dc2d0119d9d1b9a9d8

Observation d1cb1622-3927-4ef0-8546-d8c72aa83fd5 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.127336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.127336Z digest=sha256:a14cfa9a401f273d7a817d8b5a0c624e2097cebc3637388d7e5778ea7d2aaf0b

Observation 3fd118da-1829-4887-8de4-76839afb8443 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:26.534644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:26.534644Z digest=sha256:0e8625fbd06f78873e9ce7099c79423dec0ec5189c6b03b926c9fdd4d12dd23d

Observation dbceb4dc-a267-431b-bb0f-2776d01a6a1e · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:19:27.466265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:e4e505d7c056b31b554ea682d6ec406857340c3ac993a59f464329c2eae18f71

Observation 38133877-a792-4574-a582-bc635b36b31a · inbound

Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics cites this paper.

Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:37:57.184917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T09:52:38.167538Z digest=sha256:4c99384604186d9e86783f7745c6968af27058ccc9193a833bb64620876b9423

Observation a4e4bb3b-96a9-41fa-b64a-769f3af999c7 · inbound

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining cites this paper.

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:10:09.135670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-25T19:04:11.976747Z digest=sha256:77b8ce95d6a07b2f479fedf114d11c047ed9cb07092bde05a7f9af4251b1886d

Observation c4ea9473-5c1b-4fbe-83dd-040053d71e34 · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.445070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:fc3af743de5c4114e1ae0abfdf77b461d875d891fb25c348c9a131c62dd376ca

Observation 1fc5d40f-f130-4a86-9e41-b585abd7572d · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:20.625282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:20.625282Z digest=sha256:1eca6f03fcae1e9770c34a4668a3e2de0d6855731ada87ba7b184ace3a6ac147

Observation 66f4baed-efb2-4bff-8050-86e767bad1c8 · inbound

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement cites this paper.

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T01:15:07.227138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:15:07.227138Z digest=sha256:703d9031a6d06d3e47180b9754fd8cd6c9c173c6ff5037214bddd2b8d1fe9221

Observation f80a9f01-918f-46e5-9b9c-5cc441644f97 · inbound

Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization cites this paper.

Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:34:17.221331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:34:17.221331Z digest=sha256:8dfc0f7f560dfd533dd25d4e037858c3ab1606fb6aa6346e0e3a9f1ca2502c1b