Pith. sign in

Paper Citation Record · LEDGER

Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2403.16952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.16952 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:05:26.614578Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:19:57.811085Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 55cf0bf7-c786-460d-8a03-aada044e2867 · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:00:53.498646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T18:00:53.389420Z digest=sha256:647bc55a8f8356aee13d846ad519c391ae2d01d0a7a7b8588d0e56c2e79cf1cb

Observation f3992b5e-320c-44e4-97a6-1b4e7a591ed3 · inbound

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training cites this paper.

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.608695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T21:12:22.201810Z digest=sha256:6f79c3b5c1e0866f293f21433bfa91b2dbbd019f677d705d531e151d9e751204

Observation a2699ca4-c700-4482-b3f7-f4859ccf8536 · inbound

Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training cites this paper.

Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:37:00.829310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:36:50.366757Z digest=sha256:c845f31776165f641e049fc0f59ccea0b67786448ed28dcd70cbb78a26c09161

Observation 486b6b78-bd24-4905-841b-b6487ca3b1de · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:26.614578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:26.614578Z digest=sha256:0839a9aff9cf284119a2a77a3e15e9f1c9b7b754df634879a16c768bcedd77f6

Observation 89741ca0-f45b-4bd2-b4e4-1bafc3b8a197 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 165

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:06.999694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:ed23413e0a2eda2c4abff49e9a001850411290cb6897ad1efaf79d0bc0044b18

Observation fef432e1-3217-462a-b59d-b0ca16be6637 · inbound

Knowledge Transfer Scaling Laws for 3D Medical Imaging cites this paper.

Knowledge Transfer Scaling Laws for 3D Medical Imaging Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:45:51.622263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:30:50.324360Z digest=sha256:8ae227f3aa24a087fd922dca72e35e3decc25820c3bb965b2f98a7b44aaadde1

Observation 0fb22db8-e891-4c72-bfaf-c83f20e2eb62 · inbound

On the Invariance and Generality of Neural Scaling Laws cites this paper.

On the Invariance and Generality of Neural Scaling Laws Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:56.035239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:34:14.087140Z digest=sha256:2b838bb136030cc3483d39349e2c6841d9563602ae0ca77b45983e40a73e5aff

Observation 854f6c11-7a77-439d-b3c0-053bd22b24d3 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:48:00.954431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:1dc814151fa80f43dd74ad3bab543e0c62c6636647f24d59e7bc6f8837341f92

Observation a69322f0-547a-4290-83bf-0595bd573e66 · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.486738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:6620e089365c259361be4d5a226bcb65706f9a70d5649dc87cb06ac7e4ab5510

Observation 1e6235db-4944-4a0b-bcb3-0c2bd92e20cb · inbound

D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training cites this paper.

D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.231130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:41:12.453128Z digest=sha256:ab32eadb53abd27a4c2cd97d01a79439f36d29019aa736e698a7297cf6275d2e

Observation 8f41a52d-1312-4e67-9349-18a28f33a787 · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:36:44.897985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:c032652dc8ead0d244fa9cdd00d6fae513136919988eb211ffdff88bc6af4235

Observation 60a76257-86aa-4335-9722-75b17b21b171 · inbound

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning cites this paper.

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:19:57.812566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T00:46:13.587037Z digest=sha256:4c8ab405d1fbb0401d35fff7ad10a1ed6004476367e8c20864af9b6b780a8aac

Observation 7cefba8d-d3b0-4c9e-b540-1fdc4c684421 · inbound

Data and Evaluation Closed-Loop for Model Capability Enhancement cites this paper.

Data and Evaluation Closed-Loop for Model Capability Enhancement Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:25:48.051684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T01:33:08.134048Z digest=sha256:d45e66778aa557eb5910c0f4c6d1d77fa9bdaf4c83981dc529cf38be0915cc15

Observation a752bcf8-4bee-40a3-8504-d02272321cae · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:48:39.450809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:c713942f3d45ff69883593e3c2f34d84acd7575b30d9568c08219c1ff1515226

Observation 808aa92c-8e34-43b1-9180-58d2f58b0605 · inbound

Co-Adaptive Multi-Task LoRA: Transfer-Aware, Label-Free Control of Domain Participation cites this paper.

Co-Adaptive Multi-Task LoRA: Transfer-Aware, Label-Free Control of Domain Participation Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T01:54:44.256835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:54:44.256835Z digest=sha256:a257181d68cee60cfdbff7ccc1ae01cfb9bd7a796515fd8b8351b07bed5cd14f

Observation 72a3f635-59e7-4db1-9179-1a7d6df13ef4 · inbound

Domain-Aware Scaling Laws Uncover Data Synergy cites this paper.

Domain-Aware Scaling Laws Uncover Data Synergy Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T07:24:27.255815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:24:27.255815Z digest=sha256:5e4cfbd906862f1521ebae6c95c4d0235a3de71764ef1ec1d3c0206005099c29

Observation 22de8f3c-08d4-4808-a226-b81939f715aa · inbound

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation cites this paper.

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:24:46.118299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:24:46.118299Z digest=sha256:ba6dcd064a10d3c06a98c5f995c60cf98813c30051ee6183cc2f7687f0593787

Observation ddca782f-f0ad-4a5e-b0fb-b577d2f1ceab · inbound

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models cites this paper.

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T10:34:41.807413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:34:41.807413Z digest=sha256:05b8a74fa17de5e59c64e27f66aca10f97a4a04a6088de21624e5617958d0219