Pith. sign in

Paper Citation Record · LEDGER

DataDecide: How to Predict Best Pretraining Data with Small Experiments

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2504.11393.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.11393 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:25:24.256379Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:17:40.210239Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 09c9c98c-958e-41bd-a2f9-16422b9b76e7 · inbound

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models cites this paper.

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:30.902227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:19:30.902227Z digest=sha256:c12f66d9b88dbd837c29822c562f37beb3662d4a86b4e5f7588ba21158948cfe

Observation 7b80c2bb-db38-4371-8bc0-d7b1e7857db7 · inbound

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions cites this paper.

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:24.256379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:25:24.256379Z digest=sha256:9b57435334a6af259f5dd3d1bd3fa9adc6244699b6cecfff49aaad385ae8050a

Observation ae2ccad7-69e2-4d31-8b2c-bf19e6d83711 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:11.514648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:11.514648Z digest=sha256:7a0c03ead13f8691549dbcfcf2a6ec7458df196ed5f11b7b43a6321b6b10c42f

Observation cf1d38eb-a514-43f7-9cc3-7fb4574ecc31 · inbound

A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't) cites this paper.

A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't) DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T23:11:03.148848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:11:03.148848Z digest=sha256:0d87dcbe659ebd952871016c18aca7a62f17ae61ec50a257066cd13e60baac37

Observation 7b9de1d4-0dee-4613-90d8-98f86f0b72fa · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:36:44.884445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:cc365581b419b49e394cec889c7c97916f1211b58372e12407e95f2b05361ac8

Observation 7929d72c-ea1a-4822-b48e-47c15d0e4895 · inbound

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation cites this paper.

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.627096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T22:48:20.193699Z digest=sha256:f3685d712e69fbe95e05b9813f8b069fb9b5d3c3a6ec4efecbd6e87d8b39b6bb

Observation baf43501-86a2-4479-b116-8abcaa119d67 · inbound

Small Experiments, Cheaper Decisions: A Case Study in Staged Promotion for Micro-Pretraining cites this paper.

Small Experiments, Cheaper Decisions: A Case Study in Staged Promotion for Micro-Pretraining DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:17:40.212442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T13:20:43.753744Z digest=sha256:fd9599c1ad2262716670ce52904c234d5092f14430ef35d72daf275db80e9182

Observation fc358018-e038-4a34-b346-0d2b169b2c71 · inbound

Domain-Aware Scaling Laws Uncover Data Synergy cites this paper.

Domain-Aware Scaling Laws Uncover Data Synergy DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T07:24:27.255815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:24:27.255815Z digest=sha256:2e6b18c899c46b62b0583d5d3ff7d22a76d4e68b2df46208f3d24fda7dd6ed1e

Observation 1b2d7a09-31c8-4455-a1fa-1df0d0750217 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining DataDecide: How to Predict Best Pretraining Data with Small Experiments

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:04.586864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:04.586864Z digest=sha256:496e50f978b4c80478f88cf9776a2fc1902c30d92cfc0d3984f5afe111aac628