Pith. sign in

Paper Citation Record · LEDGER

Predictive Data Selection: The Data That Predicts Is the Data That Teaches

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2503.00808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.00808 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:58.609898Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:47:05.647227Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da652bd2-770b-44a9-9884-0a3d163566cd · inbound

Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training cites this paper.

Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:58.609898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:58.609898Z digest=sha256:d5f61acdc1218854ea705974b7a3fa5119dd73a02e012902cc89fa673ea641bf

Observation 17216e23-cf4a-4cb2-8bc2-680b7e14c519 · inbound

BlueLM-2.5-3B Technical Report cites this paper.

BlueLM-2.5-3B Technical Report Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:55.930316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:20:55.930316Z digest=sha256:dd5c351ec019f983b3cf499361dedaf1fc02f1559ac03b9ad4fd977910510f80

Observation 7183ee2b-8b22-4e68-bad4-1f69c26df013 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:13.692690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:13.692690Z digest=sha256:b1bd57b3199610023ff615b2d0ab097466c3a8ab507116cbcadc7c889b1b2956

Observation 04727587-5435-47c6-9a54-2a8a9b07f2fd · inbound

Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation cites this paper.

Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:21:06.812208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:06.812208Z digest=sha256:0ab177b605cc423b5f3a943ae76e82d39933db51d4c69cbe3ffb35eb0c62b10d

Observation ce4c0162-ad61-47b9-8de7-a77b68c24dc5 · inbound

Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis cites this paper.

Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T16:57:07.095586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:57:07.095586Z digest=sha256:76c8e123d05593b4a78f99bfeea781e46659116fd2b5ed99610c78435ad357ed

Observation 2237c73b-6452-4317-a938-f70c0f8f7c9c · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:18.455452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:74c778819a2543a7a1aad316ba3a3742203e8fabf57570319a694c736c0fac06

Observation a8eaf7d2-9d56-4e53-99bc-0e1d8eff1275 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:54.537057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:f23900814717c52a1ba48394b0b09dceb6d3a05b6c5bfb207dd117a678f48de9

Observation ee7886de-bc9a-43f1-a59a-17a12ff0518d · inbound

CausalMix: Data Mixture as Causal Inference for Language Model Training cites this paper.

CausalMix: Data Mixture as Causal Inference for Language Model Training Predictive Data Selection: The Data That Predicts Is the Data That Teaches

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:47:05.648948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-02T15:45:02.415577Z digest=sha256:172e124bf404e5048f1dd599cdfbc41f4df2b45497b1a2beb27d08d23410a11e