Pith. sign in

Paper Citation Record · LEDGER

Improving Pretraining Data Using Perplexity Correlations

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2409.05816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.05816 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:41:46.201649Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:36:44.914742Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b7b5ad2f-8d97-4ddb-b2d7-b0e890848129 · inbound

Predicting Emergent Capabilities by Finetuning cites this paper.

Predicting Emergent Capabilities by Finetuning Improving Pretraining Data Using Perplexity Correlations

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T13:41:46.201649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:41:46.201649Z digest=sha256:a6d3a25a0451e23ae5e27e6e043631d088935ee674eac379efbf4c46a6ed0c5d

Observation a9e89c38-b9d7-4b7f-bcb6-ddb850603f69 · inbound

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights cites this paper.

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights Improving Pretraining Data Using Perplexity Correlations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:02:00.873801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:02:00.873801Z digest=sha256:5cb97fb62140cc3a360054ba42bd61b964ff5f208c19b6604a183eb276ae04fa

Observation 3df25c12-fb7d-4560-8dc5-d25785fc236f · inbound

Optimizing Pretraining Data Mixtures with LLM-Estimated Utility cites this paper.

Optimizing Pretraining Data Mixtures with LLM-Estimated Utility Improving Pretraining Data Using Perplexity Correlations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:00:05.072473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:00:05.072473Z digest=sha256:4fa9084e00535aec99b9cd4d17a86f07945563ae6b4a332524a5e5d5866582a9

Observation f1f33d30-4c5d-4dd7-b682-eedca03e3d38 · inbound

Universal Model Routing for Efficient LLM Inference cites this paper.

Universal Model Routing for Efficient LLM Inference Improving Pretraining Data Using Perplexity Correlations

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-07T23:48:01.242684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:48:01.242684Z digest=sha256:13b49a28164dd0e006c562815cd7bb95cb9d68d4d00b67260f7a51aedcc50ed9

Observation 856fe6c8-6933-4bf3-b96a-3d9e5e54b44a · inbound

Energy-Based Transformers are Scalable Learners and Thinkers cites this paper.

Energy-Based Transformers are Scalable Learners and Thinkers Improving Pretraining Data Using Perplexity Correlations

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T20:42:36.952219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:42:36.952219Z digest=sha256:ffeada6d31a86747d2807fc54c6631737a8840deffbed0e588abb5ead1714f0b

Observation 4ab91ae9-c325-47e2-aefc-7ef900bc5b8b · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks Improving Pretraining Data Using Perplexity Correlations

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:14.396958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:14.396958Z digest=sha256:6c443a771733a595e6bc77b0649417ef4d194b9cc07735946e5e4f33752821a5

Observation 124a7159-eea0-4513-8ba0-610f38f028b9 · inbound

Next-Latent Prediction Transformers Learn Compact World Models cites this paper.

Next-Latent Prediction Transformers Learn Compact World Models Improving Pretraining Data Using Perplexity Correlations

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:25:29.476508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T07:21:04.684347Z digest=sha256:5c53944e910f6d6472536efcd60934e44fa3c0e729231fdd03fc78699f7b5e99

Observation 86b6826c-895e-4d93-b1ea-d4d75d436fa2 · inbound

Next-Latent Prediction Transformers Learn Compact World Models cites this paper.

Next-Latent Prediction Transformers Learn Compact World Models Improving Pretraining Data Using Perplexity Correlations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:57.036301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:27:57.036301Z digest=sha256:01d0ecbb2019bb51457396e50d6827431f86da093ce31121aad73166de236bdf

Observation dae64c4b-c549-4cb5-a979-17a62d8ef254 · inbound

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization cites this paper.

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization Improving Pretraining Data Using Perplexity Correlations

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:57.946769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:06:46.131725Z digest=sha256:bc8589a882590d2d99ebe1d569dc2fd6fde9ff39e457820b672f1f3f1862841f

Observation 65602620-2be1-4622-aae7-38cf84d45d5e · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research Improving Pretraining Data Using Perplexity Correlations

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:36:44.916180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:481d435f10cf77e149d4f08c42b35482dd7d078900f90d1d1fae5ceff2a6ea7c