Pith. sign in

Paper Citation Record · LEDGER

BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2508.10975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10975 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:33:58.134824Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T16:54:58.206077Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f0433471-6744-4cde-8e03-160a390f6dfa · inbound

Controllably Efficient Language Models cites this paper.

Controllably Efficient Language Models BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.134824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.134824Z digest=sha256:545f62c05b57b671e620c45bb076eec9d9c8aefad106aee925b003d96a0a59d0

Observation 62560294-e859-4955-9fd9-d16622566338 · inbound

Action-guided generation of 3D functionality segmentation data cites this paper.

Action-guided generation of 3D functionality segmentation data BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:51:31.948881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T04:49:06.269256Z digest=sha256:84a86aa26a8496b47d6cef425db20f8a2886df65e8d1c10f997d76f3f8cc75aa

Observation 4e028ef2-af96-4ef3-8b1b-581fe2349e5c · inbound

Linguistics and Human Brain: A Perspective of Computational Neuroscience cites this paper.

Linguistics and Human Brain: A Perspective of Computational Neuroscience BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 197

Resolution
unresolved
no resolver link, observed 2026-08-03T03:22:46.300889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:22:46.300889Z digest=sha256:e95aadf50c33aebd29751b7f37042a8c967aefe849c5dd88a6382bc51b151464

Observation 0f7db37f-fd7f-4ca6-ab9c-c886dea62ed6 · inbound

Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale cites this paper.

Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:48:15.130856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T11:44:58.613981Z digest=sha256:cb68031ac554c9a1a2337edb79a380d8ff9b5db6a6186739f1cd07e9e1035814

Observation b032f9a5-3ba8-490f-998e-2fa7cfccb265 · inbound

Understanding Data Temporality Impact on Large Language Models Pre-training cites this paper.

Understanding Data Temporality Impact on Large Language Models Pre-training BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 2

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T16:54:58.207894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T16:53:01.583415Z digest=sha256:cfdb6bd6f8f3ae852b590fdf9f333f5eb47757c6df026b77748b1b7ff22b78f8

Observation 4f668e0a-0926-489d-8397-14e58271fe9e · inbound

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention cites this paper.

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.203279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:39:14.327838Z digest=sha256:09d903bbee067cedc2b852da4b088ccbfd04a408a53c135876c63e4ec100cf02

Observation 04a89781-b13e-4781-bf90-88f658abbf78 · inbound

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data cites this paper.

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T07:01:45.805431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:01:45.805431Z digest=sha256:29c1d5d2a2b9f53a047152147226900e1a1030589971d189b94326e33cf205ac

Observation afc736d1-078d-4fbd-a7b6-8bed5d1cf7f6 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:07.696790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:07.696790Z digest=sha256:089db7eced942017586b66f1fd1457d0bc4d785813eb2906ae3dfe2bdf9132bc