Pith. sign in

Paper Citation Record · LEDGER

Training Trajectories of Language Models Across Scales

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2212.09803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.09803 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:41.259204Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eca6dce9-0940-423c-bb4c-a39f29078657 · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Training Trajectories of Language Models Across Scales

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:45:17.903448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:a3730baf6bbf0cd6a2db15d078beb53fd8c76a99e09bc9b1cb3e1326ffb51dea

Observation 50b8710f-c728-4e18-9c14-b2f87676ecc5 · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Training Trajectories of Language Models Across Scales

Reference 180

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:45:17.674417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:0b9dfc2b6a8631db0afb87f1e724cf3620738bf7200a1cdc622db7612b12f50c

Observation 7e2a865d-2e55-4b95-a984-98098a30b8f6 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Training Trajectories of Language Models Across Scales

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:35:21.479135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:51e54af6a06907ff140799d64012a8dcda947f8547866a5498cc3e4dc0c5f037

Observation 9bd4b689-2dfa-4773-81db-8605e8572abc · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Training Trajectories of Language Models Across Scales

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.501537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:417388eaef8a71df5812d037c5b40e9974bb58d9c4203085abc0eb75f2642c94

Observation 50842ce7-c236-41b2-a326-fb7d0b61af28 · inbound

How much do language models memorize? cites this paper.

How much do language models memorize? Training Trajectories of Language Models Across Scales

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:41.259204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:41.259204Z digest=sha256:ff214f0b0a25da5a0f77da42a5ce4db10b8a2fb9ed4fdf2f434b8320eda22fe7

Observation 472b5137-f1ae-45ce-90c0-e4df3b90f4dd · inbound

Fairness Dynamics During Training cites this paper.

Fairness Dynamics During Training Training Trajectories of Language Models Across Scales

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.924492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.924492Z digest=sha256:b9a6808b63b9cea9c0698701010c30cbfef92e7ad055b4a325065456ae0c3162

Observation 87b1c8c5-13f9-4c96-aac4-392dacadbba9 · inbound

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law cites this paper.

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law Training Trajectories of Language Models Across Scales

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:02.700113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:02.700113Z digest=sha256:80fc63ffb41d618ab92fe71e47758193e7780f27480d42ba2b18c577a43ca6ba

Observation e71ca646-0868-4ca5-873d-b501270a77b0 · inbound

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent cites this paper.

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent Training Trajectories of Language Models Across Scales

Reference 226

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:32:55.873440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T01:29:14.555216Z digest=sha256:2f712e587deee8b6b679abacf0a1ab7e85ea4158dadcdfc359e8d69b895680e0

Observation 0a19b94e-d449-4f28-9d41-7f6768ef0bdb · inbound

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent cites this paper.

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent Training Trajectories of Language Models Across Scales

Reference 226

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:40:24.922360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-25T06:39:16.246591Z digest=sha256:8a1b05ed151e78eec1c082fd5af2b9caba924e177076beadb231eded2754d5b7