Pith. sign in

Paper Citation Record · LEDGER

Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2002.11794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2002.11794 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:58:13.116663Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:58:13.282745Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 285851e4-2cd8-4642-921c-64130be57f60 · inbound

Scaling Laws for Autoregressive Generative Modeling cites this paper.

Scaling Laws for Autoregressive Generative Modeling Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:49:43.839011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T07:49:43.711653Z digest=sha256:e65ec8f026a737299c8a55e9a3fea8c91595f604811b3f133abfc5e58bff5cd0

Observation 9cefc9e6-0440-4577-a2c9-b3ffe93a2901 · inbound

Scaling Laws for Transfer cites this paper.

Scaling Laws for Transfer Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:58:13.285137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:58:13.116663Z digest=sha256:aa25439b3a18baa8191bfe7f90c3315b314ca0a1ecf035033dfc0592ea1cf624

Observation f66a64cd-c2c3-45cd-b812-bc3bcf822ec4 · inbound

A General Language Assistant as a Laboratory for Alignment cites this paper.

A General Language Assistant as a Laboratory for Alignment Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:22:59.520797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T14:22:57.925354Z digest=sha256:a0b82b203cef87e5a4ee12bdde96e246988e8fb6a26b70cab2f6b3477b722eb0

Observation d36dce0d-ee55-4c0f-8fb9-33899f190b82 · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:42:47.765535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:f24417c433a41f777e4ec4d7293e7f20ca2609023ed958cae83e7b598e1b4962

Observation b448464b-74a8-40b0-8ca0-e15f4d3bdf7c · inbound

Multi-Aspect Knowledge Distillation for Language Model with Low-rank Factorization cites this paper.

Multi-Aspect Knowledge Distillation for Language Model with Low-rank Factorization Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:38:14.644305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:35:27.085495Z digest=sha256:c3be92ce47a47e64e884c2a97f7003904536039d27ff2eb9475080f5ba8c49af