Pith. sign in

Paper Citation Record · LEDGER

Memory Analysis on the Training Course of DeepSeek Models

As of 13 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2502.07846.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07846 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:59:38.706310Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09d9dc3f-e77c-4f6f-b7a9-556239caeadf · outbound

This paper cites DeepSeek-V3 Technical Report.

Memory Analysis on the Training Course of DeepSeek Models DeepSeek-V3 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:59:38.681544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:59:38.681544Z digest=sha256:d708c30286bac40ebe3dd7996007328a5dd31b58153a53c9ba6cad3beaf80f17

Observation 4c0eb8d3-4657-462d-9d52-f98af83a59b4 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Memory Analysis on the Training Course of DeepSeek Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:59:38.687594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:59:38.687594Z digest=sha256:4c5bfea706230299a59ec70579c2a18889ca689b775e2d1b0a52bb0d087d52fe

Observation 9263eebc-ae51-4a8e-8591-911435a72915 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

Memory Analysis on the Training Course of DeepSeek Models Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:59:38.692695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:59:38.692695Z digest=sha256:813973f4ce098e6611451bb972d71d0ad04314510009cf4c899488b366da32b8

Observation 9e92a002-3c9e-427d-b385-5d97e53a27a8 · outbound

This paper cites Deepspeed: System opti- mizations enable training deep learning models with over 100 billion parameters.

Memory Analysis on the Training Course of DeepSeek Models Deepspeed: System opti- mizations enable training deep learning models with over 100 billion parameters

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:59:38.798318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T12:59:38.697008Z digest=sha256:ce10c5dcf01e0efa5f77bd0ef9376a1be661d8262d2f8944695156c215908c9b

Observation f5b9f1dc-2205-403c-bb22-f4a19ed493d5 · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models.

Memory Analysis on the Training Course of DeepSeek Models Zero: Memory optimiza- tions toward training trillion parameter models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:59:38.701541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:59:38.701541Z digest=sha256:16e547b2b6344e58873cd71b9b965ae3743b4719df80423b3bb6859ed6994140

Observation 6f548adc-7125-4d6f-8b0b-df466421eee7 · outbound

This paper cites Reducing activation recomputation in large transformer models.

Memory Analysis on the Training Course of DeepSeek Models Reducing activation recomputation in large transformer models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:59:38.773620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-08T12:59:38.706310Z digest=sha256:6114203797af7e7688e1af1fc1bccd7b0b234b5ce34ff9a6ceabc506a23597c1

Pith citing papers

No inbound Pith citation observations are available.