Pith. sign in

Paper Citation Record · LEDGER

M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2110.03888.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.03888 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:16:09.245076Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:07:22.287771Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cdf3becd-21f6-4168-9b75-c5a3413776d3 · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Reference 165

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:22.289313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:f5281fc7a51b15d4b545872acdd09634ba75cbc879a42d6f99ff6d9fdbeb4f09

Observation 59ff5f7a-1151-49fb-9e77-13748d6ceb5d · inbound

An Integrated Framework for Contextual Personalized LLM-Based Food Recommendation cites this paper.

An Integrated Framework for Contextual Personalized LLM-Based Food Recommendation M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-16T10:16:09.245076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:16:09.245076Z digest=sha256:6512a5ebd79f8df0aede5256ca8528d4d431acb75186947a6eef4d5aa4c4a237

Observation 1c64c3ed-d511-460e-a6db-c20dfd1dabe7 · inbound

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration cites this paper.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.187026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.187026Z digest=sha256:8025fee5e096da301cf3a36dbec3b61a06576a1abc3253122dccaf81e47db9c9

Observation 945e0245-891f-4d70-8b7a-ab9036e49b01 · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:17.626282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T19:46:13.015064Z digest=sha256:b8fe0ba5cbb710b37e675dd3c721a586f2deb67ee4faf6e444a6be47c5f1e628

Observation d54c0e61-df97-4609-97d9-7f633cab5ec1 · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:21:30.086287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-12T05:17:09.793360Z digest=sha256:81e5c05a9a8ef2e5118ef239a61708b029df23c5de339d532a0cedc635b9c9a2