Pith. sign in

Paper Citation Record · LEDGER

M6: A Chinese Multimodal Pretrainer

As of 22 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2103.00823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.00823 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-24T12:10:49.690618Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T12:14:26.627052Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3f8851c0-dbbc-416f-9e39-fb38c754b508 · inbound

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model cites this paper.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model M6: A Chinese Multimodal Pretrainer

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.631436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:9830201dec1223f2e2f1033d5c4d521e26999a82cb189e4070f330045500ce1f

Observation 150afee1-1804-4672-9037-75b7b00aa375 · inbound

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models cites this paper.

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models M6: A Chinese Multimodal Pretrainer

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:50:12.821851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-11T22:50:06.399707Z digest=sha256:6a5ffba33d7ebed83544ba4b3eebb34188f282b412e60c5a1c733e486d42cb9f

Observation c5aad8ab-b82c-4816-89e0-6987c51d71d6 · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs M6: A Chinese Multimodal Pretrainer

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:16.889390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-09T19:46:13.015064Z digest=sha256:a3f5a0af5cc01cd154b5ec952d7c99b3b18aae634b9865d8d06a2397cf2f7ed4

Observation 3478a0dc-58d4-4ae1-bcaf-5dc5b91b53f2 · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs M6: A Chinese Multimodal Pretrainer

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:21:30.508049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-12T05:17:09.793360Z digest=sha256:365ca3497d4cea5d4dd301c170b8193dcb5af7e7a3903b13147719fbf5a2a6f9