Pith. sign in

Paper Citation Record · LEDGER

Reducing Activation Recomputation in Large Transformer Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2205.05198.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.05198 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:30:56.938309Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

55
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 23efadb5-39a3-4564-96fa-05b7a3431662 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Reducing Activation Recomputation in Large Transformer Models

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:19:46.710952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:c5a58ecea4079b76c2315d94d56fe53e351e6a247cefd481080551a3409bd564

Observation 5a993fea-6aa6-42ef-ab50-9fbbc7c9c94a · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Reducing Activation Recomputation in Large Transformer Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.812840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:234566ca7bedc82d76731cd5faca92db5ca9ea6af8fbab579d7bd5092901b0a2

Observation e3169a5d-bf93-40d0-b0f4-2af1cd38ed7b · inbound

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel cites this paper.

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel Reducing Activation Recomputation in Large Transformer Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:15:20.079382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:15:20.027659Z digest=sha256:785676f6bc1e5e430ef47b79d79a008f57a9fa17acac99f5fb1c1e6a7cce9b2f

Observation cba71a44-9ef7-4c93-85f8-b271d07509da · inbound

Ring Attention with Blockwise Transformers for Near-Infinite Context cites this paper.

Ring Attention with Blockwise Transformers for Near-Infinite Context Reducing Activation Recomputation in Large Transformer Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T19:28:28.258495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T19:28:28.201789Z digest=sha256:dceef1c73191714d0a394130394851ccc2b6113ac549a2e9c43db640daa970cc

Observation 691d0b08-e397-4417-a090-0b2b67fe70d9 · inbound

An Empirical Study of Mamba-based Language Models cites this paper.

An Empirical Study of Mamba-based Language Models Reducing Activation Recomputation in Large Transformer Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:31:03.874186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T10:31:03.777169Z digest=sha256:b711f996f59400d6ef8aefb2a615d8279a7532c286fdf6d14edf9365c3caaaf9

Observation 6c13164e-123b-4bca-9aec-6c6c21527958 · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Reducing Activation Recomputation in Large Transformer Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:07:14.441789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:150e529cca7e2bb6e3a4a310cb7fcd91e207d743efdb941b3d02798d2458643e

Observation c7c516d8-0213-4b7b-924e-62402ec17b46 · inbound

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training cites this paper.

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training Reducing Activation Recomputation in Large Transformer Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:15:09.565107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:12:22.201810Z digest=sha256:0d5d96fb9f9244b2d29b33bb9e074ca5a72771fd2fc3206e41def7249a468e8e

Observation 6cf91668-8f56-47a2-9b89-8c4066a1a1d8 · inbound

Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project cites this paper.

Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project Reducing Activation Recomputation in Large Transformer Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:05:09.524204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:04:11.859563Z digest=sha256:cd81739279c674dda9b1cb89fac1ece6bc5fb76fe9ea45503cb46286519e7ece

Observation 05cf9cdb-0e18-47b3-8b67-145d5b4307f4 · inbound

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences cites this paper.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Reducing Activation Recomputation in Large Transformer Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.938309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.938309Z digest=sha256:b3b57359898b893ccfe821dc650a1bd56d9e754ef38e6b67738d643c6d689b7d

Observation 56cb0289-22aa-40af-8685-0e92e1d25712 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report Reducing Activation Recomputation in Large Transformer Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.846985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.846985Z digest=sha256:766f238b31a3ef69091fda6b966f5340df4bd81b2696fe2ef926acf300c2f17b

Observation ab079cb3-4ffd-4be9-96f4-13dc456ad6b2 · inbound

Photonic Fabric Platform for AI Accelerators cites this paper.

Photonic Fabric Platform for AI Accelerators Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:20:28.750375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:20:28.750375Z digest=sha256:ea3db7a8436c0cd7cd49e0c67cc4ee37eadbb544cd51fcbad4e9c2ef990f20aa

Observation 29feca7a-4ba7-4d1c-870a-f553e4fdb36f · inbound

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling cites this paper.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Reducing Activation Recomputation in Large Transformer Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.758677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.758677Z digest=sha256:0bf8331d9af97d934e8fe85986f4c2aa45348f64b60006a1f4e4f884747a73ab

Observation 2310f253-5f7a-4c4f-b35c-c2b3c75cacbb · inbound

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection cites this paper.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Reducing Activation Recomputation in Large Transformer Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:41:50.661629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:715feb1b3cac70776a88d27d34b84328ad2f740c7a2cedfa504b0f337f89ac30

Observation 1599cf5d-4ed1-448c-9376-f64d6d84fb65 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:51:45.802088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:60e05d063286a912969246b10a5de2c0c3a6a0209dc37508295030a2a5ef0333

Observation e994574e-f6b7-406f-9d3a-6140912e951d · inbound

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training cites this paper.

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:06:27.277349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:04:31.017142Z digest=sha256:e889db2a9b1f6fb5be14cf7ef78655d84a4ffd57dea2df9980e19675b6c3cc2e

Observation e0fd7fba-7f23-4a1d-9063-d8b7f7787471 · inbound

Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs cites this paper.

Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs Reducing Activation Recomputation in Large Transformer Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T22:28:46.339471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:28:46.339471Z digest=sha256:217dee55b1bbc50014f6462b0a06b66aa162fe99482e7606ea7cc6254450e4b2

Observation 12782da5-748a-4f8c-b556-b6116eb53347 · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence Reducing Activation Recomputation in Large Transformer Models

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:40:42.615219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:7424e917e1ea651be8e490e071cfd124345492657b17069d9e4fc873e312129a

Observation 25597e34-7a25-4688-a37f-d0c9ac86d4fa · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Reducing Activation Recomputation in Large Transformer Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.812534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.812534Z digest=sha256:1b14ba43fb2091817831298f722060b8a97cce6f894fa97d969f3f3615aaab09

Observation 44a5279e-e3ff-4c8f-aa1e-c570a6561d63 · inbound

Efficient Scaling of LLM Training with Flexible Context Parallelism cites this paper.

Efficient Scaling of LLM Training with Flexible Context Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T20:59:24.724128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:59:24.724128Z digest=sha256:583b09c8eaec1375022b742c7fd5d9dc9519683c9e5a564c3e6cf22f3689a8bf

Observation b2543f4e-0d6d-47cd-9f31-2669ecf6a87b · inbound

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism cites this paper.

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:24:20.560853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:23:19.236803Z digest=sha256:f800360e0d5f673f4a88dc6cdf4a7b4b9c99b477ed04bb60c93af085f7f57067

Observation 570aaedd-d074-469c-ab98-e26078be9db5 · inbound

Decoupled DiLoCo for Resilient Distributed Pre-training cites this paper.

Decoupled DiLoCo for Resilient Distributed Pre-training Reducing Activation Recomputation in Large Transformer Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:49:15.683024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T22:20:21.090246Z digest=sha256:3e85122133391f1881e5f8dcf732a1a79d7dae5a836e48d32fe096505d042da1

Observation 01b389ca-2465-4751-aeb1-731fb8ab9bfa · inbound

Efficient Training on Multiple Consumer GPUs with RoundPipe cites this paper.

Efficient Training on Multiple Consumer GPUs with RoundPipe Reducing Activation Recomputation in Large Transformer Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.808957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T10:37:22.251566Z digest=sha256:9bd595924d92907ce8659f0e4e51eba38d1b7912b934855070925c05c63a8b51

Observation 626393f0-cf06-4885-8f1f-e18a1ba5e337 · inbound

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism cites this paper.

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 40

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T17:51:07.974000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:04:02.418499Z digest=sha256:f0fad94bd360aa2ce37f429355d536de85ab2448def99063570ed2480a6201b1

Observation 8b587fa3-48b1-4433-8d58-456f31566783 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Reducing Activation Recomputation in Large Transformer Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:06:15.053308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:8c8fd7c0f7500de56da2f796753846875600dad5efc9fe44865902234f2bc721

Observation 99be7701-2fdb-4847-b917-fbca9bf0467d · inbound

Instant GPU Efficiency Visibility at Fleet Scale cites this paper.

Instant GPU Efficiency Visibility at Fleet Scale Reducing Activation Recomputation in Large Transformer Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:43:54.741443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T02:43:09.077121Z digest=sha256:43372d626d49fa673db145af75eddc07e567603bfa6fc6a2f70f50ec8ba1d190

Observation f8d65840-193a-47fe-8a6c-495610f149ab · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.529353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T20:20:53.599064Z digest=sha256:adf90869317bb978ed54e1027680482a23b723537a68d53f22d80f9ffe1a363e

Observation 0c493078-284b-448d-af6a-03d51e9a3e57 · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T10:54:16.434702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:54:16.434702Z digest=sha256:e8fe882c9ad404c22defcfa5121c951523592cec8f041febc4192467a29dcb11

Observation 8d91bf9b-23e5-4d5a-98c0-f507144fa1b9 · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T12:13:46.609563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:13:46.609563Z digest=sha256:b87887405160e29eba7c4383641403824e53571511c52dfff2854dcc65f8a945

Observation faa07337-c6ca-41dc-a16f-b1657d9e5bd8 · inbound

The Cost and Network Limits of Space-Based AI Compute cites this paper.

The Cost and Network Limits of Space-Based AI Compute Reducing Activation Recomputation in Large Transformer Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:46.028286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:46.028286Z digest=sha256:ca9c9a4e066181bfddbdbe677d40bed5f60944e4b38e54ae42379a3caaa4883b

Observation a31fe6a3-500d-467f-9f96-095c7d2e8fba · inbound

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget cites this paper.

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:45.462772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:45.462772Z digest=sha256:ccea9197842e7018cb1eaa1c3fc5919219f51b03e949323339befed06d4afec4

Observation 58bacf8a-3c44-477c-ad9a-a7ebad505fcb · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Reducing Activation Recomputation in Large Transformer Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.263658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.263658Z digest=sha256:1846cb74dc8cf7357d49885b6c92d0031e6fc05ca3216388f5c91e03803fedfd

Observation 2b777e7d-da11-4b27-8c90-c084795671c0 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Reducing Activation Recomputation in Large Transformer Models

Reference 194

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:44.020572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:44.020572Z digest=sha256:944ff6f9dd142b676969191c95f773d6da4e1ae74a8baef5999eeaa8de169648