Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2205.05198.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:30:56.938309Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
55
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 23efadb5-39a3-4564-96fa-05b7a3431662 · inbound
BloombergGPT: A Large Language Model for Finance Reducing Activation Recomputation in Large Transformer Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a993fea-6aa6-42ef-ab50-9fbbc7c9c94a · inbound
A Survey of Large Language Models Reducing Activation Recomputation in Large Transformer Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3169a5d-bf93-40d0-b0f4-2af1cd38ed7b · inbound
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel Reducing Activation Recomputation in Large Transformer Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cba71a44-9ef7-4c93-85f8-b271d07509da · inbound
Ring Attention with Blockwise Transformers for Near-Infinite Context Reducing Activation Recomputation in Large Transformer Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 691d0b08-e397-4417-a090-0b2b67fe70d9 · inbound
An Empirical Study of Mamba-based Language Models Reducing Activation Recomputation in Large Transformer Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c13164e-123b-4bca-9aec-6c6c21527958 · inbound
Wan: Open and Advanced Large-Scale Video Generative Models Reducing Activation Recomputation in Large Transformer Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7c516d8-0213-4b7b-924e-62402ec17b46 · inbound
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training Reducing Activation Recomputation in Large Transformer Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cf91668-8f56-47a2-9b89-8c4066a1a1d8 · inbound
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project Reducing Activation Recomputation in Large Transformer Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05cf9cdb-0e18-47b3-8b67-145d5b4307f4 · inbound
Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Reducing Activation Recomputation in Large Transformer Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56cb0289-22aa-40af-8685-0e92e1d25712 · inbound
RoboBrain 2.0 Technical Report Reducing Activation Recomputation in Large Transformer Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab079cb3-4ffd-4be9-96f4-13dc456ad6b2 · inbound
Photonic Fabric Platform for AI Accelerators Reducing Activation Recomputation in Large Transformer Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29feca7a-4ba7-4d1c-870a-f553e4fdb36f · inbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Reducing Activation Recomputation in Large Transformer Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2310f253-5f7a-4c4f-b35c-c2b3c75cacbb · inbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Reducing Activation Recomputation in Large Transformer Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1599cf5d-4ed1-448c-9376-f64d6d84fb65 · inbound
SpikingBrain: Spiking Brain-inspired Large Models Reducing Activation Recomputation in Large Transformer Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e994574e-f6b7-406f-9d3a-6140912e951d · inbound
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training Reducing Activation Recomputation in Large Transformer Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0fd7fba-7f23-4a1d-9063-d8b7f7787471 · inbound
Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs Reducing Activation Recomputation in Large Transformer Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12782da5-748a-4f8c-b556-b6116eb53347 · inbound
NVIDIA Nemotron 3: Efficient and Open Intelligence Reducing Activation Recomputation in Large Transformer Models
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25597e34-7a25-4688-a37f-d0c9ac86d4fa · inbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Reducing Activation Recomputation in Large Transformer Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a5279e-e3ff-4c8f-aa1e-c570a6561d63 · inbound
Efficient Scaling of LLM Training with Flexible Context Parallelism Reducing Activation Recomputation in Large Transformer Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2543f4e-0d6d-47cd-9f31-2669ecf6a87b · inbound
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism Reducing Activation Recomputation in Large Transformer Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 570aaedd-d074-469c-ab98-e26078be9db5 · inbound
Decoupled DiLoCo for Resilient Distributed Pre-training Reducing Activation Recomputation in Large Transformer Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01b389ca-2465-4751-aeb1-731fb8ab9bfa · inbound
Efficient Training on Multiple Consumer GPUs with RoundPipe Reducing Activation Recomputation in Large Transformer Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 626393f0-cf06-4885-8f1f-e18a1ba5e337 · inbound
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism Reducing Activation Recomputation in Large Transformer Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b587fa3-48b1-4433-8d58-456f31566783 · inbound
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Reducing Activation Recomputation in Large Transformer Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99be7701-2fdb-4847-b917-fbca9bf0467d · inbound
Instant GPU Efficiency Visibility at Fleet Scale Reducing Activation Recomputation in Large Transformer Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8d65840-193a-47fe-8a6c-495610f149ab · inbound
Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c493078-284b-448d-af6a-03d51e9a3e57 · inbound
Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d91bf9b-23e5-4d5a-98c0-f507144fa1b9 · inbound
Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa07337-c6ca-41dc-a16f-b1657d9e5bd8 · inbound
The Cost and Network Limits of Space-Based AI Compute Reducing Activation Recomputation in Large Transformer Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31fe6a3-500d-467f-9f96-095c7d2e8fba · inbound
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Reducing Activation Recomputation in Large Transformer Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58bacf8a-3c44-477c-ad9a-a7ebad505fcb · inbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Reducing Activation Recomputation in Large Transformer Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b777e7d-da11-4b27-8c90-c084795671c0 · inbound
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Reducing Activation Recomputation in Large Transformer Models
Reference 194
Source-reported events for the cited work
Unavailable: canonical work link unavailable.