Pith. sign in

Paper Citation Record · LEDGER

Sequence Parallelism: Long Sequence Training from System Perspective

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2105.13120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2105.13120 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:55.653397Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.806179Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b7f674bf-6ba0-432b-ad68-b49b25746a3f · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:46:10.164904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:5c0c23e4d5f35767aa80bd98ded87672104408f54d8b2a276a484483a1bb4b17

Observation 03e34ab8-7ec6-42be-a8fa-b3947023d55f · inbound

World Model on Million-Length Video And Language With Blockwise RingAttention cites this paper.

World Model on Million-Length Video And Language With Blockwise RingAttention Sequence Parallelism: Long Sequence Training from System Perspective

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:36:57.307966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T06:36:57.165551Z digest=sha256:b0e973e8e6788ffc403a4213abf898d007d2c52af75750b5805cc3b48c6753fa

Observation 664367ad-fd64-46f2-95fe-9a4d17909af8 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Sequence Parallelism: Long Sequence Training from System Perspective

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:47:27.874603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:8972b04de7943ed8572cfa59fc07824ce7e8be9d3bca4e3d5cb5549bd323aef9

Observation 1c57709c-5045-4694-aebb-90fad6e68c7b · inbound

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning cites this paper.

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning Sequence Parallelism: Long Sequence Training from System Perspective

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:55.653397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:55.653397Z digest=sha256:d91bd9f8f39a0cd88484c3c2eea3bcc24a57dd9d294b1b7d8107ce33a9b11af2

Observation 519d403c-33f9-416e-a771-6143d673a16c · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.986399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.986399Z digest=sha256:f76d06ab6b840f158611554917139939330399d230038c5bf89641656b1e1111

Observation 7b19a446-827a-4ff6-87aa-f8fe4cbc08b9 · inbound

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? cites this paper.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Sequence Parallelism: Long Sequence Training from System Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.288961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.288961Z digest=sha256:35d0489fd7646fd09c64afaa2876cfffab073b4e6aa42312f212a0833c0bb03a

Observation 8316e377-e9ca-4bba-99fc-99db6a301f91 · inbound

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs cites this paper.

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:36.333821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:36.333821Z digest=sha256:8d51ba48ad2f35e62440646abf1984dc32a2958b7dcbdf5e73b4d01527f8dea1

Observation aa75cc33-de9a-4262-beba-b1b0be105839 · inbound

ContentV: Efficient Training of Video Generation Models with Limited Compute cites this paper.

ContentV: Efficient Training of Video Generation Models with Limited Compute Sequence Parallelism: Long Sequence Training from System Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:32.085563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:32.085563Z digest=sha256:f688062b5a6ff4e9908cd6497d8bc49204c9721f5ef3e62e21dc2fb0d28f7590

Observation 4397b7e9-5a51-4462-97ae-4f0f935b4ab0 · inbound

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library cites this paper.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Sequence Parallelism: Long Sequence Training from System Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.922834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.922834Z digest=sha256:4e0dd02c6d4b388f1171809a1f2bddba2fd95046392b470d43cac9cfa9d7f725

Observation 9786b61f-a944-441d-8011-a71f08b31f5f · inbound

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences cites this paper.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.105856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.105856Z digest=sha256:2261b646f992e399f1f5ba1aaa1e94d7f894b11d37876f09af8f47aaf0352a60

Observation b5dd0957-cac0-43ba-8316-4e8de9778731 · inbound

Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs cites this paper.

Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:08.905826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:08.905826Z digest=sha256:bbacb3c2a8a2e3bb3a02c86780ec64014dc66dc3ba862ec06acb36658484c8bc

Observation dc945350-74c7-40d7-a7fe-22216310527c · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:11.234042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:11.234042Z digest=sha256:f04a725fdd476f9dff4675d7dd66cf962731011b6d2460f95543f3815062ddb0

Observation 5b76fa2b-c631-4d1f-91e3-164d2e7593a7 · inbound

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference cites this paper.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.794986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.794986Z digest=sha256:db67aa8857deef3087cc95f89e8a1a45199c4820d9d212a0735aa7f47b854a55

Observation 547516d7-5037-43eb-a675-17713898ed44 · inbound

TetriServe: Efficiently Serving Mixed DiT Workloads cites this paper.

TetriServe: Efficiently Serving Mixed DiT Workloads Sequence Parallelism: Long Sequence Training from System Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:10.837066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:10.837066Z digest=sha256:c63842acb24806497fff9590a162d6964a781c0226a21059c719938fb6ee3699

Observation 0a064b01-f126-40e6-bf92-a8c2d30f1b24 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Sequence Parallelism: Long Sequence Training from System Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.032377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.032377Z digest=sha256:ddb69d7bf30757699879e98c4a64e73c99389da0f542b43acb5114db127b215a

Observation 02222285-4611-4fc1-8a99-ed855fe0cc04 · inbound

Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems cites this paper.

Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:14.135198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:41:03.457680Z digest=sha256:c71171b918a3154f8a1e0b24dd0208d9a8d7ac65e2ea16df3cd331fcc7f66c5a

Observation a43b33c6-db05-495b-975a-439585037437 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.770043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:55:31.298030Z digest=sha256:fd0d05b60d3bc261e03f565f0a15a6f90ec40420afe405d7f77140e91a622bb5

Observation 442c23a1-c1ce-4b8c-b020-79353d910f77 · inbound

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing cites this paper.

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing Sequence Parallelism: Long Sequence Training from System Perspective

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:53:24.750202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:53:11.761843Z digest=sha256:6933d56407d919b23039229cbd187ecac78a31057d99d5c4a6c970edcede0b06

Observation 9525b37c-6ba3-41ad-81f2-6b08c1687abc · inbound

Online Dynamic Batching with Formal Guarantees for LLM Training cites this paper.

Online Dynamic Batching with Formal Guarantees for LLM Training Sequence Parallelism: Long Sequence Training from System Perspective

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:29:35.935578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T15:57:18.544274Z digest=sha256:0e6cb444bcc41ca8a885c10ee29020e0e51bbaa77a39701cdeb04c4423e14890

Observation a02cb5f7-b5af-431a-a4ad-2b877e49476e · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.807624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:9c403f596a47661a65978c89dd0548ac80942e59794c04196e3e4e98df304828

Observation 69a5c0ea-6441-4a69-8384-16c246a3f958 · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:28:43.956670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T17:26:07.870260Z digest=sha256:b94ef50224b8c5446b59c1ac9d30a9304683ce178608ad2aaab192977e831606

Observation a648ba3b-41c3-4a22-b4a4-38b97f72cd32 · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T08:40:29.554554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:40:29.554554Z digest=sha256:6decb35a35717f67ac8ed7722e564b55caacce4583aa342e1ee4502ff02bb140

Observation f8b4783f-dd2b-4cb4-8f77-fd28935a3bad · inbound

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention cites this paper.

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention Sequence Parallelism: Long Sequence Training from System Perspective

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:27:41.863772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T06:22:23.985308Z digest=sha256:8d7ce5dad67740a6fff6867c8992fba4b551ba701c73327ca69b625ee1daea27

Observation 6712b88c-b529-44d7-87b3-51e4c5d5f696 · inbound

Design-CP: Context Parallelism for Design of Protein Nanoparticles cites this paper.

Design-CP: Context Parallelism for Design of Protein Nanoparticles Sequence Parallelism: Long Sequence Training from System Perspective

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:52.764344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:29:52.764344Z digest=sha256:92de6d40a42da3d4b03c1c7e0b4ef9bd5f0c1d3ea692ab80da31d5cbdea8ca02

Observation ad4c00d3-51d0-48fa-ab5b-91bdf8133d7a · inbound

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving cites this paper.

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving Sequence Parallelism: Long Sequence Training from System Perspective

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:14.069087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:54:14.069087Z digest=sha256:ff615f7b4391099a97fc0434d6ffff0d4341e01d0776e9c7601d9b45c69133a9