Pith. sign in

Paper Citation Record · LEDGER

Sequence Parallelism: Long Sequence Training from System Perspective

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2105.13120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2105.13120 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.348204Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.806179Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b7f674bf-6ba0-432b-ad68-b49b25746a3f · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:46:10.164904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:5687d0d6c82897c0510044ac368010c77b2ea1a146cd176e3dd39a7bc94672fe

Observation 03e34ab8-7ec6-42be-a8fa-b3947023d55f · inbound

World Model on Million-Length Video And Language With Blockwise RingAttention cites this paper.

World Model on Million-Length Video And Language With Blockwise RingAttention Sequence Parallelism: Long Sequence Training from System Perspective

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:36:57.307966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T06:36:57.165551Z digest=sha256:c96b5f1b981878f4cc28bc8a139e20d7607f0a4a22e824756794182f363f49b4

Observation 664367ad-fd64-46f2-95fe-9a4d17909af8 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Sequence Parallelism: Long Sequence Training from System Perspective

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:47:27.874603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:9f62ffa455d269e71c195f5d9faf8afc8e49dbbc4d597d2d43d63c87c1d24fde

Observation 85eae24b-9be0-41e4-9ced-6facda989775 · inbound

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models cites this paper.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.348204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.348204Z digest=sha256:1dba35a20132110138e723c72125199e118b0468974969e57c619cee76513983

Observation 89950387-b391-4274-89af-6762633be046 · inbound

When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training cites this paper.

When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training Sequence Parallelism: Long Sequence Training from System Perspective

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:30:21.231024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:30:21.231024Z digest=sha256:9a2a05e033a900456eb0ceeba3b361a219029e78204c5f8a721daa2cc00da0b2

Observation 3a937a68-e871-4302-8c80-6e19c7abce6f · inbound

Movie Gen: SWOT Analysis of Meta's Generative AI Foundation Model for Transforming Media Generation, Advertising, and Entertainment Industries cites this paper.

Movie Gen: SWOT Analysis of Meta's Generative AI Foundation Model for Transforming Media Generation, Advertising, and Entertainment Industries Sequence Parallelism: Long Sequence Training from System Perspective

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:04:47.556883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:04:47.556883Z digest=sha256:f014b55d483b0b685487afb2aeacf8a1a7c8020d8d60d1f17d1fbca7d88fcf02

Observation 8458eaa1-9adc-4e5f-8467-ae62d3015b5d · inbound

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication cites this paper.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Sequence Parallelism: Long Sequence Training from System Perspective

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.920005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.920005Z digest=sha256:f55bdba165179f770a44208fe554ce51f47a4229ef7e7c57b9373eecbcbf6e74

Observation 253fd9db-d347-4c64-ad4f-2a0c238ffe64 · inbound

Automatically Planning Optimal Parallel Strategy for Large Language Models cites this paper.

Automatically Planning Optimal Parallel Strategy for Large Language Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:59:59.133758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:59:59.133758Z digest=sha256:3cdc0257e39047fac462a7f6e45751efbc8ec1c64c6671bc730a155674c2e983

Observation 9e654174-9035-441a-afb5-2289664958c6 · inbound

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation cites this paper.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Sequence Parallelism: Long Sequence Training from System Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.508011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.508011Z digest=sha256:9d68262352f24e7c7c0337932ee8bdbed5396c6106090da95ff272051343bdf9

Observation 2745746e-9eee-48bf-99c8-1b28bb474655 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.343801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.343801Z digest=sha256:fb30a24f268c82a753ea3525101083f62dac3bd47228c3ddacbbcc348689a956

Observation 1c57709c-5045-4694-aebb-90fad6e68c7b · inbound

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning cites this paper.

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning Sequence Parallelism: Long Sequence Training from System Perspective

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:55.653397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:55.653397Z digest=sha256:289fc45e0f582b927bb26745e4c849c1cd1c8f88f07ec114bf0e8475ce1cbb8d

Observation 519d403c-33f9-416e-a771-6143d673a16c · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.986399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.986399Z digest=sha256:5ceb96f215e5f93f947b37747202f17cbf39bc0ca5133c5d00903fc9b9db3a0d

Observation 7b19a446-827a-4ff6-87aa-f8fe4cbc08b9 · inbound

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? cites this paper.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Sequence Parallelism: Long Sequence Training from System Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.288961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.288961Z digest=sha256:682f2dce357815c171038a9919b1e3e758a1564c62fccf0990a18564f0f7ca1e

Observation 8316e377-e9ca-4bba-99fc-99db6a301f91 · inbound

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs cites this paper.

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:36.333821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:36.333821Z digest=sha256:06e81430baf747048c47d365bc1ddaff011bce528254e8ec9ad540379666a9d4

Observation aa75cc33-de9a-4262-beba-b1b0be105839 · inbound

ContentV: Efficient Training of Video Generation Models with Limited Compute cites this paper.

ContentV: Efficient Training of Video Generation Models with Limited Compute Sequence Parallelism: Long Sequence Training from System Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:32.085563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:32.085563Z digest=sha256:aa9c33d76eec74d84746e5f8c398b81d43118495ba17fbf26cd77ce315a445fa

Observation 4397b7e9-5a51-4462-97ae-4f0f935b4ab0 · inbound

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library cites this paper.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Sequence Parallelism: Long Sequence Training from System Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.922834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.922834Z digest=sha256:4eccb1eeab6d5524b793e61c87e5822b81cf32ae146ac9f77e16b1150c3fb84a

Observation 9786b61f-a944-441d-8011-a71f08b31f5f · inbound

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences cites this paper.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.105856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.105856Z digest=sha256:d29ad5334b11507e33d1926b2e81dc95f8c98f622f89bdfddce4fe550c6f4e3d

Observation b5dd0957-cac0-43ba-8316-4e8de9778731 · inbound

Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs cites this paper.

Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:08.905826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:08.905826Z digest=sha256:76db220b1dbecaedec8471e6ffe9ba94fdd27b336aecad62fa83d330ac00c4dc

Observation dc945350-74c7-40d7-a7fe-22216310527c · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:11.234042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:11.234042Z digest=sha256:b25bc64a0ddd5760a07faa09eaf5082bf99339fdd5075a1c8acbf8beaeaff83f

Observation 5b76fa2b-c631-4d1f-91e3-164d2e7593a7 · inbound

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference cites this paper.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.794986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.794986Z digest=sha256:3ded3e61ada9212b91e448d523f880fcacadd546352d6a50c49cd45673daf2d2

Observation 547516d7-5037-43eb-a675-17713898ed44 · inbound

TetriServe: Efficiently Serving Mixed DiT Workloads cites this paper.

TetriServe: Efficiently Serving Mixed DiT Workloads Sequence Parallelism: Long Sequence Training from System Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:10.837066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:10.837066Z digest=sha256:2a3711c94f65bb20ad8266ad5985368d879ebd8ab0a303ae135db4b0275a9d5a

Observation 0a064b01-f126-40e6-bf92-a8c2d30f1b24 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Sequence Parallelism: Long Sequence Training from System Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.032377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.032377Z digest=sha256:a3002a7d9ff901c73ba5401e7c3493767595bc289fa2564c4f893b6232146442

Observation 02222285-4611-4fc1-8a99-ed855fe0cc04 · inbound

Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems cites this paper.

Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:14.135198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T04:41:03.457680Z digest=sha256:023405ca81cf6166a2e2c5ec65e320b904216dc08df06d8b25906e0603ef661a

Observation a43b33c6-db05-495b-975a-439585037437 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.770043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T08:55:31.298030Z digest=sha256:275da8f99dd4dd723693993016a419a932b3771ad0f8c024341b9bc804d2840c

Observation 442c23a1-c1ce-4b8c-b020-79353d910f77 · inbound

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing cites this paper.

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing Sequence Parallelism: Long Sequence Training from System Perspective

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:53:24.750202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T07:53:11.761843Z digest=sha256:adbf9b4ce18e23c28872fd85b46dc6920d1f226074faf642573f4b22acdcb891

Observation 9525b37c-6ba3-41ad-81f2-6b08c1687abc · inbound

Online Dynamic Batching with Formal Guarantees for LLM Training cites this paper.

Online Dynamic Batching with Formal Guarantees for LLM Training Sequence Parallelism: Long Sequence Training from System Perspective

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:29:35.935578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T15:57:18.544274Z digest=sha256:67155a58178b0884f6b22378c6d41d1a46d50bde91cb6e353a5e4a946a0dea45

Observation a02cb5f7-b5af-431a-a4ad-2b877e49476e · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.807624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:f7b4edbb21b45434c6d3fa6db84700aef12d3faa7c7fd9ebc8ff885d9e443d97

Observation 69a5c0ea-6441-4a69-8384-16c246a3f958 · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:28:43.956670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-03T17:26:07.870260Z digest=sha256:df4a55f28e19fa5bf66ac6be84a15312baac9ad94659dac87ed70c1e529c390d

Observation a648ba3b-41c3-4a22-b4a4-38b97f72cd32 · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T08:40:29.554554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:40:29.554554Z digest=sha256:3e32250ef5f42d9176fefce7a0052df792bd5840a740c021ed2e7afde8b8458d

Observation f8b4783f-dd2b-4cb4-8f77-fd28935a3bad · inbound

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention cites this paper.

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention Sequence Parallelism: Long Sequence Training from System Perspective

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:27:41.863772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-03T06:22:23.985308Z digest=sha256:3373446e000377ec6498d054ba2d045cc946f8406f77214e21d82b8619d28914

Observation 6712b88c-b529-44d7-87b3-51e4c5d5f696 · inbound

Design-CP: Context Parallelism for Design of Protein Nanoparticles cites this paper.

Design-CP: Context Parallelism for Design of Protein Nanoparticles Sequence Parallelism: Long Sequence Training from System Perspective

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:52.764344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:29:52.764344Z digest=sha256:f8a636ba69c9b659fabf912077b8acc5fb53e078e2b18ce51184a33013b46543

Observation ad4c00d3-51d0-48fa-ab5b-91bdf8133d7a · inbound

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving cites this paper.

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving Sequence Parallelism: Long Sequence Training from System Perspective

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:14.069087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:54:14.069087Z digest=sha256:1a86714f2c18a947db003a3ad1977a9108c6e5426eee7b0f09b39fd6594acf5d