Pith. sign in

Paper Citation Record · LEDGER

Striped Attention: Faster Ring Attention for Causal Transformers

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2311.09431.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09431 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:33.063209Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 19e865a1-8439-4196-a980-63f72b256b16 · inbound

Gated Linear Attention Transformers with Hardware-Efficient Training cites this paper.

Gated Linear Attention Transformers with Hardware-Efficient Training Striped Attention: Faster Ring Attention for Causal Transformers

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:15:14.087284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T01:15:13.991219Z digest=sha256:b79e88ee63e2d34a6f5b4d3bc96d05f6666fff17118362e0c147be3e8c9d331d

Observation 47441c1d-c0c7-4f0b-b054-e6069f9a8e85 · inbound

World Model on Million-Length Video And Language With Blockwise RingAttention cites this paper.

World Model on Million-Length Video And Language With Blockwise RingAttention Striped Attention: Faster Ring Attention for Causal Transformers

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:36:57.210663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T06:36:57.165551Z digest=sha256:3df7758344c9b145b0f861f79f68135515752a9890e9d046e1d01533bbeac94f

Observation 5adb0167-f8de-4cb1-a868-e75ec101b984 · inbound

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality cites this paper.

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality Striped Attention: Faster Ring Attention for Causal Transformers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:16:25.699945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T12:16:25.390683Z digest=sha256:a1dcb50e218812aae890881894563f1d778cfffb9cdcd94b428adbb06d5aefd8

Observation 0edaa502-2638-4306-af23-19d83fadc09b · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision Striped Attention: Faster Ring Attention for Causal Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:45:36.459905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:2af57eb10872fae295e8079015d2cb8545ac1ebf63720702174efa3684e76ea0

Observation 6f211ab9-c326-409b-8d95-3e4a206842ef · inbound

FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism cites this paper.

FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:24:48.791822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:24:48.791822Z digest=sha256:c7183b4ead68836037de08324518b2a435097df127479fdf089959b14d11ec54

Observation 6bdd5a1b-c521-4d2b-a238-f346c1343156 · inbound

TurboAttention: Efficient Attention Approximation For High Throughputs LLMs cites this paper.

TurboAttention: Efficient Attention Approximation For High Throughputs LLMs Striped Attention: Faster Ring Attention for Causal Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:17.585230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:17.585230Z digest=sha256:3e97b7ecc0c8243e074bf0788db4f5e48354245dabb8789fcd1cf16f8bcf2340

Observation a5f4dc18-95ea-4a28-9974-68f6d893dfd0 · inbound

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication cites this paper.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Striped Attention: Faster Ring Attention for Causal Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.892673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.892673Z digest=sha256:50a0fa8d3a21342e05fdbd268844f56ba30ffeb371e37d4b9efaf2d7927f1a6d

Observation 441df307-bd25-4c85-acb1-14b589bc797c · inbound

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models cites this paper.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.430129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.430129Z digest=sha256:591e81f5b0f03c29e5b1d2c30b7e8f7f4c6d5a00afdf37cd6ed2a5892a455386

Observation 747eacf6-a6bc-42f1-adce-fabdc4395184 · inbound

SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training cites this paper.

SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training Striped Attention: Faster Ring Attention for Causal Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:33.063209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:33.063209Z digest=sha256:c63bab50166cd691ddd77ebba171187c51a65fa00fe82ca22350872cfe28f73c

Observation 091c1abf-aee0-4067-8db9-777408e2aedd · inbound

Efficient Pretraining Length Scaling cites this paper.

Efficient Pretraining Length Scaling Striped Attention: Faster Ring Attention for Causal Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.576790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.576790Z digest=sha256:0c3e267a6042ecf4091f3e0048177d3c6bb6fbdad67f2782bb185c6c28397d4d

Observation 0dd6a633-91c9-4e0e-a58a-42cb02ec5a8d · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving Striped Attention: Faster Ring Attention for Causal Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:09.697978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:09.697978Z digest=sha256:b58ed2d6f72d1a61928ab37dade15239974c9b29f4ceee48b142fed7ce932c9c

Observation 9100d0d1-f3ee-4ac5-b576-db660990928e · inbound

Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations cites this paper.

Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations Striped Attention: Faster Ring Attention for Causal Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:32:40.677570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:32:40.677570Z digest=sha256:e655998a3fa4cfde5f0c182ce7008669c47fe916234ab40dc5e8541bf2bd624b

Observation 828659ad-2bf4-48ac-892d-6e32b7519cd2 · inbound

DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving cites this paper.

DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving Striped Attention: Faster Ring Attention for Causal Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:05:06.265472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:05:06.265472Z digest=sha256:581e6d2496f89c48d74f4e2bc6a5b14a541fdcf8c68d11223c02ea512b4861aa

Observation 4cf2c8ee-c34e-4f71-a054-6050612b2b61 · inbound

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences cites this paper.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.281690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.281690Z digest=sha256:610fdd3c72b4cc323d02668d3fc9d67d0c9c54f2866fdb594d29ed85d1af079e

Observation 60063694-1e0a-425f-8e3c-e3a782aa45dc · inbound

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training cites this paper.

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training Striped Attention: Faster Ring Attention for Causal Transformers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:06:27.320262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T14:04:31.017142Z digest=sha256:252856f3ba3485cd155b0270a42b91ad7f1c6256a00050a1d8c309318c096043

Observation d419d5d6-aaa8-4dda-a17d-4368d07e05c0 · inbound

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training cites this paper.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Striped Attention: Faster Ring Attention for Causal Transformers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.650627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:a3287f2be5cc8ae561dd5b04d09475228df000780acf4697787677fb54bf45fa

Observation 4a38f240-de5f-4dd4-8401-72a530ab8b19 · inbound

DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training cites this paper.

DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training Striped Attention: Faster Ring Attention for Causal Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:51:58.967510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:51:58.967510Z digest=sha256:b96aa887275bd6c65872d9ba3d45326d716a4500505780e8a0effcdb8807aaaa

Observation 8369c83e-4273-4f61-97ee-f9ec7db5fb9b · inbound

Efficient Scaling of LLM Training with Flexible Context Parallelism cites this paper.

Efficient Scaling of LLM Training with Flexible Context Parallelism Striped Attention: Faster Ring Attention for Causal Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T20:59:24.267345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:59:24.267345Z digest=sha256:f145698b430a4a978f2b20bbb5b7d7480aeef10b7b2ea27940d680145c9e40f8

Observation 9f8c18c4-ec94-4b1d-a563-956f55225fb4 · inbound

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies cites this paper.

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies Striped Attention: Faster Ring Attention for Causal Transformers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:43:00.722585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:40:55.343169Z digest=sha256:435e917d73ce559625693ea2527da462610a3562930585894587696020b7ed8a

Observation a2efe191-b917-4261-a5c9-faead9eb34dc · inbound

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism cites this paper.

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism Striped Attention: Faster Ring Attention for Causal Transformers

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:20.531274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T10:23:19.236803Z digest=sha256:9a968da93a456217fc6abaf7969c6d4280b562dfc24032a3c8459aad6928b3ec

Observation c671cc5d-3024-46e1-b9f4-bd300002b7c1 · inbound

HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware cites this paper.

HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware Striped Attention: Faster Ring Attention for Causal Transformers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:00:55.656684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T03:00:19.355357Z digest=sha256:59723c6aa7471605feb355218caf17ed47345b8b612d8a3b28fbcba661551cca

Observation ec20bb38-ed73-44bd-b14f-3849e8e7b49c · inbound

Kaczmarz Linear Attention cites this paper.

Kaczmarz Linear Attention Striped Attention: Faster Ring Attention for Causal Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:23.166304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:15:58.330766Z digest=sha256:8d9cac3c92736ab2a7298062e8c1110b71015c95f0ada0550966e93219f6945e

Observation 3a744266-4670-4c50-8d2e-c6614c944168 · inbound

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design cites this paper.

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design Striped Attention: Faster Ring Attention for Causal Transformers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.420530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T12:06:46.806138Z digest=sha256:c3cb96752bdeffc95371f77458a1fd9ad68a2876d0e08948a483c4086f4f5b89

Observation 3f95ae7d-bc46-45c3-a17e-399dac150bd6 · inbound

Towards Distributed Inference of LLMs on a P2P Network cites this paper.

Towards Distributed Inference of LLMs on a P2P Network Striped Attention: Faster Ring Attention for Causal Transformers

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:15:07.942418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T23:13:57.705287Z digest=sha256:f7fad66e4f86e9ad234a5d4403fe75f8244c82ae2bc2895e006eccbfbf9008ff

Observation 7583702a-b018-4649-b582-ea1fa8804ad2 · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.802503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:faaf01b5f0c231fd4e55905bfd3c09bb28eb4ff6315b2c81ec0f84096ab89502

Observation 3ecc045f-c542-4b29-8ee2-0d29a8121101 · inbound

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training cites this paper.

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training Striped Attention: Faster Ring Attention for Causal Transformers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:37:41.994384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T06:27:48.935774Z digest=sha256:13f1853859858c339de85224a09e6fa123a69206d6379570a597d7f911e400fe

Observation 7fcde856-0cf4-4090-97fe-29b541eb4e23 · inbound

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention cites this paper.

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention Striped Attention: Faster Ring Attention for Causal Transformers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:27:41.858999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T06:22:23.985308Z digest=sha256:29de16ed2f3385c3274da01bdbaaf88859903318cb452dee75ac1de660a7b891

Observation 1e3b4bab-c379-4bc9-98ad-e1d195f3e73d · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Striped Attention: Faster Ring Attention for Causal Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:6f9b95029187c2ce379c0352d6dd4da5cc618d6ee781c79719f24b0d12961bb5

Observation 48fc9efd-7438-4d22-9406-373f497c5deb · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Striped Attention: Faster Ring Attention for Causal Transformers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:36.816550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:36.816550Z digest=sha256:ec8dca3b993f1b4ce4dd3e17d5ae5aafb78d0695ffdcd9d666e5d8f9dea898f4

Observation 6f0bc92b-1dce-4bd6-8387-3c2dacb2ea9f · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Striped Attention: Faster Ring Attention for Causal Transformers

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:43.565950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:43.565950Z digest=sha256:8b074ff5bb53bbf8a1c2ef29ac1608170ab65728a676ecd8dbb6ca12798125f2

Observation 34d83ea0-6ca3-400e-babc-b547b3e963da · inbound

Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool cites this paper.

Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool Striped Attention: Faster Ring Attention for Causal Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:06.819097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:06.819097Z digest=sha256:8c940719173e97d127353b3f4ea1c6b6722abcac199c174ce56b6cf5fbd0762a

Observation 199e88d6-ca7d-46ff-9779-2cd69ff64e3f · inbound

Motif 3: Technical Report cites this paper.

Motif 3: Technical Report Striped Attention: Faster Ring Attention for Causal Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:12:09.873268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:12:09.873268Z digest=sha256:c0949f8da9caae1adb04570cccb649afa9298cc58570763c635f77d6cae17a72