Pith. sign in

Paper Citation Record · LEDGER

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2506.13996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13996 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:30:57.837571Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:11:48.095184Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40367eda-1cf8-4460-8d22-17d7ec9af6c2 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.473062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.473062Z digest=sha256:0e81e74e40148d00ee6c402516feeba9fe9f5eb32c974fbff8d4b29c0dbaa837

Observation 97957d92-9f4b-4298-bd7f-1b37b5bb3d3a · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.703220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.703220Z digest=sha256:b6eb8cae9f161a264c90c8d42897cbdc6884aeebcb26f9e3de7981b48e0e987c

Observation 05cf9cdb-0e18-47b3-8b67-145d5b4307f4 · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Reducing Activation Recomputation in Large Transformer Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.938309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.938309Z digest=sha256:b3b57359898b893ccfe821dc650a1bd56d9e754ef38e6b67738d643c6d689b7d

Observation 9786b61f-a944-441d-8011-a71f08b31f5f · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.105856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.105856Z digest=sha256:54021c785bac5cba9bc087e401626804082864ae351eba0a6351b851a1c387f1

Observation c4684dc0-5dd9-423f-b97d-1dde205943c9 · outbound

This paper cites DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.230466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.230466Z digest=sha256:0d761c0fb4246e4e87c8d23156789781af2ca43448feba9e0db81189ba9eb446

Observation 4cf2c8ee-c34e-4f71-a054-6050612b2b61 · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.281690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.281690Z digest=sha256:5926398919c3678b9824ebf08efdd453172271435abd491936e6ab759870b995

Observation 9d9e0298-ab80-4bb6-8aba-7eadca48f60d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.362359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.362359Z digest=sha256:55b48895d2f90a81e783dc7900795443bac0739c6920a6663efd5b628d6aa36e

Observation 3d33a8a5-0944-4486-9b2e-f980395fd95f · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.510416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.510416Z digest=sha256:189ee109a976f044edda66d9b92f7a67e0ba568942a7157453fd5ca164bff15c

Observation 56cbb9a8-b30c-4123-8fb3-a696124d50ce · outbound

This paper cites LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.602093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.602093Z digest=sha256:5bacbd2f12313af43a101afd0f9b370fc274ea72613c0f9cb33b1108a025e63b

Observation 6ba52778-8e90-491f-a35d-e8bedcd57bde · outbound

This paper cites Liger Kernel: Efficient Triton Kernels for LLM Training.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Liger Kernel: Efficient Triton Kernels for LLM Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.663570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.663570Z digest=sha256:ce43d70a6da565ade2469f79648cbf95f33dff27c03abb8b3d23816968c0e3f2

Observation dd1ff85e-5006-4493-ac35-44d4b275336f · outbound

This paper cites Datasets: A community library for natural language processing,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Datasets: A community library for natural language processing,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.266013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:30:57.742706Z digest=sha256:fbc5b0bb544cf55f531b3fbd918c3a58091e449953f75e15f54fb2a712f12d5f

Observation fd638e83-7865-4f24-b9b6-1e00006d7cc6 · outbound

This paper cites ArcticTraining: Simplifying and accelerating post-training for large language models,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences ArcticTraining: Simplifying and accelerating post-training for large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.250240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:30:57.748571Z digest=sha256:b004bbfcadec9ce1ab70f6207cb3c44caf232c656bbb1935255639c3c8e5cad0

Observation 622566c1-dd5c-448b-9a97-c4efac6fe13b · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.786155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.786155Z digest=sha256:9673818baccedfd462bec7c5b727b0e12df5f0552dd871345f9391c0fae1fb2d

Observation ca5d832f-b92b-41ae-a63e-3db5732e2929 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Transformers: State-of-the-art natural language processing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.233466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:30:57.805996Z digest=sha256:39a9cf98222810f94b0c79fd32f0f76c8f89caa7cbb947d690663b3ddf5c513a

Observation 26203125-424d-4dc8-ac48-05c8314a0a3d · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.822249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.822249Z digest=sha256:e1ac7ec1a7cced885b6b308c33ee6bf586f055d02f2c3e4d7fe100f6703539eb

Observation 19ccebd1-f9c0-45b0-91fe-236b9f080d5e · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.832992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.832992Z digest=sha256:032ceaf3c784abe6b3e379a36bccf2a9bb1ede00f958e2a4f36369b3957f408d

Observation b8d73282-b716-4528-85b6-a516858600f3 · outbound

This paper cites The Case for Co-Designing Model Architectures with Hardware.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences The Case for Co-Designing Model Architectures with Hardware

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.837571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.837571Z digest=sha256:58dde3157cc727ccdf70277a3a6817e54d4604636243db10b04ebf46ec5d4419

Pith citing papers

Observation f68b75ca-1d76-4926-b339-af583cf09ac8 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.095184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.095184Z digest=sha256:e4b1c9797baf5f0ce423c36474868613b254c0dc184adbffe9b631dee1c4383e