Pith. sign in

Paper Citation Record · LEDGER

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences

As of 23 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2506.13996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13996 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:30:57.837571Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:11:48.095184Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40367eda-1cf8-4460-8d22-17d7ec9af6c2 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.473062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.473062Z digest=sha256:b0478a6e9362af72e0a8634363206995b989e1b420b41d6e36283cc763283d5f

Observation 97957d92-9f4b-4298-bd7f-1b37b5bb3d3a · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.703220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.703220Z digest=sha256:daedf43cc6f306c7fe0bcf04f077faceb1856a3a3eb9fc934be2f5877b0cd717

Observation 05cf9cdb-0e18-47b3-8b67-145d5b4307f4 · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Reducing Activation Recomputation in Large Transformer Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.938309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.938309Z digest=sha256:47601e77166a4f63a6cc7d9d9a053ce1b590f84f3381ea873c21727c234744b4

Observation 9786b61f-a944-441d-8011-a71f08b31f5f · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.105856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.105856Z digest=sha256:ff625cc3a37d20c8ee148e4cdeb1c162bc8669ea227fc91b6a01c3f3c3515214

Observation c4684dc0-5dd9-423f-b97d-1dde205943c9 · outbound

This paper cites DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.230466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.230466Z digest=sha256:a676446c9ad73c4314e5e73288a0b0cba2228203b639921331bb18b773c0aeaa

Observation 4cf2c8ee-c34e-4f71-a054-6050612b2b61 · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.281690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.281690Z digest=sha256:53402b77c4681c1492ab0270f41235100bbd84e95c819e950227dd07a02cb9da

Observation 9d9e0298-ab80-4bb6-8aba-7eadca48f60d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.362359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.362359Z digest=sha256:a7fe0a3df33b8055991f89d855bca26a4f9f0c6f02c6f0fe117d7a009538bd11

Observation 3d33a8a5-0944-4486-9b2e-f980395fd95f · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.510416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.510416Z digest=sha256:c42aaecbc5a3903bcc78bbafd38a5e12c7b0abfddb6cbe2980377cae13abe53c

Observation 56cbb9a8-b30c-4123-8fb3-a696124d50ce · outbound

This paper cites LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.602093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.602093Z digest=sha256:58b44a59a13fa5ed7ea3724477c04b1c2f5e273d4d1ebcaae77d37c5a4151e33

Observation 6ba52778-8e90-491f-a35d-e8bedcd57bde · outbound

This paper cites Liger Kernel: Efficient Triton Kernels for LLM Training.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Liger Kernel: Efficient Triton Kernels for LLM Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.663570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.663570Z digest=sha256:2f2d2fad796243ffe38f02aba82e3ac92c8be637a7b593b23364197a2924c096

Observation dd1ff85e-5006-4493-ac35-44d4b275336f · outbound

This paper cites Datasets: A community library for natural language processing,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Datasets: A community library for natural language processing,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.266013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T00:30:57.742706Z digest=sha256:2771bb9569b2e0c15b44fb8c3ffdf34edc224ae6cfc82d74b4742f55be088c80

Observation fd638e83-7865-4f24-b9b6-1e00006d7cc6 · outbound

This paper cites ArcticTraining: Simplifying and accelerating post-training for large language models,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences ArcticTraining: Simplifying and accelerating post-training for large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.250240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T00:30:57.748571Z digest=sha256:e08d37fa03f90581209ba020b18bbb5a44c67398e12f0824470b108770ee59e1

Observation 622566c1-dd5c-448b-9a97-c4efac6fe13b · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.786155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.786155Z digest=sha256:9d894842f50eb78b27cf2dc0df7d7f83a41ac6234f6da2101e295bbacee56eb0

Observation ca5d832f-b92b-41ae-a63e-3db5732e2929 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Transformers: State-of-the-art natural language processing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.233466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T00:30:57.805996Z digest=sha256:251edf1a354b645150f52811d674ffbf2b6b9b53bfcfc2f6f207f1b1c75da0c3

Observation 26203125-424d-4dc8-ac48-05c8314a0a3d · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.822249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.822249Z digest=sha256:83d467b3b71cff3ee56aeaba32c1ad5d604c00b9f23cd6726cdb55dc0f63ee1b

Observation 19ccebd1-f9c0-45b0-91fe-236b9f080d5e · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.832992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.832992Z digest=sha256:20e9d1b7e290df1d2fc56c22dd0736670b8cbb5338a7f048b565e881ec738847

Observation b8d73282-b716-4528-85b6-a516858600f3 · outbound

This paper cites The Case for Co-Designing Model Architectures with Hardware.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences The Case for Co-Designing Model Architectures with Hardware

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.837571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.837571Z digest=sha256:7938248e2bf6b321495f4c74122916b5b59b3396caf5abb4b9f39181c4467709

Pith citing papers

Observation f68b75ca-1d76-4926-b339-af583cf09ac8 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.095184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.095184Z digest=sha256:8f42410763dfc704d858e932cad788e670cae5a36ac5d51024a5fe69873a9f70