Pith. sign in

Paper Citation Record · LEDGER

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking

As of 11 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2602.21196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.21196 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:11:51.292724Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c99560fe-254d-461e-b429-c9517ab8539e · outbound

This paper cites write newline.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:47.906779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:47.906779Z digest=sha256:48ce2a41bb9dbcc918983c2d5ace9e906800fc7442c1bc694f07c29f1c90d68b

Observation 6b7a624e-99ee-42be-84d2-49e96250f034 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.019583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.019583Z digest=sha256:92658144e069f6acc78acc82f7c9d63abd3cd4a0d2786f736a9b78cbf6c39336

Observation f68b75ca-1d76-4926-b339-af583cf09ac8 · outbound

This paper cites Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.095184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.095184Z digest=sha256:d27a4aaa20a0ad1cda57a823f605f7513566d07575a2dc158c6d3f273b481d57

Observation afd8f254-9ddd-4193-9ac6-e7ee3efeba62 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.153620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.153620Z digest=sha256:19b7e6a7ab91c81cacfa3cdab7a4e87ceee1955970466339dedb89b5899e2810

Observation 96885c7e-b90e-4767-8fde-157aae7cddaa · outbound

This paper cites M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.243215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.243215Z digest=sha256:7f592a1efcbe807a9b41b745e46d7dc7ebd92c5033674b07d272d3ca04c2c12b

Observation a1e2d119-88ba-4606-bd02-3953d6af281a · outbound

This paper cites Y., Ermon, S., Rudra, A., and R \'e , C.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Y., Ermon, S., Rudra, A., and R \'e , C

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.328966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.328966Z digest=sha256:29f5552dbc95d78899d1c664fa9f338156f04c03279ec51acc68555cc99d2e10

Observation 8cda7e97-6288-4297-9187-bcb71248e09d · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.416657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.416657Z digest=sha256:ff9e918aac852675c9dc1f6bcfe0f30f91b7479a980820d76c14d649aed3d841

Observation e1f768da-03cd-436e-8b4a-4fb7ae1a4679 · outbound

This paper cites Gemini 3.0: A new era of intelligence with gemini 3.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Gemini 3.0: A new era of intelligence with gemini 3

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.587505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.587505Z digest=sha256:6edf9f6ad831ecd34c516dc8a463d62bdcc7198b3c5b473b1d07354b40559a02

Observation 55cd6d9f-3cba-42f5-8de9-4db081f938f5 · outbound

This paper cites The Llama 3 Herd of Models.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.708496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.708496Z digest=sha256:b08cae3304e085f9de14ef569508821056a6e1d8287d552e740a4371507b9b58

Observation 3b8a64e6-f660-40a4-99ac-8949748bea35 · outbound

This paper cites Unsloth, 2023.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Unsloth, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.861211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.861211Z digest=sha256:42c1d6ab369a793fdba90257ed77ad8235f19890736106ca1f805cfa193a1217

Observation 2dcfedae-0636-4195-a19d-59acb3f1eb3f · outbound

This paper cites Advanced Long-context End-to-end Speech Recognition Using Context-expanded Transformers.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Advanced Long-context End-to-end Speech Recognition Using Context-expanded Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.953565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.953565Z digest=sha256:3857faf1c5442f1362f38cfc816eee997bb90185d8fb06a2e9b2e19b1f0ac1bc

Observation 4190dcf3-bdb3-4e09-a596-5bd0f0cdfc3c · outbound

This paper cites Liger-kernel: Efficient triton kernels for LLM training.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Liger-kernel: Efficient triton kernels for LLM training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.060784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.060784Z digest=sha256:fae70872ef958a3d8d7bd13eef6cb87759d2bfb0f0ef5f19a5e241af67f8219e

Observation 001bd1f1-5071-4097-8ae6-d3ae4e2eaf8c · outbound

This paper cites Qwen2.5-Coder Technical Report.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Qwen2.5-Coder Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.195785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.195785Z digest=sha256:3e8e0420f8d3ce2e00c071c3f8492afc0d329e1469ec2aed0c57a62008c87eef

Observation ea053d39-524b-43bb-b29e-3471a748a529 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.350958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.350958Z digest=sha256:0cca27fde29fd64bbb359f3bea176371e197695fb287f3b13a9870381b739555

Observation c96becb5-636f-4967-96d9-4265da21b165 · outbound

This paper cites Mistral 7B.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Mistral 7B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.493389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.493389Z digest=sha256:81b38091429ba035563f93bcc55d300f8d979155f0a2b059b46a8798ba23e50b

Observation 25749982-9c90-41c1-a6d2-5bdb6013d582 · outbound

This paper cites LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.632698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.632698Z digest=sha256:8d8bb100e18c484707f9c9b62541984c015468d319c8366a6cd2a0c4b3b9eb53

Observation 8edc2120-d570-4a4a-8b49-e52572cbe760 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Kimi K2: Open Agentic Intelligence

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.698578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.698578Z digest=sha256:19acdb88cd970a0ba98d2be6fd46221406ce5add5c9b238f9ca7615b93748b85

Observation 25597e34-7a25-4688-a37f-d0c9ac86d4fa · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Reducing Activation Recomputation in Large Transformer Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.812534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.812534Z digest=sha256:9182167b1b40cb003e3952decc7d0e06ab21ec9e4bc029af231ec8eb6354a573

Observation 5563856c-5726-4244-98e2-77337331e676 · outbound

This paper cites StarCoder: may the source be with you!.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking StarCoder: may the source be with you!

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.946228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.946228Z digest=sha256:dfcdab13fb6cd4fe6254c42300a3123d6b0cbcb3174369723296d97e90622238

Observation 0a064b01-f126-40e6-bf92-a8c2d30f1b24 · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Sequence Parallelism: Long Sequence Training from System Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.032377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.032377Z digest=sha256:a3002a7d9ff901c73ba5401e7c3493767595bc289fa2564c4f893b6232146442

Observation 67abdf03-4eaa-42b9-a064-545937294de6 · outbound

This paper cites Sequence parallelism: Long sequence training from system perspective.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Sequence parallelism: Long sequence training from system perspective

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.141072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.141072Z digest=sha256:419ed3d4c90df9b947921881ffb8859b621fde13cdf3890569b423ed4abeff0f

Observation eee07d3f-0a17-4640-afd9-3d4765fae34e · outbound

This paper cites Torchtitan: One-stop pytorch native solution for production ready LLM pretraining.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Torchtitan: One-stop pytorch native solution for production ready LLM pretraining

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.244468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.244468Z digest=sha256:53585db15f88728a07a84ce19e82f7e672d672315a7cfd29a1c36c29eade7973

Observation 2082e951-a090-4af2-8b19-a5badeed4f86 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.310017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.310017Z digest=sha256:8cc45a11296e3c766fb44834bc8979b7e7f947088fdd522573b6521f8ff39f5d

Observation f50b51a8-d40d-4473-ad95-6f265343e380 · outbound

This paper cites MAGI-1: Autoregressive Video Generation at Scale.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking MAGI-1: Autoregressive Video Generation at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.391061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.391061Z digest=sha256:b8bc21ee45526ee1a8e872f06ff5004a7e596ad73abdcbc3d96d669474154b30

Observation a9875024-4b17-4f39-b112-889bd3585726 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.475375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.475375Z digest=sha256:5d44f56288cc62866e5f4baa1c365500a9430b454d8a576e87021c42171efb90

Observation fb5aa8ca-ed68-452b-8ff0-a1db0ab6c9c8 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.563806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.563806Z digest=sha256:bbdb63748099dde52616ed071ee8dd5e994e4db5e265b397b952b3f0d5636596

Observation c3ff2873-f83a-469e-ade1-6898833286c3 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Wan: Open and Advanced Large-Scale Video Generative Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.645700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.645700Z digest=sha256:68cd35ad418f436330ffd463756a63dc05fe1389ce84eff8c48222ea89407e45

Observation 6f46242a-be0f-4f88-a349-e6cf3f157d53 · outbound

This paper cites N., Kaiser, L., and Polosukhin, I.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking N., Kaiser, L., and Polosukhin, I

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.774863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.774863Z digest=sha256:1554833f664278ce1c60ddec8531abceacb3f6491f2243e683ff7981ca13f963

Observation 83960181-6e2b-4ede-8d1f-f2233a667b44 · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking HunyuanVideo 1.5 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.925574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.925574Z digest=sha256:75eeacd52879dc0e66ef0282398fba7a4c7cdeec027b32a2f7e37dc5bdf40efe

Observation 7cb410f2-0c61-40f6-b4b6-97a6acbd66e3 · outbound

This paper cites Qwen3 Technical Report.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:51.044030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:51.044030Z digest=sha256:6a4cd259ed7452d4a6c5ca105e7e5a6b26cf163b59d33be7db70560dd2caa7a6

Observation b78d2d07-028f-4631-9aff-1f837a23dda8 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:51.164939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:51.164939Z digest=sha256:fabb3af8adda600a430fe2a28b7f6c6bfe3deac3fd219ad26c14322030e1c961

Observation f7629dfe-3273-4b4e-95e1-5320c0571101 · outbound

This paper cites Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:51.253437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:51.253437Z digest=sha256:4dcdb8ed88068578165db0f3b0a36d595d0c363f55af9044cb86695fb966fd41

Observation f4cfafaf-795e-49b0-bafe-c0b38fc668cf · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:51.292724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:51.292724Z digest=sha256:f40438552ee555215a3651779e84f30a26400eeaf086a6735286cf35d5d273d3

Pith citing papers

No inbound Pith citation observations are available.