Pith. sign in

Paper Citation Record · LEDGER

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

As of 5 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2510.18830.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.18830 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T19:44:04.833504Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T10:50:12.926232Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T10:53:13.511836Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact26
  • verified fuzzy40
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5156b2c3-1e9e-49dc-8745-21a4ec1cc503 · outbound

This paper cites Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.617697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:cbdf2f230df47f190823c936f53029c2ffd68f9e76fab1ca60ed435b001d28d2

Observation 8b59d568-3960-45b1-a816-be2561950209 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.671447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:0ac8c57cead77cff78b8ab44c4f88af65e788822859fc29ce59f1905b6a897f2

Observation 11946e18-9bf2-405b-92bb-d7f9f5539dc8 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.681275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:e9c6414eebca024653877a04b89d6c42fd330f11514d8dd269f5640f6d0cea71

Observation af577f49-b0c7-482a-beaf-219c289e88cf · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.814875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:3fbf53943ba0251d1169a0b87f9894ab1b82aa6ff491105bb716b613d22972fc

Observation 1b1a1cc1-1fed-4755-a618-e6a580edc27d · outbound

This paper cites Introducing deep research.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Introducing deep research

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.796990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:845472a3de4651c8f0e2beef486abd2c8e2ad9f45afd4e94e5e672e58b982a8b

Observation 501e3336-eb8a-4efd-ba3f-236e59fe70fc · outbound

This paper cites Leave it to manus.https://manus.im/.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Leave it to manus.https://manus.im/

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.770143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:3c7df1063ffa10003a8b88b240a7df20b859716837c19766231f92ca302ffe5a

Observation 3967738b-0bcf-4abc-b923-f82fb79bbcf5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.666724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:22df66d4ac082c3243484fd0d1eb34284fa1cf42828d386ea55d2817ed61bcd4

Observation 3c2098a1-d92b-47a9-9f8a-72f3df8d1add · outbound

This paper cites Qwen2 Technical Report.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Qwen2 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.701904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:b6b0099526ed83c23ece2f3cd92ed1a10b92f68f4080b616d2946ac6af93af63

Observation 078993bc-e5d7-4298-ae2b-49d7427f839d · outbound

This paper cites DeepSeek-V3 Technical Report.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training DeepSeek-V3 Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.684088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:cebed33b09ab9cde2373adceafe0e2f8b0682a46b790c6ed38db6990e236853a

Observation 365dd40c-f9b7-4744-b502-aafdfd5e2728 · outbound

This paper cites Qwen2.5-1M Technical Report.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Qwen2.5-1M Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.692165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:c4a318aafeb43bc1d9dcfa6b9f91fbcd27248c473bb9257f03f7c301c8bf2780

Observation c66bbf04-33bf-48a5-887a-ef0cbac2436b · outbound

This paper cites The Llama 3 Herd of Models.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training The Llama 3 Herd of Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.647423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:bb5ba44c0c8f3b363f657efdabdf3b9e059b8a7258bb4bd6d8f65b196cc3d6bd

Observation 5e1ac85b-cf30-4fff-a625-a69ab91f625e · outbound

This paper cites How to train long- context language models (effectively).

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training How to train long- context language models (effectively)

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.660631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:429e0dcaafab2eb71769125fe1a79fbae6cdd12cf6db8d2c329708533eab23f0

Observation db7b691a-988a-4892-99b9-015b8950cfb1 · outbound

This paper cites QUEST: Query-aware sparsity for efficient long-context LLM inference.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training QUEST: Query-aware sparsity for efficient long-context LLM inference

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.810291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:7ab7b69aededb1901f78af038a2d3c166154cb3635ba6bb2a98eaf8c7b69b1b8

Observation e3d1aef4-6538-43b8-a3c4-70a452dda11b · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.764076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:5a7c34ae2cbaf7398fab5ccb4d80336c0e8be24255f6001327689b6fa7ecf369

Observation 99781537-8690-4c73-b4ae-59911756f801 · outbound

This paper cites Flexprefill: A context- aware sparse attention mechanism for efficient long-sequence inference.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Flexprefill: A context- aware sparse attention mechanism for efficient long-sequence inference

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.753640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:df0e9d544c0a1103443f20e9aa83e1df9ad4f4ccdcdf7a8eb623a055876a92ee

Observation 2530f5e5-728d-42bd-84b2-79dab862de57 · outbound

This paper cites Sparq attention: Bandwidth-efficient LLM inference.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Sparq attention: Bandwidth-efficient LLM inference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.767266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:9de63265d6e02b76e9739c8d735eed0368945a1101c3309083339c5a525bf3cf

Observation 7620b104-d95d-4ebb-8ee0-b469f1ffc6c4 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.618794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:ad79d5c5e93680fa559a5950c2f311c4832e87d2d71998dbc209af8be48461f3

Observation 52c40fb4-c89c-4206-aa23-b0b1a830503d · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.643653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:295961e78602be2e64bd645da6fba3a30687fb9a956ae789d1857ac1994bc47c

Observation 0bedc2a2-31f4-41fe-8952-7f1b8aa92efb · outbound

This paper cites Ring attention with blockwise transformers for near-infinite context.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Ring attention with blockwise transformers for near-infinite context

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.816601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:c7f892986c6bfe47c186dbb0ef4ff3680609d1f5cf6b0bfb1c2fc09b46711821

Observation d419d5d6-aaa8-4dda-a17d-4368d07e05c0 · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Striped Attention: Faster Ring Attention for Causal Transformers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.650627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:7c68fa00920767b93c277cde0635bec5496999adf27db1bfbc56a8362980b5b5

Observation 5ea97251-5024-4118-89b3-d8afe3ffeca7 · outbound

This paper cites Qwen2.5 Technical Report.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Qwen2.5 Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.695297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:aacd239851b7fd462efd488124cb7d18ead98c56d8d78cc0f9cbbe9c0afa2c56

Observation 8f620287-2ec2-4a1d-a770-223e185746ac · outbound

This paper cites RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.771894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:ff9df244fb4550165a9f17fd01b689473fd13875782f022368b39fb090743a96

Observation 3c7d8006-3b1c-4c2d-bbb3-47c47c934954 · outbound

This paper cites Needle in a haystack - pressure testing llms.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Needle in a haystack - pressure testing llms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.789323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:bed034d646016c93afa6853d28bf002db26b458d933c9f5e4cd24f1bd98962d8

Observation b4744c0a-c978-43a1-9840-5f7f976ccc23 · outbound

This paper cites Infinitebench: Extending long context evaluation beyond 100K tokens.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Infinitebench: Extending long context evaluation beyond 100K tokens

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.803646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:b0fe4d43db837c292c51aca43b36d75bcaae2f870370d0e44b268dd29f351ec5

Observation 9acf4671-7194-4cb0-a65e-e8227b7d8ab9 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Compressive Transformers for Long-Range Sequence Modelling

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.647760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:fc695040b081369cea0b8a6098fce478e6632940016ca806f7a5535054ddf533

Observation 0f2178f6-c1cc-4d47-819c-b161ceb74655 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344–16359.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344–16359

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.812799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:750e7fad24b30e23d87a90ca6ca8f9ef9d11b5454622e38253fe71f44c40fe9a

Observation e350ce1f-9611-4814-8d8a-287d81cd19c2 · outbound

This paper cites [Feature request] balancing computation with zigzag blocking.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training [Feature request] balancing computation with zigzag blocking

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.826498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:7d67a2c2c9c673d699428edec51153008850bc4e3cfbb967775f2e0125ef2447

Observation 123ba3fd-bb6b-49f8-ae64-83c38e133c4c · outbound

This paper cites XAttention: Block Sparse Attention with Antidiagonal Scoring.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training XAttention: Block Sparse Attention with Antidiagonal Scoring

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.675195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:db02bddaebe2edfeae235535047e84a4aa905714575022855679637b2da09fe6

Observation f6d17fc7-3803-45a0-98c1-4b4fb9709656 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.819660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:e8208a2e6078aab3cea639361701830479df7666baa4af4de3dbd90ac62329e0

Observation 3cb32a3f-1115-4ce4-ab42-b1adb59684f9 · outbound

This paper cites LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.688737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:687bebb7873bd10431c2f0ab5db80a90f19926ca118fffdfc7cad4c60f68050e

Observation 0e24abef-8f4c-4ab8-93b9-45eab6506218 · outbound

This paper cites {nnScaler}:{Constraint-Guided} parallelization plan generation for deep learning training.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training {nnScaler}:{Constraint-Guided} parallelization plan generation for deep learning training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.821999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:33e436138eb84865faa380ede0e7754e3587d7aff1669f05a29b26397fc29b7c

Observation 8ccb14f7-86a5-4fd3-92ec-5bebc232e287 · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Zero: Memory optimiza- tions toward training trillion parameter models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.733256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:3a0dd70a7ff48bdd86e9c229bff18cf079a5f552c54586a0aaacdcfa4d53505a

Observation 2dd9909d-2552-46a2-981d-bc4b8af800c1 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.Advances in neural information processing systems, 32.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Gpipe: Efficient training of giant neural networks using pipeline parallelism.Advances in neural information processing systems, 32

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.801230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:aa9bd925ad9e245a3f279bc89b5bb98bdbecebbd37bc9f789e70282a11e7cf67

Observation 9088f2ef-f293-4b82-a8bc-c10788acca88 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Training Deep Nets with Sublinear Memory Cost

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.680299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:8684d4b42178c601f7fb3e25888a93f66f8950d5872ccf5375233000e84acca3

Observation a5154ded-68f1-4877-831e-01769cf4a23f · outbound

This paper cites Block Sparse Attention.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Block Sparse Attention

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.808064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:d747e55fc5e3ef71a2b428d4989cb8962968ec625b9de23436e6b2102d192b67

Observation 2aa37b1f-138d-46e6-93a4-5e1f82d646f4 · outbound

This paper cites Pit: Optimization of dynamic sparse deep learning models via permutation invariant transformation.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Pit: Optimization of dynamic sparse deep learning models via permutation invariant transformation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.824343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:af8981cd9137a0c8f03228146a8b6dee172dfb78a44c5691837ad6004aff1a15

Observation a867570d-452f-4a05-9915-f395e76c49b7 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Efficient memory management for large language model serving with pagedattention

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.799983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:c6bf0624100ce3b5a2d2606c48b201b2af79a8f23e96c01888e29ca2ee8e3fc2

Observation 2796d3b7-1526-4251-8dc9-035bb939000a · outbound

This paper cites YaRN: Efficient context window extension of large language models.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training YaRN: Efficient context window extension of large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.790405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:1fa068fda181885014d14bfa61369a281a3bf756147948b20a0373043aab3f50

Observation d618a994-714a-4444-a140-113e5d4776ce · outbound

This paper cites Reducing activation recomputation in large transformer models.Proceedings of Machine Learning and Systems, 5:341–353.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Reducing activation recomputation in large transformer models.Proceedings of Machine Learning and Systems, 5:341–353

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.805533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:99e01f6c79c251e549dd5fe43905a092e84422182598da52f5131703b89c1784

Observation e84b3ddf-c141-4043-b1c8-4c10a6c5de85 · outbound

This paper cites System optimizations for enabling training of extreme long sequence transformer models.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training System optimizations for enabling training of extreme long sequence transformer models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.759334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:e962eddf951874a64c2037b494524327ddabdc3ce792eaf32be7efc49fe9c4a8

Observation 1ca7e239-0078-40c4-b6b1-3d624cb28fb0 · outbound

This paper cites Mini-sequence transformers: Optimizing intermediate memory for long sequences training.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Mini-sequence transformers: Optimizing intermediate memory for long sequences training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.738305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:5b675a120c21f5dd84c3528dcdc275af86427dd92f06f78e1729f513aad22c64

Observation 9e833785-1c43-43b2-879b-0cf4c406a57b · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.604126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:db7913626f82533abf1937bdf3b3559d8fd494e9e0112a3277a1805d75196e64

Observation a492123d-5b69-401e-b410-6bbfaff96408 · outbound

This paper cites ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.608821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:3f2f1b9802a41b7647260f8e1b3a72792e5e3f57fd2d60595730a6172feb396b

Observation 4079208c-6b8f-41cd-8e90-5c2bfba5487f · outbound

This paper cites WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.687243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:d588035c634748457a9d5ade36f8c6b4689d50670668a814cb2da0fe03745ce6

Observation 07d888c3-aa41-4f36-a53d-1a6db1f2b7b8 · outbound

This paper cites Flexsp: Accelerating large language model training via flexible sequence parallelism.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Flexsp: Accelerating large language model training via flexible sequence parallelism

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.782393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:960a9affe8e391e908a19f4c9c8518789c47fd6bc2043d6a8929dbaffc5aeeb5

Observation 85f52a22-9f59-4f47-bb47-6d587fae16c4 · outbound

This paper cites Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T19:44:19.661086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:b06936a32c601beafa63e7cd410c43a17183f63d60b1df0bec60446dbe251b3a

Observation fb9b6527-eff2-4a36-b865-b0ae63916a20 · outbound

This paper cites Magiattention: A distributed attention towards linear scalability for ultra-long context, heterogeneous mask training.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Magiattention: A distributed attention towards linear scalability for ultra-long context, heterogeneous mask training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.819381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:59ad3406f4c2692af5b9db8fd01ee890f99e83a0d471612e191ecb2952c73369

Observation 9e51610f-379a-4ad8-9c12-778355d60636 · outbound

This paper cites LongroPE: Extending LLM context window beyond 2 million tokens.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training LongroPE: Extending LLM context window beyond 2 million tokens

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.802683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:77fce3c1d621671f7bb89a66176ec1173151e847d5dd8ab9b02a0553839e4223

Observation 43b0a93d-7881-4d41-81e9-e4939ed19ed7 · outbound

This paper cites A length-extrapolatable transformer.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training A length-extrapolatable transformer

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.805934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:a3f33b0b44e5ca8524d23a6f7882c1fc5ce024e81c8071c3761f65f42cd0e097

Observation 2d489ac2-221e-424d-92a4-a2b5c881d464 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Extending Context Window of Large Language Models via Positional Interpolation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.683897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:0c95df28c6f840c0a189cdc6f685d6a0286b0b2e8ad63189f16013041040e737

Observation 04eddd50-1c84-4e98-8eba-b86dbdd17d7d · outbound

This paper cites Training-free long-context scaling of large language models.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Training-free long-context scaling of large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.817334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:b94b3e4306b497eec0dd6ad73289bd0fdab63a0e682387926c3cf0f6f446f312

Observation 1226a374-0537-4db2-9916-087ad7edf812 · outbound

This paper cites Why does the effective context length of LLMs fall short? InThe Thirteenth International Conference on Learning Representations.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Why does the effective context length of LLMs fall short? InThe Thirteenth International Conference on Learning Representations

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.793674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:ff4bd06178f1f30c03aa45d3f11cfd6e8c67b2a50f9cf5d521af99539fb0b297

Observation 1657870b-8491-4ce4-8d88-0f92c0073135 · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.798957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:4aac4191ca13e9bffe42f1e6a2a805e867ba7e824222915f09647aedad89f423

Observation 17587785-99ab-439f-bd14-2e24dcfaa86f · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.750835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:2ac44711fa638bff2a7e17b6b025731b83ccabb7bc5a86730d1788dd82605fdc

Observation f37217c9-de4e-469f-8398-79b64e6c5e05 · outbound

This paper cites You only cache once: Decoder-decoder architectures for language models.Advances in Neural Information Processing Systems, 37:7339–7361.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training You only cache once: Decoder-decoder architectures for language models.Advances in Neural Information Processing Systems, 37:7339–7361

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.784456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:0d74d9395460fab7d0198877fb5829d30407dcd0ee63404b375c8e4a7067a8ba

Observation 1b7e8f8a-947e-478a-a1ca-5aab8f7c45cf · outbound

This paper cites GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.690254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:d4dc3dd28515b06179948f7596b9846b0a0705497979edcdf3994237f297dc6e

Observation 732e8330-7756-4c70-827f-fa8c363f6026 · outbound

This paper cites GQA: Training generalized multi-query transformer models from multi-head checkpoints.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training GQA: Training generalized multi-query transformer models from multi-head checkpoints

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.832231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:2ef797787a904eeef50d6d1827380529904ef285cf8018d80f55f0afca3417b5

Observation 6e481a43-a90c-4b04-a399-d51c2e9bab06 · outbound

This paper cites DHA: Learning decoupled-head attention from transformer checkpoints via adaptive heads fusion.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training DHA: Learning decoupled-head attention from transformer checkpoints via adaptive heads fusion

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.826744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:0400509e4d8e8a58660d7a745726bb144d6ffc9ea19654c57d7e1ba2f48787c5

Observation defbfae1-5b82-4ccf-8ac0-f3e207e1445f · outbound

This paper cites LLM maybe longLM: Selfextend LLM context window without tuning.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training LLM maybe longLM: Selfextend LLM context window without tuning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.776759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:86d4bf69b02ad3fa34548e124c23a594a3a6c5421dabdb8599641632adb9deca

Observation 5e2aee28-565d-4b14-ae5b-1d7657f00efd · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.829391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:3f2626dd21a52cae86e104d6ce63bb9e2dde96f4bd1590f72b4d23f9e602a1ee

Observation 9400a663-a1d5-455e-b551-eb84a272d2b1 · outbound

This paper cites Longformer: The Long-Document Transformer.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Longformer: The Long-Document Transformer

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.651616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:7c4a07d476766282e5642555a07bca2f6868e0979268c6dd67b8289c5720324e

Observation 48e379e1-52db-4310-be03-da0911e128f4 · outbound

This paper cites Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33:17283–17297.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33:17283–17297

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.791761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:1421eac9fec8648b2749da04e88a4096d9daffdf64157d2919b159bdc0b389e8

Observation c982ce7e-a165-476f-92fa-b86a024a784b · outbound

This paper cites Reformer: The Efficient Transformer.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Reformer: The Efficient Transformer

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.698753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:8c3d154ae948cc425d5c20efd1c95e5f3dd2cc53aa1ea817186bbda0c593158d

Observation de00f9e1-e847-4bb3-8830-214186114207 · outbound

This paper cites Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, and Lili Qiu.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, and Lili Qiu

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.813750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:8718eee719923b85dca4af32faec0a2b8221586fab9df2193b351447eee2b617

Observation fbee031b-c71d-4f29-81f0-8865182a58f8 · outbound

This paper cites Spargeattn: Accurate sparse attention accelerating any model inference.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Spargeattn: Accurate sparse attention accelerating any model inference

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.678313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:de3d26f49f6ab03fe6dc7f8a593b073002527e2e1f42b932511f0c8353e3ba78

Observation 652ae4b6-c055-46ab-8b27-c726766dbc2f · outbound

This paper cites MagicPIG: LSH sampling for efficient LLM generation.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training MagicPIG: LSH sampling for efficient LLM generation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:44:19.808308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:a39fc1329caefeea2f96a158dac4a32c53cf1a05c6329f5260fcd579e02b1b5e

Observation 0b81fa29-922c-456b-ab01-4c17a00644ac · outbound

This paper cites Adam: A Method for Stochastic Optimization.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Adam: A Method for Stochastic Optimization

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.667903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:69a99b3c537f67e1efc2a30c0fc16b67f08972e0693908443a4c6ecb0ab41fda

Observation f82f0226-62b5-4383-b62f-734d4564d5cd · outbound

This paper cites an unresolved cited work.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-21T19:44:19.824066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:d750a1fb851a9cc9aa2571b73574a3bfe1d4c1a2b8337da52b06376948883f81

Pith citing papers

Observation a8990c60-f268-4918-920b-394a6500e38c · inbound

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention cites this paper.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.513638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:4d93434e5ea8bcbab63f15c59098aec6d7e23177e8d57c7dc325aa52150028f3