Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T19:44:04.833504Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2510.18830.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T19:44:04.833504Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T10:50:12.926232Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-20T10:53:13.511836Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5156b2c3-1e9e-49dc-8745-21a4ec1cc503 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8b59d568-3960-45b1-a816-be2561950209 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11946e18-9bf2-405b-92bb-d7f9f5539dc8 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af577f49-b0c7-482a-beaf-219c289e88cf · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b1a1cc1-1fed-4755-a618-e6a580edc27d · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Introducing deep research
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 501e3336-eb8a-4efd-ba3f-236e59fe70fc · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Leave it to manus.https://manus.im/
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3967738b-0bcf-4abc-b923-f82fb79bbcf5 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c2098a1-d92b-47a9-9f8a-72f3df8d1add · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Qwen2 Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 078993bc-e5d7-4298-ae2b-49d7427f839d · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training DeepSeek-V3 Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 365dd40c-f9b7-4744-b502-aafdfd5e2728 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Qwen2.5-1M Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c66bbf04-33bf-48a5-887a-ef0cbac2436b · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e1ac85b-cf30-4fff-a625-a69ab91f625e · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training How to train long- context language models (effectively)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db7b691a-988a-4892-99b9-015b8950cfb1 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training QUEST: Query-aware sparsity for efficient long-context LLM inference
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e3d1aef4-6538-43b8-a3c4-70a452dda11b · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 99781537-8690-4c73-b4ae-59911756f801 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Flexprefill: A context- aware sparse attention mechanism for efficient long-sequence inference
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2530f5e5-728d-42bd-84b2-79dab862de57 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Sparq attention: Bandwidth-efficient LLM inference
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7620b104-d95d-4ebb-8ee0-b469f1ffc6c4 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 52c40fb4-c89c-4206-aa23-b0b1a830503d · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0bedc2a2-31f4-41fe-8952-7f1b8aa92efb · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Ring attention with blockwise transformers for near-infinite context
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d419d5d6-aaa8-4dda-a17d-4368d07e05c0 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Striped Attention: Faster Ring Attention for Causal Transformers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ea97251-5024-4118-89b3-d8afe3ffeca7 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Qwen2.5 Technical Report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f620287-2ec2-4a1d-a770-223e185746ac · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c7d8006-3b1c-4c2d-bbb3-47c47c934954 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Needle in a haystack - pressure testing llms
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b4744c0a-c978-43a1-9840-5f7f976ccc23 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Infinitebench: Extending long context evaluation beyond 100K tokens
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9acf4671-7194-4cb0-a65e-e8227b7d8ab9 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Compressive Transformers for Long-Range Sequence Modelling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f2178f6-c1cc-4d47-819c-b161ceb74655 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344–16359
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e350ce1f-9611-4814-8d8a-287d81cd19c2 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training [Feature request] balancing computation with zigzag blocking
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 123ba3fd-bb6b-49f8-ae64-83c38e133c4c · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training XAttention: Block Sparse Attention with Antidiagonal Scoring
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6d17fc7-3803-45a0-98c1-4b4fb9709656 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cb32a3f-1115-4ce4-ab42-b1adb59684f9 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e24abef-8f4c-4ab8-93b9-45eab6506218 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training {nnScaler}:{Constraint-Guided} parallelization plan generation for deep learning training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ccb14f7-86a5-4fd3-92ec-5bebc232e287 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Zero: Memory optimiza- tions toward training trillion parameter models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2dd9909d-2552-46a2-981d-bc4b8af800c1 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Gpipe: Efficient training of giant neural networks using pipeline parallelism.Advances in neural information processing systems, 32
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9088f2ef-f293-4b82-a8bc-c10788acca88 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Training Deep Nets with Sublinear Memory Cost
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5154ded-68f1-4877-831e-01769cf4a23f · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Block Sparse Attention
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2aa37b1f-138d-46e6-93a4-5e1f82d646f4 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Pit: Optimization of dynamic sparse deep learning models via permutation invariant transformation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a867570d-452f-4a05-9915-f395e76c49b7 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Efficient memory management for large language model serving with pagedattention
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2796d3b7-1526-4251-8dc9-035bb939000a · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training YaRN: Efficient context window extension of large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d618a994-714a-4444-a140-113e5d4776ce · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Reducing activation recomputation in large transformer models.Proceedings of Machine Learning and Systems, 5:341–353
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e84b3ddf-c141-4043-b1c8-4c10a6c5de85 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training System optimizations for enabling training of extreme long sequence transformer models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1ca7e239-0078-40c4-b6b1-3d624cb28fb0 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Mini-sequence transformers: Optimizing intermediate memory for long sequences training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e833785-1c43-43b2-879b-0cf4c406a57b · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a492123d-5b69-401e-b410-6bbfaff96408 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4079208c-6b8f-41cd-8e90-5c2bfba5487f · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07d888c3-aa41-4f36-a53d-1a6db1f2b7b8 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Flexsp: Accelerating large language model training via flexible sequence parallelism
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85f52a22-9f59-4f47-bb47-6d587fae16c4 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb9b6527-eff2-4a36-b865-b0ae63916a20 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Magiattention: A distributed attention towards linear scalability for ultra-long context, heterogeneous mask training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e51610f-379a-4ad8-9c12-778355d60636 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training LongroPE: Extending LLM context window beyond 2 million tokens
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43b0a93d-7881-4d41-81e9-e4939ed19ed7 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training A length-extrapolatable transformer
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d489ac2-221e-424d-92a4-a2b5c881d464 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Extending Context Window of Large Language Models via Positional Interpolation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04eddd50-1c84-4e98-8eba-b86dbdd17d7d · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Training-free long-context scaling of large language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1226a374-0537-4db2-9916-087ad7edf812 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Why does the effective context length of LLMs fall short? InThe Thirteenth International Conference on Learning Representations
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1657870b-8491-4ce4-8d88-0f92c0073135 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training KIVI: A tuning-free asymmetric 2bit quantization for KV cache
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 17587785-99ab-439f-bd14-2e24dcfaa86f · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f37217c9-de4e-469f-8398-79b64e6c5e05 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training You only cache once: Decoder-decoder architectures for language models.Advances in Neural Information Processing Systems, 37:7339–7361
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b7e8f8a-947e-478a-a1ca-5aab8f7c45cf · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 732e8330-7756-4c70-827f-fa8c363f6026 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training GQA: Training generalized multi-query transformer models from multi-head checkpoints
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e481a43-a90c-4b04-a399-d51c2e9bab06 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training DHA: Learning decoupled-head attention from transformer checkpoints via adaptive heads fusion
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation defbfae1-5b82-4ccf-8ac0-f3e207e1445f · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training LLM maybe longLM: Selfextend LLM context window without tuning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e2aee28-565d-4b14-ae5b-1d7657f00efd · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9400a663-a1d5-455e-b551-eb84a272d2b1 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Longformer: The Long-Document Transformer
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48e379e1-52db-4310-be03-da0911e128f4 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33:17283–17297
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c982ce7e-a165-476f-92fa-b86a024a784b · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Reformer: The Efficient Transformer
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de00f9e1-e847-4bb3-8830-214186114207 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, and Lili Qiu
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fbee031b-c71d-4f29-81f0-8865182a58f8 · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Spargeattn: Accurate sparse attention accelerating any model inference
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 652ae4b6-c055-46ab-8b27-c726766dbc2f · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training MagicPIG: LSH sampling for efficient LLM generation
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b81fa29-922c-456b-ab01-4c17a00644ac · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Adam: A Method for Stochastic Optimization
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f82f0226-62b5-4383-b62f-734d4564d5cd · outbound
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8990c60-f268-4918-920b-394a6500e38c · inbound
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.