Pith. sign in

Paper Citation Record · LEDGER

MoBA: Mixture of Block Attention for Long-Context LLMs

As of 4 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 51 inbound Pith citation observations for arXiv:2502.13189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.13189 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T06:15:46.085555Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:02.987885Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact13
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch38

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12e2308a-217e-49c4-a647-8e6d0a944d20 · outbound

This paper cites Scaling Learning Algorithms Towards.

MoBA: Mixture of Block Attention for Long-Context LLMs Scaling Learning Algorithms Towards

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.322837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:8a11b83d7dca7e13ac35bbd1656982bef9ff5ac97555bc59736cab22db06b847

Observation 5792de7e-bc5b-4143-8d42-b9f68f66d312 · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

MoBA: Mixture of Block Attention for Long-Context LLMs and Osindero, Simon and Teh, Yee Whye , journal =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.300897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:789b9a7f0d43dd73aa3333409c5eb4a9cbbbbf9be76f80a69e0afc02f318d027

Observation 96afa291-3871-48f8-9a08-24a026a7469e · outbound

This paper cites 2016 , publisher=.

MoBA: Mixture of Block Attention for Long-Context LLMs 2016 , publisher=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.303270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:1ba64b5915d10dff0c62e454c491afa5df08f5173ca094441b8994fed29d3cc4

Observation 34cfb8e2-a36a-4b7a-87d6-af8e8e28c49c · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

MoBA: Mixture of Block Attention for Long-Context LLMs Generating Long Sequences with Sparse Transformers

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.210118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:136f107e810c4e6cd290e6764119fdfe6c34c9da97ad87f5429d558c354d50a9

Observation 2f741f92-27b3-4a70-a768-9cbfaa59e447 · outbound

This paper cites LongT5: Efficient Text-To-Text Transformer for Long Sequences.

MoBA: Mixture of Block Attention for Long-Context LLMs LongT5: Efficient Text-To-Text Transformer for Long Sequences

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.213645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:b7bee4d218cadc3cfb52e125e05052fc7b9520743e0c2dbe1249f8f4dcc84841

Observation 3b5588be-29d5-4f98-a344-922e38631ada · outbound

This paper cites Advances in neural information processing systems , volume=.

MoBA: Mixture of Block Attention for Long-Context LLMs Advances in neural information processing systems , volume=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.311157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:4c8c2af8d8ff369d96df23feb9433f91e863ca926ee419f8c24f15a65ab7b21f

Observation 514cd672-df6f-4d31-93dc-794f790a38cc · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

MoBA: Mixture of Block Attention for Long-Context LLMs LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.217139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:adb4c23aebdb6cf2f8e8a251e6e53a8d710a51471dd410d4f59eb13092c60040

Observation 9b2bbb8b-ab59-4153-8108-5f7111b2aac5 · outbound

This paper cites Axial Attention in Multidimensional Transformers.

MoBA: Mixture of Block Attention for Long-Context LLMs Axial Attention in Multidimensional Transformers

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.220433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:6ca98b52c6bd6f5ca45922b97ac10e86a0fdc509b47f6f12e00468b7a2357f86

Observation d2b1cde8-cb34-4e30-b1d1-25e47ff9ed5c · outbound

This paper cites 2024 , url =.

MoBA: Mixture of Block Attention for Long-Context LLMs 2024 , url =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.318192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:22fc50629ea9fbab0a74ef7c7f5d75c5bf74d6d15cfe9d3c4eb19f24edd6d1bb

Observation 6a879985-4ec8-4282-9e4b-f298c61796f2 · outbound

This paper cites Longformer: The Long-Document Transformer.

MoBA: Mixture of Block Attention for Long-Context LLMs Longformer: The Long-Document Transformer

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.223489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:5ce646e9245a50cd22e7f41c3dfed52f6cddf4ad002cdf2b2d10a4384d76cf6a

Observation ead3599f-f9cc-418f-9179-1c87a3e95737 · outbound

This paper cites ETC: Encoding Long and Structured Inputs in Transformers.

MoBA: Mixture of Block Attention for Long-Context LLMs ETC: Encoding Long and Structured Inputs in Transformers

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.226877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:be2c2eb855a72000889027dadbe37779bb0dba58e52eeaa9da448294b67d6d11

Observation 57751407-4cf9-44f8-80a4-c1e854fe1993 · outbound

This paper cites GMAT: Global Memory Augmentation for Transformers.

MoBA: Mixture of Block Attention for Long-Context LLMs GMAT: Global Memory Augmentation for Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.230102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:90ce639e348bd3add3f5dfefa53e437c02c030364cc178158b576449060dac69

Observation efa61b2d-2c40-4970-904f-55f59d316ff6 · outbound

This paper cites Star-Transformer.

MoBA: Mixture of Block Attention for Long-Context LLMs Star-Transformer

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.233414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:d7da8fcfb008e9ecabdfb7144a7ddfc672caa59fec235a2c70800ce47b36b3a1

Observation b25c3b8a-7bf2-427f-b295-d86fed6eab4a · outbound

This paper cites Blockwise Self-Attention for Long Document Understanding.

MoBA: Mixture of Block Attention for Long-Context LLMs Blockwise Self-Attention for Long Document Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.236730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:1838c17cfeb2b48ac96e4645fa9d023cb027ab0c4bbfb9c6ddf079628f4eb054

Observation d6d3db81-2cd1-499b-83b4-646e7fe5ceef · outbound

This paper cites Reformer: The Efficient Transformer.

MoBA: Mixture of Block Attention for Long-Context LLMs Reformer: The Efficient Transformer

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.240130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:62ea763644e98cb07b7404bd347a74eeab5b31833ff669f6adb055dbd3c8fd24

Observation cfa15bc7-0a08-4c1f-ac1f-bd161827775f · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

MoBA: Mixture of Block Attention for Long-Context LLMs Transactions of the Association for Computational Linguistics , volume=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.333948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:e405e68144f35aa15119b5478e7fb6050175210352907f5d3537d9ec2d54b905

Observation e9d33f26-8fdf-4aea-8ac7-b0d98759d246 · outbound

This paper cites Memorizing Transformers.

MoBA: Mixture of Block Attention for Long-Context LLMs Memorizing Transformers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.243639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:44c893d26a047d8908a7a86c1b44fbdbad3559c7126dfdbfa0b39b3827b7eab5

Observation 1a12abfc-82e2-4c6c-ab43-34d8fe56e6f9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

MoBA: Mixture of Block Attention for Long-Context LLMs Advances in Neural Information Processing Systems , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.338489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f7bfb108fe3c18521bab2d7aa3f427212deb6e8f513340c6e14c60fc582ca99f

Observation a07c9720-4fe3-42fb-b5fb-18333ad56c1e · outbound

This paper cites CoLT5: Faster Long-Range Transformers with Conditional Computation.

MoBA: Mixture of Block Attention for Long-Context LLMs CoLT5: Faster Long-Range Transformers with Conditional Computation

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.247140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:640e5b424b7bdd6de610fc06ad27f82cd192dfdf63f9c1643c2e2885001d0ba2

Observation 0375d885-57bf-477f-8c30-f5df2fc931df · outbound

This paper cites International Conference on Machine Learning , pages=.

MoBA: Mixture of Block Attention for Long-Context LLMs International Conference on Machine Learning , pages=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.343436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f419a14fd17fb902d849186f5937f3c2523ae5baf1776e4adc46ef877f1a853a

Observation 67719e0c-3f05-4aeb-b7a4-1b0c4d3a5174 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

MoBA: Mixture of Block Attention for Long-Context LLMs Advances in Neural Information Processing Systems , volume=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.346114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f7984ef95304d4325de941bdca6b0f158a83c8ad96ee90f24192deee9cdc5633

Observation f15e4b8f-8298-4f3f-b212-0ed879b3308c · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

MoBA: Mixture of Block Attention for Long-Context LLMs Efficient Streaming Language Models with Attention Sinks

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.250197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:5d137567761f23df764148a79325206583b7f3f7effb1e27ea8987d0ed5ee9b3

Observation 7277d24d-9352-4285-820c-0cb3f6496af7 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

MoBA: Mixture of Block Attention for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.253626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:9a87dc31c93ab7f8ca20f2b1e15300cd13c4658a431a287220643d945fd64f15

Observation 00ac9eb8-ae5a-4f83-9310-e27de39e5693 · outbound

This paper cites Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909.

MoBA: Mixture of Block Attention for Long-Context LLMs Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.256755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:3fe9fd2fe8dc63c823c6f0475e0dea00abf310910a6c74bade51f51df234e3fb

Observation 4bfcbe89-ca11-47ad-ac6f-e8349a2c8a76 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

MoBA: Mixture of Block Attention for Long-Context LLMs SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.260047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:a0af8a68ace309494fdebe93998b969171c42d832cb0022724bd87ee2317e48a

Observation 91644674-c33f-4c00-aba0-99523a188179 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

MoBA: Mixture of Block Attention for Long-Context LLMs Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T11:11:21.895994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:d2cf4ecdd7710fb69b8183d3a490b188bf2aecd917139c85a9e43c7dd3663e0d

Observation a7878e5d-c591-4aa5-8ce4-a30a81ee994a · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

MoBA: Mixture of Block Attention for Long-Context LLMs Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.266513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:a064ba387cc9852f0688889ee2c49d561c39c5736bd4a59a1726df3ffb2ccff3

Observation f8187306-dc17-460f-b73e-463d1aec41f5 · outbound

This paper cites Transformers are Multi-State RNNs.

MoBA: Mixture of Block Attention for Long-Context LLMs Transformers are Multi-State RNNs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.269719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f36cf49a0c88bb7ad2785b93bf37f612bb8bb2afa8266d2995a570217b2ba35f

Observation 00c7f58e-c130-4208-9fb8-98c99e972f56 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

MoBA: Mixture of Block Attention for Long-Context LLMs Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.272655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:964a9e101aac87a403818fa16c3cbdac3a3a9e66cafca21d43fecd980c305db7

Observation 39690b6a-cd0c-4abf-b951-7edd59aa7551 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

MoBA: Mixture of Block Attention for Long-Context LLMs RWKV: Reinventing RNNs for the Transformer Era

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.276017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f4e14c39457de259731cda70fbfbd0b6d7c9826bbe79d4e0ba17cafac5dadad8

Observation 532b35ac-d941-4e80-ae57-94dedef9d6f4 · outbound

This paper cites International Conference on Machine Learning , pages=.

MoBA: Mixture of Block Attention for Long-Context LLMs International Conference on Machine Learning , pages=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.327299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:07354cfce724adab0db95f75c0ba13eff81fcab58d31ef306820f78bdc46185a

Observation 6edff3ea-9471-49b3-83c6-ad6e3d32e8ee · outbound

This paper cites Rethinking Attention with Performers.

MoBA: Mixture of Block Attention for Long-Context LLMs Rethinking Attention with Performers

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.278924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:014343d81d72f2d586edc86f3817c70e21bf10315fba93ab4e25e8f933bb45c8

Observation 2cf07895-9ff1-4b48-a39b-d3f94ccd1aeb · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

MoBA: Mixture of Block Attention for Long-Context LLMs Linformer: Self-Attention with Linear Complexity

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.281866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:0cc1341b3667d7b300cdf14d5f59a4754bebba6fd08e72948483ac409004261f

Observation 8084c340-6425-4c01-8ea5-5bf795723cc1 · outbound

This paper cites International conference on machine learning , pages=.

MoBA: Mixture of Block Attention for Long-Context LLMs International conference on machine learning , pages=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.336297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:847599a67dd23dc460ef0fb9d2058de2db8a0d6d682ec0f1eeedd3a0a54e0406

Observation 5de5ed35-0373-454a-a6b3-dcf414336252 · outbound

This paper cites Efficient Transformers: A Survey.

MoBA: Mixture of Block Attention for Long-Context LLMs Efficient Transformers: A Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.285388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f87137149a27f9a94cac18928cb61896d6705e9690de165ab710ce8f3d84e430

Observation 981ac2b2-1042-4276-8a1f-d4e46aff7130 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

MoBA: Mixture of Block Attention for Long-Context LLMs Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.288484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:2f174dc1e291d453451aceb5417596769f9daaee509e85e13e33074d3e1d9e8f

Observation f2cac3bf-2d49-43ec-b3a8-2e358773441f · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

MoBA: Mixture of Block Attention for Long-Context LLMs DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.291664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:fc5470a923c28c9263ec8d50e773b03f5229d02d5b05774127fad87b5001c603

Observation ac5bd988-1512-42d5-b2d3-911fb07bb9f9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MoBA: Mixture of Block Attention for Long-Context LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.294698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:c3477038bb6ce1e09f14d7dee5272d78622580ee76d629f8adcff5e95bfa2e2e

Observation 301216a2-0a7d-46ef-a9f6-cbc5a762a757 · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

MoBA: Mixture of Block Attention for Long-Context LLMs USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.298224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:950ef62f460353f8df11b0af2673c69edc5b0f6e7762ce23ffbe10d0455081b9

Observation 3de63877-21a0-4fe7-b85d-29e2c110e9a0 · outbound

This paper cites Effective Long-Context Scaling of Foundation Models.

MoBA: Mixture of Block Attention for Long-Context LLMs Effective Long-Context Scaling of Foundation Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.122456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:ec8bcd4950887c2f83ac3af999ce97939dd805a07735df5a01e30ee160e0dea1

Observation 4c8bd24e-30ee-4361-9f20-a39bb6978cf8 · outbound

This paper cites Training Compute-Optimal Large Language Models.

MoBA: Mixture of Block Attention for Long-Context LLMs Training Compute-Optimal Large Language Models

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.127680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:6933e8600e33bbcf8a977a5633476eb3e4f181302c131e5d4e7e0c605c9a16af

Observation 5eb8880f-4595-4d0a-a18b-38c610de0259 · outbound

This paper cites Why Does the Effective Context Length of LLMs Fall Short?.

MoBA: Mixture of Block Attention for Long-Context LLMs Why Does the Effective Context Length of LLMs Fall Short?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.132633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:aa2f7c39002adda48e0526aa63a6e5d87d49afd7dacf3d6f94c3bb4ca5ee5a4d

Observation 7f7f9b04-8ec8-4d8a-998e-f1f76f9359ef · outbound

This paper cites Breaking the Softmax Bottleneck: A High-Rank RNN Language Model.

MoBA: Mixture of Block Attention for Long-Context LLMs Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.137946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:eeffc166e758f2bc612d40c82f8191ba207b6c38cf030d7e1ec13619bddff922

Observation 710b9ef6-14fc-48bd-a274-2ae6fb78c45d · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

MoBA: Mixture of Block Attention for Long-Context LLMs DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.142836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:9c431bcdcee290113ecff501617937ed97c202e3d544c65528d08a83c6858b83

Observation e79fe5f8-5b77-45b0-8c8e-162d08d2de25 · outbound

This paper cites Qwen2.5 Technical Report.

MoBA: Mixture of Block Attention for Long-Context LLMs Qwen2.5 Technical Report

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:15:46.147240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:8eb1f92e26df7af03e0dcdff1f065c4e3f1ddb52a15b125358df6d91551ff9b1

Observation 8667b04c-c3ce-40f6-ac36-06504cc9e4a0 · outbound

This paper cites 2023 , journal=.

MoBA: Mixture of Block Attention for Long-Context LLMs 2023 , journal=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.331669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:e5dfdbe9d8ff58c210c353d87c6a1b77b6298b34d582282efe059cb5dc41442a

Observation 8c9f3010-3668-49b1-8002-b70b6b3ef302 · outbound

This paper cites LongHeads: Multi-Head Attention is Secretly a Long Context Processor.

MoBA: Mixture of Block Attention for Long-Context LLMs LongHeads: Multi-Head Attention is Secretly a Long Context Processor

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.151606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:35a7b69eb5a4738bf342253f9a4917f7377b2ada3766480bb2e3117ecbd30840

Observation d2868931-587e-43ee-933d-eebd682c6fba · outbound

This paper cites LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation.

MoBA: Mixture of Block Attention for Long-Context LLMs LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:04:05.393377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:8609ea4a1cf4f9b42fcc664624f5684480fb91bb140efaa39e045cf15b0ec9d6

Observation 3f941a00-a54c-4ee6-979b-bb26acf1bf0d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MoBA: Mixture of Block Attention for Long-Context LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:15:46.159064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:c13ba9e82f2b4256eb1a87c62a6e29ec6972a7cfcd71567ddaac45bc2cbe30ce

Observation 44924d88-cb1c-4f33-938c-4cb66f2ec26c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MoBA: Mixture of Block Attention for Long-Context LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.162426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:c2e198a8e8d7c9ea0832dc509ff168bff620a1e4a4bd49e6a0edc5acaedf95ff

Observation 6ef9a4ad-f9b4-4eed-abe4-dde4c5d50bf1 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

MoBA: Mixture of Block Attention for Long-Context LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.165967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:31e2f89d22ebedeeb2c3ad9e20f0ebe077d8352ad178a0f8b3e901dfab7a537d

Observation 2b4ea7c9-1078-4610-abd1-e1717b026f56 · outbound

This paper cites Linearizing Large Language Models.

MoBA: Mixture of Block Attention for Long-Context LLMs Linearizing Large Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.169517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:36ec3c03f0637def1ea206bbb3b5dbd6900c9f448a650e530d270a925eee65a8

Observation 18c6b876-afea-4a2c-a198-9d8cb2d8336a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

MoBA: Mixture of Block Attention for Long-Context LLMs Advances in Neural Information Processing Systems , volume=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.315969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:58ae0109ceedd46e21dbee025b17805a837d6ab83dbcc3764828203d1fc1c7ac

Observation 89a7a419-4dd7-40de-9e05-ed267e17c94f · outbound

This paper cites The Mamba in the Llama: Distilling and Accelerating Hybrid Models.

MoBA: Mixture of Block Attention for Long-Context LLMs The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.173069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:090499de8e8f51628b053b9986af1aa073916a7bd5b528961f4d434136fb5aef

Observation 91c17f74-ab0e-4972-8bf7-96dfd231179d · outbound

This paper cites LoLCATs: On Low-Rank Linearizing of Large Language Models.

MoBA: Mixture of Block Attention for Long-Context LLMs LoLCATs: On Low-Rank Linearizing of Large Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.117218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:ff2e08db7240a8ffef6b49033c8f81b8004f02e4bc9a4d907ad7713414f61690

Observation 0e5c1054-ab0f-4330-84a0-b8c5a6461f63 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

MoBA: Mixture of Block Attention for Long-Context LLMs Retentive Network: A Successor to Transformer for Large Language Models

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.176867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:879bf60ca387ad9a2035c36c6e6620f742d0c154df94687b6025dffafd252539

Observation 4a8264cb-9a74-4e91-bb83-0174ae50ff0e · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

MoBA: Mixture of Block Attention for Long-Context LLMs Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.181106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:d83aa3bcc07852d40b93adb6baf5b12ea8659d54beb80a3389e65d86f61b4970

Observation 7d40d808-e022-46d4-9c11-43bcb4407866 · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

MoBA: Mixture of Block Attention for Long-Context LLMs Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T23:19:58.366890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:9aa1a23d9a890b2642988c18643da15dbda1682724999ea7ae6c1f2967db2857

Observation 70cbb10f-d834-44e6-917c-872bc1fcfe16 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

MoBA: Mixture of Block Attention for Long-Context LLMs MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:26:38.921226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:e9be2238c7d80272e235624dc0c0d44d62bad9fe196880ed292d4e93e1015f9e

Observation 3e10dd34-8e53-4d49-97a5-93813fe8839a · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

MoBA: Mixture of Block Attention for Long-Context LLMs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:3d584cb4f53512ffabd3cd64603046b06dcd6d0080b622cccdf474390a9d5011

Observation 1b3a2159-d31e-4435-bcc8-7f355f7a69b0 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MoBA: Mixture of Block Attention for Long-Context LLMs Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.196956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f593586097950d6c002b4230477692d1cba467e48bf9973f1c4b306c7ecbbe41

Observation f6a8de79-2a3a-4e1c-9872-cb16dad3c2b9 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MoBA: Mixture of Block Attention for Long-Context LLMs GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.200494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:9fe8924d5c926118f59714a27a803ede8b94a27bd65e00f3e4e74af77cd683e6

Observation c26987b1-df06-49d8-b5e1-144410e91d1e · outbound

This paper cites Journal of Machine Learning Research , volume=.

MoBA: Mixture of Block Attention for Long-Context LLMs Journal of Machine Learning Research , volume=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.313749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:082e3718c246e193f289d0083dffe6039d017f9a8222b429cc54583b8d23b09e

Observation 84a53930-ff56-42f4-82a0-c4d607a2c969 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

MoBA: Mixture of Block Attention for Long-Context LLMs ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.203629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:ec3230226fda351f374ab6a07b3938ec69141179c5b53154d9e2e7355ba8ced4

Observation 4a50dbf4-e967-4987-b4ba-51a2bc0e31ca · outbound

This paper cites 2018 , journal=.

MoBA: Mixture of Block Attention for Long-Context LLMs 2018 , journal=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.325086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:6e3b9eef0ede5a04bc16aa6f2b862db42e6dbc1fbe3f125f1070c4c6c84b3b9d

Observation 437eac4b-f066-4fd1-a985-d50434a82917 · outbound

This paper cites 2023 , journal=.

MoBA: Mixture of Block Attention for Long-Context LLMs 2023 , journal=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.329476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:d2dcfbdffe64870d733447e3a1a25a4b4f200540da7d3f8e618dae2b5628d5ff

Observation 0f0fb7a0-5a3a-4155-986e-3af46da9e583 · outbound

This paper cites NIPS , year=.

MoBA: Mixture of Block Attention for Long-Context LLMs NIPS , year=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.340993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:aa872180e105d5c611f546075f0aaded053dafc7401cf77e20c020b3d36eeff8

Observation 51cbc9c4-bd84-4efe-8bcb-6a04b4feca62 · outbound

This paper cites an unresolved cited work.

MoBA: Mixture of Block Attention for Long-Context LLMs Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:15:46.348702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:51c808dfd653af0bea2ee38dffe8bc87fa415a26253ef055bb834b3d0d65f9ad

Observation 6fcc58e4-9da8-495a-a43b-22e292ad3bd9 · outbound

This paper cites International conference on machine learning , pages=.

MoBA: Mixture of Block Attention for Long-Context LLMs International conference on machine learning , pages=

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.351291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:fa06bf53b58180b6f97f0593ab4b7df2dc062380beb814d0811d79794c6ed898

Observation a51c58be-91b1-4c7c-97ed-9cecd5b02c86 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

MoBA: Mixture of Block Attention for Long-Context LLMs Advances in Neural Information Processing Systems , volume=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.305836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:cec22c35620539396eef9f55aa5805a331e8ad83bce72ef07d85415d909129a8

Observation 002410d4-50e7-4c87-953f-205dc12df36f · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

MoBA: Mixture of Block Attention for Long-Context LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.206948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:3dce8876d587da37d06d2a46898c330918ae349c52a79013c118fd76e3d552af

Observation 5284c4b4-c7f4-439b-9a48-e4027324dd3a · outbound

This paper cites Cell , volume=.

MoBA: Mixture of Block Attention for Long-Context LLMs Cell , volume=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.320409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:43fca3adeb9b77c492e62372c28ea1d6b26fec119ebbb9115ae86afc4d688dc3

Observation 441115e9-38cb-435a-800a-0320414f36cc · outbound

This paper cites and Yang, Zhilin and Zhou, Xinyu and Zhang, Mingxing and Qiu, Jiezhong , title =.

MoBA: Mixture of Block Attention for Long-Context LLMs and Yang, Zhilin and Zhou, Xinyu and Zhang, Mingxing and Qiu, Jiezhong , title =

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:15:46.308534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:9a3a88d106fb4b8892a5100b963ed4d1d351fb9a545e7d8deb1fae98eb429171

Pith citing papers

Observation 91649ab1-c636-4a27-89a8-ad28dba2fafb · inbound

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention cites this paper.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:186becf8b3fdc23d9f7620d194b7fc7e138d3f2f408913e8795c6733fc123a04

Observation cbbca396-4ce9-41eb-bab1-80182d8c10ed · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:17:24.406028Z digest=sha256:cef109f98998c0d49323dfca1ea889af767cfef0fc1b6114ced509e170411a50

Observation 593aa73c-a482-4ac9-8a62-df9131d31232 · inbound

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference cites this paper.

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T22:06:52.086462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T22:03:10.316005Z digest=sha256:2d46bcf0a5f639dcb8e760cbbf019435aad21e7398f5d99dae3d78a04f480fe6

Observation 5ed79671-7dcb-4b84-950c-0656a590b46d · inbound

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference cites this paper.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.987885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.987885Z digest=sha256:38a46c9cbb491e9013c6082f9e58764f5799150eab6c9bf566cc24a3d5a881c4

Observation 4c8c4548-2e04-42be-bbb5-a61a684b5dce · inbound

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights cites this paper.

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:21:14.905213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:18:04.431436Z digest=sha256:1848fb12c89bf9da4cb0d9de4065a33b80c15624595d7a45389b33371456ab5f

Observation 52c40fb4-c89c-4206-aa23-b0b1a830503d · inbound

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training cites this paper.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:44:19.643653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:6606fcc2f11c7d9e123270d4c8d7e7746ff1e82b3a58c72e15a9e7342e97ddfd

Observation 2242e950-55d2-4337-8d75-0495dc20700b · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:43a1526b2fe4076d0a1adaa2b2ce1fbdeffdb6fe4ad3b7b14166c0548f0bbc81

Observation fce0a8bf-9db7-4480-88c1-caec5d6e0e5b · inbound

Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics cites this paper.

Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:08:39.758825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T23:06:27.535997Z digest=sha256:3d72b6ebc90fb2d999e2c6eb656ba2047fb01b6f8fc7526de79d5c5bec87641b

Observation 951f5481-db43-4982-99fb-16529f1273ca · inbound

BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential Recommendations cites this paper.

BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential Recommendations MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:26:42.182634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T07:20:46.517218Z digest=sha256:9c9981286861513ec6ce5412b55b1126887bba9cd4ca363fc9613e1ff979e09c

Observation c3fb043c-b8f9-4c19-90c8-534aa5a0ee3e · inbound

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling cites this paper.

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T15:46:09.903579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:46:09.903579Z digest=sha256:ab063b5ce1f7f2c9414c1bb0e15e5a15abecd676fbea28bf7a0990eba433a07f

Observation 9f28409e-6f4e-4de5-a9df-3a649259673c · inbound

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers cites this paper.

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T15:35:15.553731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:35:15.553731Z digest=sha256:ef12c5d3ea1f01d5723aa4c94ac61a49f4ebfe9fae509622b2058196ce9e5910

Observation 6ac34134-e03a-4309-825f-d9ee75fa499b · inbound

Mixture-of-Top-k Attention: Efficient Attention via Scalable Fast Weights cites this paper.

Mixture-of-Top-k Attention: Efficient Attention via Scalable Fast Weights MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:52:37.649073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:51:25.375847Z digest=sha256:d4d2b9344d3dfded3b8cf87ecbc1173a61a2a4a219f38bc6f1302f3b914a06d2

Observation 62b64b26-4dca-463c-9f29-07274483fd75 · inbound

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models cites this paper.

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:42.148739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:42.148739Z digest=sha256:c1edb27f8460fb4e4da37e09f7cb48c893e5a2217e1846e2a0c346f528225ecf

Observation 3dff839e-e01b-490d-8368-7674fa151a69 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:59:33.902420Z digest=sha256:5d7aca5cf255af02f763b2c5fad5102a0d4e154a3ffb11cd031e47d3ca801f3d

Observation 05a76c18-e14d-406b-8671-3280f3a5a4fe · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:50:09.458960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T12:45:27.150368Z digest=sha256:8f877ba371097b451ea58ac2d09fdbcf08039531f02cb48cb32dc09e75d5c2f8

Observation 703fce07-2eb4-422f-a5d2-54418911d808 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T22:05:43.025992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:05:43.025992Z digest=sha256:399ac1a5674336ca709bb104e00a30b620770e29ad13589c4848c30f6b137020

Observation 8b64b4c0-d99c-4bbc-8a2d-e46dfa70daa8 · inbound

Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions cites this paper.

Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T16:19:49.419805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:19:49.419805Z digest=sha256:e6abe29eef264ff14364675a24cd84018ce9c92856a5dc52eb2d8bf0d9b4b646

Observation db90e36b-b2db-409e-9a2d-ae47d675ee4b · inbound

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention cites this paper.

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:40:28.854082Z digest=sha256:df34e089317dd5897a381c4b7317bf454d68a53a6a37ca8f73f098e7558c260e

Observation e37bf27e-2c58-4303-a673-69373a4e434d · inbound

Why Attend to Everything? Focus is the Key cites this paper.

Why Attend to Everything? Focus is the Key MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T11:55:36.701888Z digest=sha256:6a45cbad745909a0e95f421fab22d98d7e2a7301179bacf47ab14f2141cd30c6

Observation 96fa5758-9f5b-41f7-a1cc-2deed830a7e0 · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:60618fbba62fc00f7bf919a790e908f179a87da08ad37a77b4644358fd87107c

Observation 70fc0503-e8dc-4fae-877f-dc822a24bd59 · inbound

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation cites this paper.

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:10:40.858525Z digest=sha256:b56f2d503e2f439ef0e85e9671bb3f53969de5105345b7ea6b0580337e5a2819

Observation 0d330da9-381f-4a9f-91e7-9705836464b4 · inbound

Simplified Sparse Attention via Gist Tokens cites this paper.

Simplified Sparse Attention via Gist Tokens MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T00:33:20.234142Z digest=sha256:c27d17c37b4f59cf3c233caa98d860e53247ba007e425b8ad4bb679662bb71e4

Observation eb1bedcb-bb67-422e-9551-47fcf22d9926 · inbound

Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation cites this paper.

Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T09:58:45.903396Z digest=sha256:7d61e03571b44efd7eed220c20470544cc35d5e522210c454a77efce7aab6881

Observation 46165425-c769-4cba-88ab-9f9a6f8c515d · inbound

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding cites this paper.

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T17:56:39.124969Z digest=sha256:cc8409a98db58969eb6e05afce63981e845972c399a6f4f31715ac2720d7c34e

Observation 80e7fe14-75c1-45c5-bdf0-f2ff75546026 · inbound

MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems cites this paper.

MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T04:29:57.548702Z digest=sha256:031adcbff5d452768473dd88670ee8bbdbf7b8747cce90ee769ccde16e87d117

Observation 59a50ae5-32e2-4822-9612-85474f9aebdd · inbound

UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification cites this paper.

UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T10:43:01.724760Z digest=sha256:4681f198b5528bdc18b1ea1ae6be702b82285e5575cfa6444aa5bc78e7732881

Observation 11ef081f-4157-4478-be15-0655e734f5cb · inbound

Long Context Pre-Training with Lighthouse Attention cites this paper.

Long Context Pre-Training with Lighthouse Attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T10:10:38.610613Z digest=sha256:9ae12c6e1d9c8b995a30af3a8b4f96576575a6477817c1c16aa7b74664ff5c20

Observation b0897f00-1ee1-4dca-8059-d00e17609658 · inbound

MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference cites this paper.

MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:27:55.991919Z digest=sha256:382535d75ac718d58f3082e30106033f146b839b9471c1831b9928c3f8b08f21

Observation a8495edb-0293-4b3b-b588-6a2420b73829 · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:09754cfe9fd6dbeea1e5ca5354047280a3ee396d565a76022ef9f734ee17e41d

Observation f66b317f-425c-486d-89b2-23f33c98abef · inbound

AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference cites this paper.

AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T04:42:26.433689Z digest=sha256:e4eefb7d8dd8c207b9e41872151a5d85db0d685630dac5b4a6a954d1c6a3b517

Observation 528c7a79-a5f3-4fd2-a759-949c7e27a743 · inbound

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives cites this paper.

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:38:21.587653Z digest=sha256:0d5325890aabdd7c9491081066a7fbe4fb3938a28ee38b3e0cf167585f418ea2

Observation 0fefc4be-0bbb-4834-9867-18a5d02daa44 · inbound

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices cites this paper.

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:43.898122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T20:00:27.987481Z digest=sha256:16273c9414899d73748aa2595a64901b82a25575470c88e59dd9d2b85e24a3b5

Observation ac3ca924-8a70-4938-bbff-cf5ff5feb2c8 · inbound

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection cites this paper.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.954338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:74f366bc434fa63cf64e585206c1046c5080776a26503f428fb566f274f2fea2

Observation 8d377bdb-ab8e-4879-bd48-443a7f4da7bb · inbound

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference cites this paper.

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:18:13.985129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:13:46.098095Z digest=sha256:7bfdef689769819fea572b1858fa3569b260decb29f4c913c02f7b63c76857a6

Observation a4afe736-77b8-4f36-9da0-83346051d342 · inbound

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference cites this paper.

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:03:59.891987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T22:03:07.685253Z digest=sha256:6bb644aaa1aba29f6465ee2926b646122f3f3240b0c14e1e770e8329960551ed

Observation 0707ab39-5de4-4e2a-8751-07a87169a7b3 · inbound

Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference cites this paper.

Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:53:31.442700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T14:44:08.335157Z digest=sha256:8442943aeb8b96dd4caae9bb32f83e6ec75b9b7284c0c7016007e0ff60d81550

Observation fa47025f-39c9-48de-9077-b6e51fe0753d · inbound

Veda: Scalable Video Diffusion via Distilled Sparse Attention cites this paper.

Veda: Scalable Video Diffusion via Distilled Sparse Attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:03:14.703191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T07:54:13.390911Z digest=sha256:03336fb2b3e4b0a5b32138b63fbc825aad079fa4f88267e5b17800cdad8add64

Observation 9d0a0992-787b-4667-b0e1-844344b5fcb3 · inbound

PithTrain: A Compact and Agent-Native MoE Training System cites this paper.

PithTrain: A Compact and Agent-Native MoE Training System MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:12:46.518850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T23:12:14.193089Z digest=sha256:35f4f341cbcf1ed91efc9824c96519d5878f430f4aca6c21493f25b3f2f976d1

Observation f82cc046-74f5-4be5-9c97-881ca97d28ce · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:36:59.493601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:e693477efd1146edb9563ef32ba92bedabc96fc597ad78eaab2c61992705bf98

Observation 98d16934-afb1-4c7a-97c6-94fa64e93098 · inbound

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing cites this paper.

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:36:59.573386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T01:06:04.896501Z digest=sha256:2522636be3bcead95abedc8d3e4bf3b5d2e1c8789d6f6a81fa62459c70fe7152

Observation 17684f85-2612-453e-9fd9-65cb7539bca8 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:47:59.831441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:ba34b3c164cdea16e2745db8e9dd3f203e5718363bf3e543df7a7fe1c7704733

Observation 10e7ae0c-ac28-4599-9b2c-29186ff03bae · inbound

ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers cites this paper.

ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:49:44.824110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T09:22:55.731241Z digest=sha256:2dce2ee1bc11f5b2605e5816f001021d2d58c12e21f0c053e1d0bbfa4322ad98

Observation cb04915f-b1b4-434e-950a-ce400a613889 · inbound

MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers cites this paper.

MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:14:17.946951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T06:11:23.742632Z digest=sha256:ee6a44023f87725dde9b2dfe6d8011879db48fd45410f8702e5d904f0822aaf8

Observation f0d533e4-affd-46cc-805b-fbec043a26c7 · inbound

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale cites this paper.

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.260804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T20:43:40.938898Z digest=sha256:4dcdddc5e6fd3dc20f7d4752cc4eb38918c619fbb61ae8c62102bee9b72b5240

Observation 5cb92a38-82cb-45da-9c9c-c78b1ab8dd38 · inbound

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling cites this paper.

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-07-12T05:40:56.557646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:40:56.557646Z digest=sha256:b0abcd297639d0429f347af70e8a22643daba91fd90500ac5db83d1c2562f8f2

Observation 60420798-09bf-424f-b3cc-063e7a00f118 · inbound

Uncertainty-gated selection for block-sparse attention cites this paper.

Uncertainty-gated selection for block-sparse attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T22:15:14.580916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T22:15:14.580916Z digest=sha256:caac75f4a8ee23003dc95f70cb323d8d00da52159a4262012acdf79f65f1f93a

Observation 261ad5a0-1847-4a54-a4c9-8b71fb2d2987 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T01:36:44.216848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:3aab56a969b0b5ec1cae78a245f6c93407f1fc91c95fff1fa855be1be3cfd1b5

Observation 4dd80a52-b285-43e4-b4ad-8148cfacc908 · inbound

Learning What Matters: Supervising Global Context Pruning with Causal Evidence Sets cites this paper.

Learning What Matters: Supervising Global Context Pruning with Causal Evidence Sets MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T07:14:36.975183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:14:36.975183Z digest=sha256:9b9ba36e4260c8882e4bbc6f65f12af46efbe20942712f3a2dfb3383b60ef4d6

Observation 78245182-d1d3-4e19-9e05-613aed85bdd3 · inbound

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding cites this paper.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.707395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.707395Z digest=sha256:0cba24434b039ea51d71c4b2ce14112a1a658e38a5cb9ea4eec123ffcf59b8cc

Observation 59aba606-4570-4574-8eb9-a9b4061065c8 · inbound

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention cites this paper.

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T11:04:04.532117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:04:04.532117Z digest=sha256:2081a92dbfe269ed2b5d766680d83ab5ecdcbce33c58fa29e96881b366ad0a48

Observation 5f33fcab-a341-4953-ae6e-5a943bd895d5 · inbound

Memory for Large Language Models cites this paper.

Memory for Large Language Models MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T02:37:54.207074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:37:54.207074Z digest=sha256:203fafb810db97bb39b7848c39bde7db395127019403d95c758ac5a4fd11ffc8