Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:50:59.403092Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2506.01969.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:50:59.403092Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T01:55:09.658053Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T01:57:51.249567Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 19bf4709-fd63-43f6-9afb-3c0118487b4e · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dc128f2d-17ac-40f6-a60a-bfb08fef5793 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs In: Burstein, J., Doran, C., Solorio, T
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8253e9a0-f5e6-491f-b181-7f8ea0794ded · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b3c921-e9d5-480f-9ad8-3a59d3f2d39a · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs In: Meila, M., Zhang, T
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4421f8a3-e26e-409a-b078-199aa00b7b5b · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1363dff7-f4a6-4965-9021-6bf35f7d3302 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd08d5fc-070f-45c5-8335-3e836b213fa2 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs DeepSeek-V3 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f9c228-b719-4657-b8a7-7e57a3684841 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f7bb8a33-a6f2-48b1-9709-46e1be3a8701 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087e3237-e088-4669-8367-26b8cc2ad5b6 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d9ae77-800f-42f4-a03d-3776e8d6ab72 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c6f7aa-3939-438c-8908-be6f87a4f08f · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 481c18d2-07bf-469b-b628-2209a5dc7896 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3cecf1e-9c74-4211-ae13-2f37a68c71a3 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f88aee9a-a2fd-4ed2-8b2a-a73c0bf564c5 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs In: Proceedings of the 29th Symposium on Operating Systems Principles
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529f7460-8570-4473-98ef-2638c834582f · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 037e9032-0e2f-4d4f-a080-7bde75c28608 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs Generating Long Sequences with Sparse Transformers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db7d9a0-7575-44a3-8414-150b50a49279 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs Longformer: The Long-Document Transformer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea0ecf8-2cd9-4218-bf73-3eacc49ff827 · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a86976-1127-4086-ae60-cfe1ee6c6f9e · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e1e6ce-0058-492a-96fe-d6a27d6702bd · outbound
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs RWKV: Reinventing RNNs for the Transformer Era
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668fa2d6-7176-4e49-ac2e-3529b01a39c7 · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 627a9baf-3ed6-436d-8de6-f76ea77c37e2 · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.