Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T05:58:03.113220Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2602.10718.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T05:58:03.113220Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:17:09.834609Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-11T21:26:16.418821Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ff04e2b7-85ac-4e7e-834c-bf1ce329ac26 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d24bc5db-c3dd-4b34-b259-5aaf256ee9a0 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining DeepSeek-V3 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9de1798a-32a1-499c-b0bb-e358b6d830b7 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Large language models: a survey of their development, capabilities, and applications.Knowledge and Information Systems, 67(3):2967–3022
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2e41304e-2d6e-4cb2-9134-275e1cd61209 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e339cbd7-a8b8-4fcd-b620-23eb6c459b6d · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Beyondaime: Advancing math reasoning evaluation beyond high school olympiads
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation da9fac27-60df-462d-9c09-4df9712bedd9 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Kvquant: Towards 10 million context length llm inference with kv cache quantization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b2c2e177-53a5-465e-ba1f-1eb02c3c8643 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Introducing LongCat-flash-thinking: A technical report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d3cd48a7-b0a4-4fed-8418-f758da0887c4 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Longcat-flash technical report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb96c667-edaf-44f5-8f66-4620ccfbe884 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f4d28fa1-cf1e-4b17-a6f4-5dbde74f30ac · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Are We Done with MMLU?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1396979d-481e-4421-9d02-8e322cfc130d · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0c31b9a6-c1d5-433a-bc64-53ff67c91f19 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Measuring Massive Multitask Language Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 028ebe34-3e9b-4e94-8cd5-cc8794a6c64a · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Hmmt 2025
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a6e294eb-2d6f-4672-8ed3-0e82722d1d48 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Flashattention-3: Fast and accurate attention with asyn- chrony and low-precision
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b8e00ad6-46fa-4c02-8815-fdad04cf581f · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a8bdbf34-4915-4981-bc9b-9fbf9317e9ef · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Flashmla: Efficient multi-head latent attention kernels
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bd162b2c-62f9-4ba4-b5a4-eefdb6b46a4d · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Fp8 quantization: The power of the exponent.Advances in Neural Information Processing Systems, 35:14651–14662
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a10f8116-8746-4f62-9278-ea6d0cbfeaec · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining A Survey on Large Language Model Acceleration based on KV Cache Management
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ea2e2da1-6c11-417a-ab5c-00f035b25fcf · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining FPTQ: Fine-grained Post-Training Quantization for Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ffb43ce5-0f4b-4e5c-bb04-6d0b42bb1dfa · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining From live data to high-quality benchmarks: The arena-hard pipeline, April 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 146f8d88-3c75-485b-a414-89bae5b4080d · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Let’s verify step by step
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation afd98214-88a5-4ff9-bb3d-3162a0aff120 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Zebralogic: On the scaling limits of LLMs for logical reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 85a44771-7cbb-4e45-8c06-ce93a4e8cbe8 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b96e52ad-ce14-4acb-ae21-6e69344598f3 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c9283ccc-49a4-4505-9a65-f03811395bf1 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining & Cai, X
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 401b5094-ff2a-45d5-863a-4030a33e5ac6 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c87961f7-0b14-4ef1-a1e1-fabfdb25302b · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Aime 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6406654b-3f19-44e9-a286-6746a9842724 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Aime 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a8535f93-a846-44a7-8091-26e1744cc9b0 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Nvidia h100 tensor core gpu
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 28da0d75-5e81-4a14-899e-9f53e5f184c9 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining FP8 Formats for Deep Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 872555fd-866a-4059-a8c7-45b562dca162 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b21af5fd-a71a-4e57-a297-453a0d1f8cd8 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 738d38ad-17eb-4351-93ff-6eef95345c56 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Unveiling super experts in mixture-of-experts large language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 75706047-98c0-465f-8482-4571787ff875 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 941a7b1b-0002-4dab-a2d8-811c82e3db2e · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ffc5e7d6-2ead-4d3b-9349-b6a1fc54ba35 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Massive Activations in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 19de3124-d426-40e6-bb7f-00ec675e5497 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Longcat-flash-thinking-2601 technical report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8115f59e-59db-4dfd-a92f-2a0e67652181 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Longcat-flash-omni technical report
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2181953e-4a3c-40af-bae4-5cebd2f7a20d · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining FP8 versus INT8 for efficient deep learning inference
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ab9cbd9a-41aa-410b-9fb2-8beda59b1d44 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4b5203a8-14c2-436b-bcdc-0a352b55ec5c · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Exploring layer-wise information effectiveness for post-training quantization in small language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8aa1dece-69ae-4dec-b02c-f03492ae9dd6 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Dope: Denoising rotary position embedding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 818aac3d-4993-424f-b1dc-d899ec3280d3 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining ParallelComp: Parallel Long-Context Compressor for Length Extrapolation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e33aaf71-3a38-47ba-93a8-a839b1471656 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Fit and prune: Fast and training-free visual token pruning for multi-modal large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8c23886-9869-4f6f-8177-b2cb59ef9e6e · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cfd79c3e-131e-4582-9060-99f4ad02e374 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Efficient context scaling with longcat zigzag attention
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation db7c344a-33ed-4e21-bf4a-efeaf9d8d115 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c10a3a8d-699b-4888-b2ff-d0a95b7cebf2 · outbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Instruction-Following Evaluation for Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1f507c2f-d205-4357-bc1c-f7dd9b5705f1 · inbound
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c274f7cc-b4a6-4258-a008-fa485603c4dd · inbound
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.