Pith. sign in

Paper Citation Record · LEDGER

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining

As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2602.10718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.10718 v3

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T05:58:03.113220Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:17:09.834609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T21:26:16.418821Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact28
  • verified fuzzy14
  • unresolved1
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff04e2b7-85ac-4e7e-834c-bf1ce329ac26 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.740867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:ad8491d2b78f31a0d4742f893f73633ba377a328df93d17227e7cd674e5515aa

Observation d24bc5db-c3dd-4b34-b259-5aaf256ee9a0 · outbound

This paper cites DeepSeek-V3 Technical Report.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining DeepSeek-V3 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.737232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:aa92e91f60a1471514c901e1dd1fbe1bb7131742ade1807eabd3fa35b30a939a

Observation 9de1798a-32a1-499c-b0bb-e358b6d830b7 · outbound

This paper cites Large language models: a survey of their development, capabilities, and applications.Knowledge and Information Systems, 67(3):2967–3022.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Large language models: a survey of their development, capabilities, and applications.Knowledge and Information Systems, 67(3):2967–3022

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.412001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:ef4ad927c20e0be5af79cfbaafff139655b7844e0ee77367e99ab6bf26ba449a

Observation 2e41304e-2d6e-4cb2-9134-275e1cd61209 · outbound

This paper cites Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213– 100240

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.409792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:2e523fb7c56b3aa23df76e7a654fcfd934400b4889298c760ab710a755d4b8a6

Observation e339cbd7-a8b8-4fcd-b620-23eb6c459b6d · outbound

This paper cites Beyondaime: Advancing math reasoning evaluation beyond high school olympiads.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Beyondaime: Advancing math reasoning evaluation beyond high school olympiads

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.405607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:733bf7e57808f288d6b23e786481a3a55622a61418c9f81b4c4f1852db3d6c8a

Observation da9fac27-60df-462d-9c09-4df9712bedd9 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Kvquant: Towards 10 million context length llm inference with kv cache quantization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.407669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:197541c5b2fb0f0fdc699f259321b0c8e9d7bb424a2f0d808f90bbde769f7dd5

Observation b2c2e177-53a5-465e-ba1f-1eb02c3c8643 · outbound

This paper cites Introducing LongCat-flash-thinking: A technical report.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Introducing LongCat-flash-thinking: A technical report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.708301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:eb101fdb83fbd7ff75fafc46a6572450affc11d51a38f95baac5274fb00e1210

Observation d3cd48a7-b0a4-4fed-8418-f758da0887c4 · outbound

This paper cites Longcat-flash technical report.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Longcat-flash technical report

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.705131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:1a7ea0eb548966341927bd156dd25c879aa80c59882f39688a4552dc27f89404

Observation eb96c667-edaf-44f5-8f66-4620ccfbe884 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.807482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:eb93bb53b68016d253db44803ff53bf02bc83c72fca3ffd5ebadfa082ca4bf12

Observation f4d28fa1-cf1e-4b17-a6f4-5dbde74f30ac · outbound

This paper cites Are We Done with MMLU?.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Are We Done with MMLU?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.711571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:faeb81c16d3e69dfc6986311dfc4460df838c6179b257a50032feeeead381647

Observation 1396979d-481e-4421-9d02-8e322cfc130d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.733974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:f893735f55b3a408f457efc437e318e83247e9888404db137bde75525a725d39

Observation 0c31b9a6-c1d5-433a-bc64-53ff67c91f19 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Measuring Massive Multitask Language Understanding

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.724731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:ffbd16955ab85df0b80183a7b011d5d0d6a1383820de8110c75f9628c3da7294

Observation 028ebe34-3e9b-4e94-8cd5-cc8794a6c64a · outbound

This paper cites Hmmt 2025.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Hmmt 2025

Reference 13

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T06:00:41.442481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:338c595b13fe85da92bfa3170f94bf78785f893455621d6324e38541621d9e98

Observation a6e294eb-2d6f-4672-8ed3-0e82722d1d48 · outbound

This paper cites Flashattention-3: Fast and accurate attention with asyn- chrony and low-precision.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Flashattention-3: Fast and accurate attention with asyn- chrony and low-precision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.440548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:a50c8a656fc208d5565692f73141ea5555859b1540f1d223580ac14627988f5a

Observation b8e00ad6-46fa-4c02-8815-fdad04cf581f · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.438225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:29e1b429fd30b6051394c525ab4d879bd99b549e9e9da61fbb7097ff74d34f74

Observation a8bdbf34-4915-4981-bc9b-9fbf9317e9ef · outbound

This paper cites Flashmla: Efficient multi-head latent attention kernels.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Flashmla: Efficient multi-head latent attention kernels

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.435586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:4e9be659a0e0a06896658dbf84414a9bda4715eb76c58e2ebc1654cdbbba90f5

Observation bd162b2c-62f9-4ba4-b5a4-eefdb6b46a4d · outbound

This paper cites Fp8 quantization: The power of the exponent.Advances in Neural Information Processing Systems, 35:14651–14662.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Fp8 quantization: The power of the exponent.Advances in Neural Information Processing Systems, 35:14651–14662

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.432977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:d84c504047f67625bc67fd50416acf600f6047cebdf9c72131cf9a2dfc61cdd0

Observation a10f8116-8746-4f62-9278-ea6d0cbfeaec · outbound

This paper cites A Survey on Large Language Model Acceleration based on KV Cache Management.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining A Survey on Large Language Model Acceleration based on KV Cache Management

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.800812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:e9c1b0b9bc883bc859887672cb578f6f0489fbb40e8c436e8de54ceef77ac2c2

Observation ea2e2da1-6c11-417a-ab5c-00f035b25fcf · outbound

This paper cites FPTQ: Fine-grained Post-Training Quantization for Large Language Models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining FPTQ: Fine-grained Post-Training Quantization for Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.790664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:4f96e962da3502430ed7aea4037605f264ef958bd8294f6ac04e616d5eef2312

Observation ffb43ce5-0f4b-4e5c-bb04-6d0b42bb1dfa · outbound

This paper cites From live data to high-quality benchmarks: The arena-hard pipeline, April 2024.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining From live data to high-quality benchmarks: The arena-hard pipeline, April 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.430158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:b2b3e569a2234f639f3f3583333e932777dae6ccd05248ce888ae5f2f97b5d90

Observation 146f8d88-3c75-485b-a414-89bae5b4080d · outbound

This paper cites Let’s verify step by step.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Let’s verify step by step

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.427723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:83b90a973f0c24727274fcd3d93bd9233607304b45fce799082a8c9f6de2fd82

Observation afd98214-88a5-4ff9-bb3d-3162a0aff120 · outbound

This paper cites Zebralogic: On the scaling limits of LLMs for logical reasoning.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Zebralogic: On the scaling limits of LLMs for logical reasoning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.425299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:79e95b39a48a5c9567d8935478257133b985d9bb7c890ff147f2874305a452e7

Observation 85a44771-7cbb-4e45-8c06-ce93a4e8cbe8 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.422560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:1500863527cd822a9a2ad2c53976f32d6eff4f99f1da3dc0018fd0ca0e9178e5

Observation b96e52ad-ce14-4acb-ae21-6e69344598f3 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.787332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:cee6c1d9a3ec7982796197b19d0c046580253ad20b142276edb8dce65ad855ea

Observation c9283ccc-49a4-4505-9a65-f03811395bf1 · outbound

This paper cites & Cai, X.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining & Cai, X

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:00:40.727875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:9aa2f52de65ad238a2508ec261528594332d57d6cb8e72db297f9bd5133abea9

Observation 401b5094-ff2a-45d5-863a-4030a33e5ac6 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.714918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:e1cae6a80d83e59e443244609b853026a2a7e801981bfc103e320356d5e69cb3

Observation c87961f7-0b14-4ef1-a1e1-fabfdb25302b · outbound

This paper cites Aime 2024.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Aime 2024

Reference 27

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T06:00:41.420002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:c079dee25446532c4205bd0308cda6499a25bd7e3ac44d38538b84b5e03fcff7

Observation 6406654b-3f19-44e9-a286-6746a9842724 · outbound

This paper cites Aime 2025.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Aime 2025

Reference 28

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T06:00:41.418155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:054c3cf4d81240a03333d5ca24b1190da2182af6f3325d059ce36431b586f6a4

Observation a8535f93-a846-44a7-8091-26e1744cc9b0 · outbound

This paper cites Nvidia h100 tensor core gpu.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Nvidia h100 tensor core gpu

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.444489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:45798bc8dd9a00138805542bf04dc0b51c54249fa20e084ff021776d729c889a

Observation 28da0d75-5e81-4a14-899e-9f53e5f184c9 · outbound

This paper cites FP8 Formats for Deep Learning.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining FP8 Formats for Deep Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.780835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:a4a12dd3ff6db3ef0e32f1787be8558a12475e9ce2bd830d5b9093a8df49d3c7

Observation 872555fd-866a-4059-a8c7-45b562dca162 · outbound

This paper cites an unresolved cited work.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:00:41.416066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:6ad7bd715b437740319e3e124e836c1d016844052ee3ec88c14384fcd731d00a

Observation b21af5fd-a71a-4e57-a297-453a0d1f8cd8 · outbound

This paper cites Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.765255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:9c90e04d406b56b7924c4c602ddc201006fb6dc3003c8475956d08ba7e20cf3c

Observation 738d38ad-17eb-4351-93ff-6eef95345c56 · outbound

This paper cites Unveiling super experts in mixture-of-experts large language models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Unveiling super experts in mixture-of-experts large language models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.769142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:7fc7cbe437630ef2e135b84c87461434f453bfac801679ba27e6bbbe5b674c5c

Observation 75706047-98c0-465f-8482-4571787ff875 · outbound

This paper cites AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.761408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:3b588c87493e644e201be8f7ce8916fee67e060f6e74be304a4dec28e4d0b05e

Observation 941a7b1b-0002-4dab-a2d8-811c82e3db2e · outbound

This paper cites KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.756876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:9e2d529cf668f3b07c8fb12a7e1ef74646b2aad54b6912e6e14194e58722352c

Observation ffc5e7d6-2ead-4d3b-9349-b6a1fc54ba35 · outbound

This paper cites Massive Activations in Large Language Models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Massive Activations in Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:130054338bef68162fd4b2ce0a5e6911d37646c895f9af9381a5f20635bed01c

Observation 19de3124-d426-40e6-bb7f-00ec675e5497 · outbound

This paper cites Longcat-flash-thinking-2601 technical report.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Longcat-flash-thinking-2601 technical report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.772983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:d4c9cd2cf2fc7063098ba67794d44432ce6d4e1cf4c33abacd983e0c952dc745

Observation 8115f59e-59db-4dfd-a92f-2a0e67652181 · outbound

This paper cites Longcat-flash-omni technical report.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Longcat-flash-omni technical report

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.784126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:8c0614b6902a82b5782b80d8a18ce3c1c78352ba314cb3d09ed7830814092c4f

Observation 2181953e-4a3c-40af-bae4-5cebd2f7a20d · outbound

This paper cites FP8 versus INT8 for efficient deep learning inference.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining FP8 versus INT8 for efficient deep learning inference

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:00:40.804585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:c4f2be949cb78b366be5d0b17f87d752a3b50baf77489fa3573abfd7a56c6c62

Observation ab9cbd9a-41aa-410b-9fb2-8beda59b1d44 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.744485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:e072d4b303687c97c663af4e905ba353e3401c05815725edeee2aa77d4291098

Observation 4b5203a8-14c2-436b-bcdc-0a352b55ec5c · outbound

This paper cites Exploring layer-wise information effectiveness for post-training quantization in small language models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Exploring layer-wise information effectiveness for post-training quantization in small language models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.794189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:3d32716859d392bb2c903286df933e0a24439a848e9354bbe2d3b2d73c9f0156

Observation 8aa1dece-69ae-4dec-b02c-f03492ae9dd6 · outbound

This paper cites Dope: Denoising rotary position embedding.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Dope: Denoising rotary position embedding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.752336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:37bf0151548dd36227f9194d7ddadc35e22ab52e06abcf609df1504ef837f023

Observation 818aac3d-4993-424f-b1dc-d899ec3280d3 · outbound

This paper cites ParallelComp: Parallel Long-Context Compressor for Length Extrapolation.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining ParallelComp: Parallel Long-Context Compressor for Length Extrapolation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.721817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:630ca4ca8695c361effad947c6b2e5c584d3f6964a12b8bd41a785deda4b1e83

Observation e33aaf71-3a38-47ba-93a8-a839b1471656 · outbound

This paper cites Fit and prune: Fast and training-free visual token pruning for multi-modal large language models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Fit and prune: Fast and training-free visual token pruning for multi-modal large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:00:41.414152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:0dd20095834e873d2b01aa8c2fa235dd15edb0a8433bfb144a4a2f234ddd644b

Observation e8c23886-9869-4f6f-8177-b2cb59ef9e6e · outbound

This paper cites RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.748727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:68a9cac1b57bec642a77c8c3a5a387eff7055f97ce8ed7a1132c244cbead5906

Observation cfd79c3e-131e-4582-9060-99f4ad02e374 · outbound

This paper cites Efficient context scaling with longcat zigzag attention.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Efficient context scaling with longcat zigzag attention

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.797923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:548da1403e484bf2cf596d6481f79af5568384286093329262ee8c1c6a228ca7

Observation db7c344a-33ed-4e21-bf4a-efeaf9d8d115 · outbound

This paper cites Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.730979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:5f8c5cac1d23ae2e006e310caa87c4da1ae73fcb82ad4af77e35e00145b5d33d

Observation c10a3a8d-699b-4888-b2ff-d0a95b7cebf2 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Instruction-Following Evaluation for Large Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.718108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:2b7530290490e176a8520e95ffec806ee084c6fda0a8935236d20a156bed3b44

Pith citing papers

Observation 1f507c2f-d205-4357-bc1c-f7dd9b5705f1 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.161185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:9355cccfc71669ba548e12410deb93b27377a1c9e64a47707e17ef780f00533b

Observation c274f7cc-b4a6-4258-a008-fa485603c4dd · inbound

Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving cites this paper.

Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:26:16.431880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T05:37:02.419657Z digest=sha256:e451948ff715dc2c2443d3c8d05a79ff90492e792512663ba58e82a6b2f3df24