Pith. sign in

Paper Citation Record · LEDGER

Hardware-Efficient Attention for Fast Decoding

As of 22 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 5 inbound Pith citation observations for arXiv:2505.21487.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21487 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:32:35.413009Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:48:01.180107Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.329269Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved69
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43121bc2-b3cf-48bd-b5ec-d989455c453e · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Hardware-Efficient Attention for Fast Decoding SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.014918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.014918Z digest=sha256:253a5a36b2a419f4a95ba52ae80ef0c711a3ec89d1e6b22643637e03b685f899

Observation 41699e1c-1716-4d8f-90e2-d7bf64aa9f0f · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Hardware-Efficient Attention for Fast Decoding GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.108759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.108759Z digest=sha256:12cad33f2d495b9aba7fe60025735e89469e66599281cd97615087457f244457

Observation c8abc7c3-bda3-4139-8f81-e0cb39023140 · outbound

This paper cites DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.

Hardware-Efficient Attention for Fast Decoding DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.198408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.198408Z digest=sha256:b312e5edf6b3416b5e1715abb217009b9c9e00b31b068bdf14071151c1959872

Observation 71edab39-71ef-451a-b4ee-82e3c796a122 · outbound

This paper cites How to scale your model.

Hardware-Efficient Attention for Fast Decoding How to scale your model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.314397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.314397Z digest=sha256:3766de734953e28b7930abb0d2014093fd4fb7714caaeb582b10d1f20f579489

Observation 006a6026-b703-4b2a-9bb7-af46222eabbe · outbound

This paper cites Round and Round We Go! What makes Rotary Positional Encodings useful?.

Hardware-Efficient Attention for Fast Decoding Round and Round We Go! What makes Rotary Positional Encodings useful?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.405631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.405631Z digest=sha256:34ae9ba433f096557b45790d074eb155106e599220ff1aad4c5ef90086aa3116

Observation 0705c79f-e322-49e3-8af2-b710e8b0cbde · outbound

This paper cites Singe: Leveraging warp specialization for high performance on gpus.

Hardware-Efficient Attention for Fast Decoding Singe: Leveraging warp specialization for high performance on gpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.516253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.516253Z digest=sha256:b29ed26363edad0368691ef25b2f9b3beef35d45b52dfd86e440e92dfcdffa8d

Observation 7c62b84a-73c0-4045-b3a2-e4789ca473d8 · outbound

This paper cites Cosmopedia, February 2024.

Hardware-Efficient Attention for Fast Decoding Cosmopedia, February 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.606445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.606445Z digest=sha256:d03d292e8bf94e41fb0d9ced64dee09c64c4cc4912ee59eef81ca33db13f27d8

Observation 52226de7-e862-4371-ba00-cd0cd4de1039 · outbound

This paper cites PIQA : reasoning about physical commonsense in natural language.

Hardware-Efficient Attention for Fast Decoding PIQA : reasoning about physical commonsense in natural language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.689226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.689226Z digest=sha256:3c2131d315210c7c3150ca1bcd5b1fdca0b43168e209dd7bf02c653e76cf40d0

Observation 53b89d84-b0e0-4fb9-b254-9faf7d4b55d3 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

Hardware-Efficient Attention for Fast Decoding GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.758624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.758624Z digest=sha256:573b4c69056e89a602b2d0f202b29d68cd6362f4c08b84654d96d95d757ab396

Observation cd97bfa7-81dd-4579-b112-b259b2a85479 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

Hardware-Efficient Attention for Fast Decoding Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.855424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.855424Z digest=sha256:2b36d024a39ed9a01d519465a907da89ec69c5cb7f285b4dfcc59ec90e321095

Observation e6dc91aa-aca2-434d-bc3b-409bccdb244c · outbound

This paper cites Language Models are Few-Shot Learners.

Hardware-Efficient Attention for Fast Decoding Language Models are Few-Shot Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.964527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.964527Z digest=sha256:6daecd309860adb42ae2a9d65bf299f38897e95226700eb1e6b96d118a3e438c

Observation 7e5ea5fb-a49e-40fb-b73a-2d4c9cb73d6a · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

Hardware-Efficient Attention for Fast Decoding Palu: Compressing KV-Cache with Low-Rank Projection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.031025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.031025Z digest=sha256:300c3eb079dfbb363306d584cd7bf657b1590065a8795b3977980fa030c0bf67

Observation 6bd2ed83-9ca6-45b0-9e61-9a775d604bad · outbound

This paper cites What rotary position embedding can tell us: Identifying query and key weights corresponding to basic syntactic or high-level semantic information.

Hardware-Efficient Attention for Fast Decoding What rotary position embedding can tell us: Identifying query and key weights corresponding to basic syntactic or high-level semantic information

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.124285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.124285Z digest=sha256:f4e4969326b951cb15f4aa0fb864cdeca4a137fd2735b1e952b4ac0a39e261f9

Observation ab1256f6-b547-4074-af20-9c032096d05d · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Hardware-Efficient Attention for Fast Decoding MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.213422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.213422Z digest=sha256:937a507753843bd7d186ad0e481da9a229b32b40f5995ed6a8fdca40ef98b4d9

Observation be7ee620-6b3d-4922-8314-c5a0505f2535 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Hardware-Efficient Attention for Fast Decoding FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.282299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.282299Z digest=sha256:4129e88b6eae24635b018ae57aa87af4e7b6a4ca4f455c18635e3bbeb0d4cecc

Observation 95c9086a-1c49-4563-97dd-31c451a00e4d · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Hardware-Efficient Attention for Fast Decoding Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.361518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.361518Z digest=sha256:9822299f3c61f79538d97c2e3c81cfdf5813fe2fed6b8217fbbb8f5cea4ca048

Observation a3c6219a-08ae-40e7-996d-9744c11d397f · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Hardware-Efficient Attention for Fast Decoding FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.427540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.427540Z digest=sha256:99cbb640d022554d585edf6d061617b86ecf7571f5364391221e12dad3a3ff6f

Observation d7970ad1-5764-47cf-ab55-18c67e2fed42 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Hardware-Efficient Attention for Fast Decoding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.499057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.499057Z digest=sha256:6b79145410de9bed1b25133c65f068564ac8a4eb08a9867ae99a07f44c1483f3

Observation 67e5787e-a7a4-4627-a326-c57ad1e02fa2 · outbound

This paper cites DeepSeek-V3 Technical Report.

Hardware-Efficient Attention for Fast Decoding DeepSeek-V3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.545757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.545757Z digest=sha256:1731af0bf8ec8fc0a2dfc484c761094145454a5e1bb97408203700f8e5834d68

Observation 7968316b-9305-4349-8005-76947f67185c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Hardware-Efficient Attention for Fast Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.660022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.660022Z digest=sha256:bc5b1fc875ed9b138532ac48965637ad06315860535aebd372ea9b730602021f

Observation 2b8ea07b-7e80-4384-8a62-49119f32b53c · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Hardware-Efficient Attention for Fast Decoding The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.694921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.694921Z digest=sha256:02d74954d0200dee8b4306b4ec2dd2c2d6d1f2b4b8aa37a5e7316116f4858591

Observation 139f2bc8-0929-4599-b437-cfbf0a2f0cfd · outbound

This paper cites AI and Memory Wall.

Hardware-Efficient Attention for Fast Decoding AI and Memory Wall

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.729115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.729115Z digest=sha256:0d67ae0ee7dc309b5652e5ee3d94ae2542046689ed6121d0f7964094e11cf184

Observation ff636470-7567-4fcc-aacd-2e29786a2d98 · outbound

This paper cites What Your DRAM Power Models Are Not Telling You: Lessons from a Detailed Experimental Study.

Hardware-Efficient Attention for Fast Decoding What Your DRAM Power Models Are Not Telling You: Lessons from a Detailed Experimental Study

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:32:36.701141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:30.766151Z digest=sha256:35277cbacecae48a77549e2c7a3c9c34c23e9908e71d89aaa7c52b95d84abc85

Observation d741e4a5-9158-4ac3-a163-4c04e4e0f317 · outbound

This paper cites Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA.

Hardware-Efficient Attention for Fast Decoding Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.813413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.813413Z digest=sha256:a7419f5b0546c9c82e8ee4d6629800ecc5927683c8922405d63317af9c948d96

Observation 0c615176-c5f5-4637-b98a-a7516149c1d0 · outbound

This paper cites The Llama 3 Herd of Models.

Hardware-Efficient Attention for Fast Decoding The Llama 3 Herd of Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.846409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.846409Z digest=sha256:0164239d64c1dfa1bfbda298da5f8a40d69f100710b4819c448a7a2d1b64b296

Observation 2955cf70-99a0-4063-95b3-5f8086e443cc · outbound

This paper cites an unresolved cited work.

Hardware-Efficient Attention for Fast Decoding Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:32:39.162151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:30.889776Z digest=sha256:c41e7a2f640716c4bcfc0b73bbb097d57fb69f1ea5a617106e471b0548fe473f

Observation 9f7dd2d3-10c8-40d3-a70c-dec91328cd2a · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Hardware-Efficient Attention for Fast Decoding Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.925220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.925220Z digest=sha256:3d4b294c47f4c622d74eba0f070486d085f0bd836fb3e47a6719d17eb2e518fc

Observation f576932a-1cf8-4630-82db-d0af173eecb0 · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

Hardware-Efficient Attention for Fast Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.932279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.932279Z digest=sha256:2fa5fb9de5cf16d3cb2fffa38d04a75ff25ac2e56cfe460cbfa8d389ef267a72

Observation 944ab807-ae86-4bad-a3b8-b1eb23e8c927 · outbound

This paper cites Measuring massive multitask language understanding.

Hardware-Efficient Attention for Fast Decoding Measuring massive multitask language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:38.970755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:30.938181Z digest=sha256:0481d8887238fdeb46b8c651fe8cca5ae75b3faef3274b9d506102f151fd8c3f

Observation b846c7a3-8c26-4820-9f30-fd34b7596dbf · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Hardware-Efficient Attention for Fast Decoding KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.954664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.954664Z digest=sha256:2b3a140733689f9d574190f2f981641d68d73c34daa1006b400d289f3f32cb48

Observation 52e12996-15d3-485f-bf64-f646d8edfe21 · outbound

This paper cites Multi-matrix Factorization Attention.

Hardware-Efficient Attention for Fast Decoding Multi-matrix Factorization Attention

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:32:36.488542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:30.968649Z digest=sha256:c5bbad41b1f47ba6589c38cf9ffb4100a6ef3e6269c9c303028df417df73939e

Observation 1be02b3e-8259-4ea6-94a7-4198e332a6e4 · outbound

This paper cites Data Movement Is All You Need: A Case Study on Optimizing Transformers.

Hardware-Efficient Attention for Fast Decoding Data Movement Is All You Need: A Case Study on Optimizing Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.978625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.978625Z digest=sha256:9ed1b3831d546d498667d40a64f6cfc0d77fa68b6bbe66168204a59beed9d2a7

Observation 2bfcb9ad-1b43-4722-8b3f-0c50f89efe30 · outbound

This paper cites Towards economical inference: Enabling deepseek's multi-head latent attention in any transformer-based llms, 2025.

Hardware-Efficient Attention for Fast Decoding Towards economical inference: Enabling deepseek's multi-head latent attention in any transformer-based llms, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.994451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.994451Z digest=sha256:9b971cc9a45cf8828bcc6aed7f914916fbcdd1e7da4e1f781b29fc9c0d3060fc

Observation e5f0562a-5d2d-4c4f-aa66-386544c71ac7 · outbound

This paper cites Weight decay induces low-rank attention layers.

Hardware-Efficient Attention for Fast Decoding Weight decay induces low-rank attention layers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.007337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.007337Z digest=sha256:730abe716a1fbf7abcd0e4365e2ffc20a6781f5b3258aa6af662287f27729c34

Observation 661446ef-c137-4989-9696-e2e1da24b08f · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Hardware-Efficient Attention for Fast Decoding Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.026027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.026027Z digest=sha256:5070eb610e8da266472c2d0601db166fe086197e581dde0721464a0313ef87b0

Observation 21995711-3da1-4c6a-8a20-0e9cfaadaa91 · outbound

This paper cites Software pipelining: An effective scheduling technique for vliw machines.

Hardware-Efficient Attention for Fast Decoding Software pipelining: An effective scheduling technique for vliw machines

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:38.784872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:31.043180Z digest=sha256:ac8e7ba1dbc9507f8c4a6c8678c3cdbf27d84c0243252e9c4237697a10a5f03d

Observation f71a8029-0f86-4186-8e08-25c87f10d93e · outbound

This paper cites Flashmla: Efficient mla decoding kernels.

Hardware-Efficient Attention for Fast Decoding Flashmla: Efficient mla decoding kernels

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:38.627480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:31.061293Z digest=sha256:f98cf18a4340955ef9a17891f8380ae3ed0f53de36fec1849c21b641e8bf2e51

Observation eff9c8b8-d14c-434b-b783-8a38e6bee5b5 · outbound

This paper cites PyTorch Distributed: Experiences on Accelerating Data Parallel Training.

Hardware-Efficient Attention for Fast Decoding PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.080828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.080828Z digest=sha256:42e3a19daa841b437643da86839f64f77b1722cb17c666267a1e73881bd33fdd

Observation 8e4c7db0-b686-4cbb-8373-b9af37adc193 · outbound

This paper cites Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models.

Hardware-Efficient Attention for Fast Decoding Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.119922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.119922Z digest=sha256:9ced9d430c7621ad29c45fcc52aaaecc6925fa721794b087ee199d6bb8ec2263

Observation f701fc1d-ca5f-4519-8b73-65db4e436b18 · outbound

This paper cites Decoupled Weight Decay Regularization.

Hardware-Efficient Attention for Fast Decoding Decoupled Weight Decay Regularization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.164678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.164678Z digest=sha256:87e520813f300a2c03b715541a7164c93b6bb26a91f944485ce2edcee6021b52

Observation 503095d2-ec09-4a0f-a98b-a455f59f9589 · outbound

This paper cites Fineweb-edu: The finest collection of educational content, 2024.

Hardware-Efficient Attention for Fast Decoding Fineweb-edu: The finest collection of educational content, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:38.436599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:31.215706Z digest=sha256:1ff4f9e99e67a478299877526d8ea8edff5c7e0afcfc3f463302c0d4e07bb37f

Observation 638a7032-594d-4b74-9745-4f0ad845af01 · outbound

This paper cites TransMLA: Multi-Head Latent Attention Is All You Need.

Hardware-Efficient Attention for Fast Decoding TransMLA: Multi-Head Latent Attention Is All You Need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.245399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.245399Z digest=sha256:538ca77d6fd04b1b4bc9db9a1b7da1adda0fde1aa331c196a2a121e7b7040f4b

Observation 8135a6be-8134-4eaf-b8c6-15fbe0896b87 · outbound

This paper cites The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation , 2025.

Hardware-Efficient Attention for Fast Decoding The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation , 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:38.252804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:31.277572Z digest=sha256:dd9023801149cb4bcd822b6af289530906ce08f79f7f222b1d206d4483cafaff

Observation 09bf71be-fb68-43d3-8ab1-3211e8692aa4 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Hardware-Efficient Attention for Fast Decoding Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.323457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.323457Z digest=sha256:1735f3fa2900d76357c29052d4caf9b87b56c75e49c7c9233eebd59deaa46170

Observation 7c0562cb-be3a-4f63-ab2d-2bbf53aedc0b · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

Hardware-Efficient Attention for Fast Decoding Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.382482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.382482Z digest=sha256:900a255ce20469a25426c91bc3b9d1c92733379bf76e1c29cb15da1df9fcbe99

Observation b98d2e69-c40f-4c0f-a376-a80a2b2b4410 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Hardware-Efficient Attention for Fast Decoding Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.424098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.424098Z digest=sha256:2f2cb549c60223c26af8ee0db002cdad653655319719b69c18ac00a9ad2c6add

Observation 6f0e748c-836d-45c4-a33c-a574081d2c92 · outbound

This paper cites NVIDIA H100 tensor core gpu architecture.

Hardware-Efficient Attention for Fast Decoding NVIDIA H100 tensor core gpu architecture

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:38.088689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:31.475639Z digest=sha256:d0b8e3ed0f234ea6d55fbfb5a386e42e830d13d49df820374f1a9b8fe91a9dda

Observation 5d24c8e9-66ce-43f3-ae8d-bca541d95507 · outbound

This paper cites NVIDIA Blackwell architecture technical brief.

Hardware-Efficient Attention for Fast Decoding NVIDIA Blackwell architecture technical brief

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:37.939319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:31.582198Z digest=sha256:e138ac47ff5287b7ccf7ea4daec5384d2f1e034deb84018b83820176e595dd55

Observation ad2625d0-5d63-4535-ab96-701d522e82d1 · outbound

This paper cites Nvlink, 2024.

Hardware-Efficient Attention for Fast Decoding Nvlink, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:37.755767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:31.690658Z digest=sha256:d0d79dc6526094ed4f09ddbea376f3d8c8b05a178674e25b0f7b786d1eab644e

Observation 96183437-57c5-4352-a08c-630affcfbdf6 · outbound

This paper cites Reducing shared memory footprint to leverage high throughput on Tensor Cores and its flexible API extension library.

Hardware-Efficient Attention for Fast Decoding Reducing shared memory footprint to leverage high throughput on Tensor Cores and its flexible API extension library

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:32:36.174838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:31.812951Z digest=sha256:d5e64ddf5db2e7da11fafffcf4f114e23e03636304cbdb7457c935cb15a2039b

Observation 2915d83a-6956-42af-aa26-f424888facb7 · outbound

This paper cites OpenAI o1 System Card.

Hardware-Efficient Attention for Fast Decoding OpenAI o1 System Card

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.887862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.887862Z digest=sha256:64c5605db6a2c2ba53a8d9fb2cd732721ab63a32bc947194ffc7270b0a3d05c0

Observation 1702db5c-1acd-43d7-b486-f84714a7ed3b · outbound

This paper cites Efficiently Scaling Transformer Inference.

Hardware-Efficient Attention for Fast Decoding Efficiently Scaling Transformer Inference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.926378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.926378Z digest=sha256:ba7ac882957a8135e167ecc13f90add204bdb0de53910766e2e7b893e66d0d9c

Observation 0bcb144d-aa00-4007-b461-f4a2a6e581ba · outbound

This paper cites Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference.

Hardware-Efficient Attention for Fast Decoding Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.995104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.995104Z digest=sha256:1359d73f632f3bef4fa41119656b2fcef7cb75836081a5c519044516972346ae

Observation 788ca0a5-0716-46b8-891c-d7d56a934921 · outbound

This paper cites Winogrande : An adversarial winograd schema challenge at scale.

Hardware-Efficient Attention for Fast Decoding Winogrande : An adversarial winograd schema challenge at scale

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:37.608442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:32.090714Z digest=sha256:332d6172d4ace586c4159f36572e286c38547a8702c4d10066237f302f8fa720

Observation ead51bd7-7598-479f-a897-c57912814f90 · outbound

This paper cites Eigen Attention: Attention in Low-Rank Space for KV Cache Compression.

Hardware-Efficient Attention for Fast Decoding Eigen Attention: Attention in Low-Rank Space for KV Cache Compression

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.138676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.138676Z digest=sha256:dcb2497527ede9fa74d853091a65bec9b26a4db7420c588c4ee8945c70413160

Observation 3e3fdf22-8490-4f41-8d96-586d4a78f0f0 · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.

Hardware-Efficient Attention for Fast Decoding Flashattention-3: Fast and accurate attention with asynchrony and low-precision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:37.483734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:32.189674Z digest=sha256:c791552482ed45da99e7a633bae7eb4c5e8552df043c771159294c0a605ebe0c

Observation 6f1e4605-2836-4295-a4ac-b2b98d1da2b0 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

Hardware-Efficient Attention for Fast Decoding FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.266167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.266167Z digest=sha256:785c88255e761720f65776a97979f6c7169331990fb8f3592358c0650770fbb3

Observation e5b00151-640b-42bb-9877-83f742b3eea4 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Hardware-Efficient Attention for Fast Decoding Fast Transformer Decoding: One Write-Head is All You Need

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.363913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.363913Z digest=sha256:878ad4f0bb7c79f084e2f7c635d615248253f4afde6594252fb60fc196442227

Observation 62940194-287f-414c-bdc0-0abb9eb19d8f · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

Hardware-Efficient Attention for Fast Decoding FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.440200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.440200Z digest=sha256:a332c410ebd78c019a1c30ff0f0c3c102374d178c2136b1fb9b8ab5b039e0edc

Observation 5adf0400-21db-4fc6-8c8a-0c87a08d3785 · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

Hardware-Efficient Attention for Fast Decoding Loki: Low-rank Keys for Efficient Sparse Attention

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.523852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.523852Z digest=sha256:a9a460fb038deebf426d6883c39323ec82a82a81d15a8a2bb98cb9d6c874db64

Observation 5153ecd2-bccb-4103-b046-66d22c674311 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Hardware-Efficient Attention for Fast Decoding Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.582152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.582152Z digest=sha256:b40ffd475ef09e1b23625ec5c2fe03c3fcc8decca608c08ce721678a05bf887f

Observation 19705616-bf2d-4dd9-9de1-b0e7f5552dd9 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Hardware-Efficient Attention for Fast Decoding RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.701752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.701752Z digest=sha256:51319d0d71a1969e40317742afc6fd5eb2987988f62376896316b0554f502eeb

Observation 97826bf4-382f-47e4-a23f-1c48ba6253a6 · outbound

This paper cites Seesaw: High-throughput LLM Inference via Model Re-sharding.

Hardware-Efficient Attention for Fast Decoding Seesaw: High-throughput LLM Inference via Model Re-sharding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.805859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.805859Z digest=sha256:3f610703aa07d18c0530f10feb6e0a93b97cbfc71f6c061e4f1f67fe7ce649ff

Observation 2393bfa6-f739-4a74-81c7-0b8bee2b5202 · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

Hardware-Efficient Attention for Fast Decoding ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:32.885907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:32.885907Z digest=sha256:9acb2fb3672e8b0c4d78bbf637ae046b2c8459a65ba48ef4ddd9746d7cbcb137

Observation f5d19750-7ca1-40a7-8855-f457f13434a0 · outbound

This paper cites CUTLASS , January 2023.

Hardware-Efficient Attention for Fast Decoding CUTLASS , January 2023

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:37.348425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:32.996785Z digest=sha256:5daaa5cf04357461762c9983916ca9f3f989139664d585c87f4373507b69c28e

Observation d1c09476-aaa9-4969-8441-a25821869572 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Hardware-Efficient Attention for Fast Decoding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.058043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.058043Z digest=sha256:bfb05c093dc54b13ecbd5932589059fd12a0e6518492376b4c6a02b960b9051f

Observation 24488eff-f07b-4f01-9b36-09c7b3a08eb1 · outbound

This paper cites Attention is all you need.

Hardware-Efficient Attention for Fast Decoding Attention is all you need

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.173385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.173385Z digest=sha256:39268615f7c922c3ebb0735bed82145ac3b5051c5536b6f5a7cdf1fc1f3ee4b1

Observation d11f10f8-e395-435c-9a61-dbafc7f0840d · outbound

This paper cites Gpt-j-6b: a 6 billion parameter autoregressive language model, 2021.

Hardware-Efficient Attention for Fast Decoding Gpt-j-6b: a 6 billion parameter autoregressive language model, 2021

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:37.172316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:33.282513Z digest=sha256:c0938cd0dc3115f2e2e7ce9b0924bce2f53dd4eb930eb601fd84b528715826f6

Observation e15629f2-7a49-41eb-85b5-df5ef3621fc5 · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

Hardware-Efficient Attention for Fast Decoding RedPajama: an Open Dataset for Training Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.391740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.391740Z digest=sha256:a00d817ae4a23b755240810a2b980e0a765cc627b297c87d5dbb8f55d88aa5b0

Observation f2c4e85d-efab-4fde-a5a5-a4e3263c2ed7 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

Hardware-Efficient Attention for Fast Decoding Crowdsourcing Multiple Choice Science Questions

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.520521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.520521Z digest=sha256:6778415cd92840a45ee04b24b26622fe541c09b2e46393ad72678a02611d6515

Observation 90e3223e-4804-4ed6-bc09-e0459f95b255 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.

Hardware-Efficient Attention for Fast Decoding Roofline: an insightful visual performance model for multicore architectures

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.655981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.655981Z digest=sha256:9f4138056124dae55914d97e18ab47d032327b0a03bfce1173afc40b8a9a29c1

Observation 692ab0ed-be0c-493f-b724-84e4e1f72e74 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Hardware-Efficient Attention for Fast Decoding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.762529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.762529Z digest=sha256:3f4f92848219d79cc45ceee8d8c2fdceb3587247383cf6bf2f25eb2dc51e06f9

Observation 9110ff98-fbaa-4836-9a8b-59490b1c9cd7 · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

Hardware-Efficient Attention for Fast Decoding Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.884257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.884257Z digest=sha256:9283d3d98910f34139d6b7fa5751675655e7cf192fd2e0c9db0be0b4e38b155b

Observation 6d6193d9-49dc-49f7-887e-729b9f1582cb · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Hardware-Efficient Attention for Fast Decoding Efficient Streaming Language Models with Attention Sinks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.994365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.994365Z digest=sha256:fca1e689a33947bcb7a8fc3926007085bd037b08ec8a911f7ed3d81a9ad120be

Observation 3f9dcf9c-f951-456e-8398-1e779737b3f8 · outbound

This paper cites Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering.

Hardware-Efficient Attention for Fast Decoding Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:34.109663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:34.109663Z digest=sha256:e951e1ef3693ecec63d6eb7b522b9025c807c14b5450035b0a2b7d647ab68726

Observation b6fa272e-353d-4eda-b049-f0269a2789eb · outbound

This paper cites Rope to nope and back again: A new hybrid attention strategy, 2025.

Hardware-Efficient Attention for Fast Decoding Rope to nope and back again: A new hybrid attention strategy, 2025

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:34.255027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:34.255027Z digest=sha256:e5e5476f62961de2901f7f1488f3c08c6b660062638ea688c37599c6baa82be9

Observation 8006c060-ad31-4f82-9c8d-43160b9684fa · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Hardware-Efficient Attention for Fast Decoding Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:34.407299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:34.407299Z digest=sha256:8c5df75409e9e916ed3a26819879b04bd810f1391aecf83443016851b4c74b18

Observation d5467938-fa24-4ea2-884c-4e003de5fc59 · outbound

This paper cites Effectively Compress KV Heads for LLM.

Hardware-Efficient Attention for Fast Decoding Effectively Compress KV Heads for LLM

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:34.524396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:34.524396Z digest=sha256:cff6e7c0598d8ea98db6d8d934dda335cc8b9dfe1771130971895d4162a27760

Observation 8762b690-7f6b-45bc-9cdc-fcf2bc0739ca · outbound

This paper cites Affordable Generative Agents.

Hardware-Efficient Attention for Fast Decoding Affordable Generative Agents

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:34.655835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:34.655835Z digest=sha256:49c269b5d043530d5622a82713dbb6c866a31167742732c340e9d6cf86cf4796

Observation 65eb3793-a028-4e6d-9bba-e34299118ba5 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Hardware-Efficient Attention for Fast Decoding Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:34.791543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:34.791543Z digest=sha256:854362bfc5a18afefc379d7f690323441ab62242fc55e867791d954ec6c0fafe

Observation 77a7190d-872d-4076-96af-5d8ba2eb9908 · outbound

This paper cites HellaSwag : Can a machine really finish your sentence? In Anna Korhonen, David R.

Hardware-Efficient Attention for Fast Decoding HellaSwag : Can a machine really finish your sentence? In Anna Korhonen, David R

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:34.945787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:34.945787Z digest=sha256:5e1a556665f3cee1b2ade9d06d81010a03a4c70b696af12e582a6e2800f392a3

Observation 58724d79-09c6-4ab7-8ef7-da461af97e40 · outbound

This paper cites Tensor product attention is all you need, 2025.

Hardware-Efficient Attention for Fast Decoding Tensor product attention is all you need, 2025

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:35.042714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:35.042714Z digest=sha256:33d1c57274413244d126e7d4b6911a65a4864d8ad35f389e96accea08c9d1054

Observation b36fd59c-01be-41ec-84e9-33fb8e59e3e3 · outbound

This paper cites H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

Hardware-Efficient Attention for Fast Decoding H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:35.147561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:35.147561Z digest=sha256:1d1160650c64926ae5683e98a34e7c36ca2f9c035e22a301479447f37786a9ef

Observation 2144f0f0-0145-4270-afb6-4cc0b32e59be · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Hardware-Efficient Attention for Fast Decoding SGLang: Efficient Execution of Structured Language Model Programs

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:35.275765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:35.275765Z digest=sha256:3d34873e0657da319d6429258e53b8bce8ba723cb20a7d202bad36d48435defb

Observation 830b33b4-b87e-4a4d-a26e-bfd33221fa38 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.

Hardware-Efficient Attention for Fast Decoding Sglang: Efficient execution of structured language model programs

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:32:36.995246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:32:35.413009Z digest=sha256:70a52d83617f37078ab3868983ec9c069672f14e99e6606d59774be629e4b19b

Pith citing papers

Observation 5f4dc038-a52c-45e7-8213-b076abe14c63 · inbound

TransMLA: Multi-Head Latent Attention Is All You Need cites this paper.

TransMLA: Multi-Head Latent Attention Is All You Need Hardware-Efficient Attention for Fast Decoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T11:48:01.180107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:48:01.180107Z digest=sha256:1b4e58059d7f65388ba76976a763cd673e687ee1b2f6a158b4e78019d190e666

Observation 292cce17-8c9a-4178-887e-7497c8c511e0 · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hardware-Efficient Attention for Fast Decoding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.737715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.737715Z digest=sha256:5c96ffc2fe7f36ebedf2d5a5704523b0a9aeed263c9e110c7f094915ba41d102

Observation 3a655ef5-cf0b-481a-a0e3-b783519b25c6 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Hardware-Efficient Attention for Fast Decoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:08.095113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:08.095113Z digest=sha256:ab0c4b9297b294b7cb42317a7f36dce9edde2562b369ed8d93d1b90be7fb3b71

Observation c97ed32f-bce7-49e3-9860-9a60cc640eea · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Hardware-Efficient Attention for Fast Decoding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.085058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:a8209c9f3013455fe1be702095a142780407b5fe020a14ce51018a711ff9a403

Observation bc4302b9-6ff6-4770-8263-b52508b85bad · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Hardware-Efficient Attention for Fast Decoding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.357138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:4a54e32b8c273cc52078630f8d2dd2829fa30be772e3fda50f75fe0441c80120