Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:32:35.413009Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 5 inbound Pith citation observations for arXiv:2505.21487.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:32:35.413009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T11:48:01.180107Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T01:57:51.329269Z
85 of 85 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 43121bc2-b3cf-48bd-b5ec-d989455c453e · outbound
Hardware-Efficient Attention for Fast Decoding SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41699e1c-1716-4d8f-90e2-d7bf64aa9f0f · outbound
Hardware-Efficient Attention for Fast Decoding GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8abc7c3-bda3-4139-8f81-e0cb39023140 · outbound
Hardware-Efficient Attention for Fast Decoding DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71edab39-71ef-451a-b4ee-82e3c796a122 · outbound
Hardware-Efficient Attention for Fast Decoding How to scale your model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 006a6026-b703-4b2a-9bb7-af46222eabbe · outbound
Hardware-Efficient Attention for Fast Decoding Round and Round We Go! What makes Rotary Positional Encodings useful?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0705c79f-e322-49e3-8af2-b710e8b0cbde · outbound
Hardware-Efficient Attention for Fast Decoding Singe: Leveraging warp specialization for high performance on gpus
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c62b84a-73c0-4045-b3a2-e4789ca473d8 · outbound
Hardware-Efficient Attention for Fast Decoding Cosmopedia, February 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52226de7-e862-4371-ba00-cd0cd4de1039 · outbound
Hardware-Efficient Attention for Fast Decoding PIQA : reasoning about physical commonsense in natural language
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b89d84-b0e0-4fb9-b254-9faf7d4b55d3 · outbound
Hardware-Efficient Attention for Fast Decoding GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd97bfa7-81dd-4579-b112-b259b2a85479 · outbound
Hardware-Efficient Attention for Fast Decoding Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6dc91aa-aca2-434d-bc3b-409bccdb244c · outbound
Hardware-Efficient Attention for Fast Decoding Language Models are Few-Shot Learners
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5ea5fb-a49e-40fb-b73a-2d4c9cb73d6a · outbound
Hardware-Efficient Attention for Fast Decoding Palu: Compressing KV-Cache with Low-Rank Projection
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd2ed83-9ca6-45b0-9e61-9a775d604bad · outbound
Hardware-Efficient Attention for Fast Decoding What rotary position embedding can tell us: Identifying query and key weights corresponding to basic syntactic or high-level semantic information
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1256f6-b547-4074-af20-9c032096d05d · outbound
Hardware-Efficient Attention for Fast Decoding MagicPIG: LSH Sampling for Efficient LLM Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be7ee620-6b3d-4922-8314-c5a0505f2535 · outbound
Hardware-Efficient Attention for Fast Decoding FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c9086a-1c49-4563-97dd-31c451a00e4d · outbound
Hardware-Efficient Attention for Fast Decoding Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3c6219a-08ae-40e7-996d-9744c11d397f · outbound
Hardware-Efficient Attention for Fast Decoding FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7970ad1-5764-47cf-ab55-18c67e2fed42 · outbound
Hardware-Efficient Attention for Fast Decoding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e5787e-a7a4-4627-a326-c57ad1e02fa2 · outbound
Hardware-Efficient Attention for Fast Decoding DeepSeek-V3 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7968316b-9305-4349-8005-76947f67185c · outbound
Hardware-Efficient Attention for Fast Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b8ea07b-7e80-4384-8a62-49119f32b53c · outbound
Hardware-Efficient Attention for Fast Decoding The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 139f2bc8-0929-4599-b437-cfbf0a2f0cfd · outbound
Hardware-Efficient Attention for Fast Decoding AI and Memory Wall
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff636470-7567-4fcc-aacd-2e29786a2d98 · outbound
Hardware-Efficient Attention for Fast Decoding What Your DRAM Power Models Are Not Telling You: Lessons from a Detailed Experimental Study
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d741e4a5-9158-4ac3-a163-4c04e4e0f317 · outbound
Hardware-Efficient Attention for Fast Decoding Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c615176-c5f5-4637-b98a-a7516149c1d0 · outbound
Hardware-Efficient Attention for Fast Decoding The Llama 3 Herd of Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2955cf70-99a0-4063-95b3-5f8086e443cc · outbound
Hardware-Efficient Attention for Fast Decoding Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9f7dd2d3-10c8-40d3-a70c-dec91328cd2a · outbound
Hardware-Efficient Attention for Fast Decoding Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f576932a-1cf8-4630-82db-d0af173eecb0 · outbound
Hardware-Efficient Attention for Fast Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 944ab807-ae86-4bad-a3b8-b1eb23e8c927 · outbound
Hardware-Efficient Attention for Fast Decoding Measuring massive multitask language understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b846c7a3-8c26-4820-9f30-fd34b7596dbf · outbound
Hardware-Efficient Attention for Fast Decoding KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e12996-15d3-485f-bf64-f646d8edfe21 · outbound
Hardware-Efficient Attention for Fast Decoding Multi-matrix Factorization Attention
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1be02b3e-8259-4ea6-94a7-4198e332a6e4 · outbound
Hardware-Efficient Attention for Fast Decoding Data Movement Is All You Need: A Case Study on Optimizing Transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bfcb9ad-1b43-4722-8b3f-0c50f89efe30 · outbound
Hardware-Efficient Attention for Fast Decoding Towards economical inference: Enabling deepseek's multi-head latent attention in any transformer-based llms, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5f0562a-5d2d-4c4f-aa66-386544c71ac7 · outbound
Hardware-Efficient Attention for Fast Decoding Weight decay induces low-rank attention layers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 661446ef-c137-4989-9696-e2e1da24b08f · outbound
Hardware-Efficient Attention for Fast Decoding Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21995711-3da1-4c6a-8a20-0e9cfaadaa91 · outbound
Hardware-Efficient Attention for Fast Decoding Software pipelining: An effective scheduling technique for vliw machines
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f71a8029-0f86-4186-8e08-25c87f10d93e · outbound
Hardware-Efficient Attention for Fast Decoding Flashmla: Efficient mla decoding kernels
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eff9c8b8-d14c-434b-b783-8a38e6bee5b5 · outbound
Hardware-Efficient Attention for Fast Decoding PyTorch Distributed: Experiences on Accelerating Data Parallel Training
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e4c7db0-b686-4cbb-8373-b9af37adc193 · outbound
Hardware-Efficient Attention for Fast Decoding Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f701fc1d-ca5f-4519-8b73-65db4e436b18 · outbound
Hardware-Efficient Attention for Fast Decoding Decoupled Weight Decay Regularization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 503095d2-ec09-4a0f-a98b-a455f59f9589 · outbound
Hardware-Efficient Attention for Fast Decoding Fineweb-edu: The finest collection of educational content, 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 638a7032-594d-4b74-9745-4f0ad845af01 · outbound
Hardware-Efficient Attention for Fast Decoding TransMLA: Multi-Head Latent Attention Is All You Need
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8135a6be-8134-4eaf-b8c6-15fbe0896b87 · outbound
Hardware-Efficient Attention for Fast Decoding The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation , 2025
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 09bf71be-fb68-43d3-8ab1-3211e8692aa4 · outbound
Hardware-Efficient Attention for Fast Decoding Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0562cb-be3a-4f63-ab2d-2bbf53aedc0b · outbound
Hardware-Efficient Attention for Fast Decoding Orca: Progressive Learning from Complex Explanation Traces of GPT-4
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98d2e69-c40f-4c0f-a376-a80a2b2b4410 · outbound
Hardware-Efficient Attention for Fast Decoding Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0e748c-836d-45c4-a33c-a574081d2c92 · outbound
Hardware-Efficient Attention for Fast Decoding NVIDIA H100 tensor core gpu architecture
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5d24c8e9-66ce-43f3-ae8d-bca541d95507 · outbound
Hardware-Efficient Attention for Fast Decoding NVIDIA Blackwell architecture technical brief
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ad2625d0-5d63-4535-ab96-701d522e82d1 · outbound
Hardware-Efficient Attention for Fast Decoding Nvlink, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 96183437-57c5-4352-a08c-630affcfbdf6 · outbound
Hardware-Efficient Attention for Fast Decoding Reducing shared memory footprint to leverage high throughput on Tensor Cores and its flexible API extension library
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2915d83a-6956-42af-aa26-f424888facb7 · outbound
Hardware-Efficient Attention for Fast Decoding OpenAI o1 System Card
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1702db5c-1acd-43d7-b486-f84714a7ed3b · outbound
Hardware-Efficient Attention for Fast Decoding Efficiently Scaling Transformer Inference
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bcb144d-aa00-4007-b461-f4a2a6e581ba · outbound
Hardware-Efficient Attention for Fast Decoding Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788ca0a5-0716-46b8-891c-d7d56a934921 · outbound
Hardware-Efficient Attention for Fast Decoding Winogrande : An adversarial winograd schema challenge at scale
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ead51bd7-7598-479f-a897-c57912814f90 · outbound
Hardware-Efficient Attention for Fast Decoding Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3fdf22-8490-4f41-8d96-586d4a78f0f0 · outbound
Hardware-Efficient Attention for Fast Decoding Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6f1e4605-2836-4295-a4ac-b2b98d1da2b0 · outbound
Hardware-Efficient Attention for Fast Decoding FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b00151-640b-42bb-9877-83f742b3eea4 · outbound
Hardware-Efficient Attention for Fast Decoding Fast Transformer Decoding: One Write-Head is All You Need
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62940194-287f-414c-bdc0-0abb9eb19d8f · outbound
Hardware-Efficient Attention for Fast Decoding FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5adf0400-21db-4fc6-8c8a-0c87a08d3785 · outbound
Hardware-Efficient Attention for Fast Decoding Loki: Low-rank Keys for Efficient Sparse Attention
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5153ecd2-bccb-4103-b046-66d22c674311 · outbound
Hardware-Efficient Attention for Fast Decoding Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19705616-bf2d-4dd9-9de1-b0e7f5552dd9 · outbound
Hardware-Efficient Attention for Fast Decoding RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97826bf4-382f-47e4-a23f-1c48ba6253a6 · outbound
Hardware-Efficient Attention for Fast Decoding Seesaw: High-throughput LLM Inference via Model Re-sharding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2393bfa6-f739-4a74-81c7-0b8bee2b5202 · outbound
Hardware-Efficient Attention for Fast Decoding ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d19750-7ca1-40a7-8855-f457f13434a0 · outbound
Hardware-Efficient Attention for Fast Decoding CUTLASS , January 2023
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d1c09476-aaa9-4969-8441-a25821869572 · outbound
Hardware-Efficient Attention for Fast Decoding Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24488eff-f07b-4f01-9b36-09c7b3a08eb1 · outbound
Hardware-Efficient Attention for Fast Decoding Attention is all you need
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d11f10f8-e395-435c-9a61-dbafc7f0840d · outbound
Hardware-Efficient Attention for Fast Decoding Gpt-j-6b: a 6 billion parameter autoregressive language model, 2021
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e15629f2-7a49-41eb-85b5-df5ef3621fc5 · outbound
Hardware-Efficient Attention for Fast Decoding RedPajama: an Open Dataset for Training Large Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c4e85d-efab-4fde-a5a5-a4e3263c2ed7 · outbound
Hardware-Efficient Attention for Fast Decoding Crowdsourcing Multiple Choice Science Questions
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e3223e-4804-4ed6-bc09-e0459f95b255 · outbound
Hardware-Efficient Attention for Fast Decoding Roofline: an insightful visual performance model for multicore architectures
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 692ab0ed-be0c-493f-b724-84e4e1f72e74 · outbound
Hardware-Efficient Attention for Fast Decoding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9110ff98-fbaa-4836-9a8b-59490b1c9cd7 · outbound
Hardware-Efficient Attention for Fast Decoding Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d6193d9-49dc-49f7-887e-729b9f1582cb · outbound
Hardware-Efficient Attention for Fast Decoding Efficient Streaming Language Models with Attention Sinks
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f9dcf9c-f951-456e-8398-1e779737b3f8 · outbound
Hardware-Efficient Attention for Fast Decoding Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6fa272e-353d-4eda-b049-f0269a2789eb · outbound
Hardware-Efficient Attention for Fast Decoding Rope to nope and back again: A new hybrid attention strategy, 2025
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8006c060-ad31-4f82-9c8d-43160b9684fa · outbound
Hardware-Efficient Attention for Fast Decoding Gated Linear Attention Transformers with Hardware-Efficient Training
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5467938-fa24-4ea2-884c-4e003de5fc59 · outbound
Hardware-Efficient Attention for Fast Decoding Effectively Compress KV Heads for LLM
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8762b690-7f6b-45bc-9cdc-fcf2bc0739ca · outbound
Hardware-Efficient Attention for Fast Decoding Affordable Generative Agents
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65eb3793-a028-4e6d-9bba-e34299118ba5 · outbound
Hardware-Efficient Attention for Fast Decoding Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a7190d-872d-4076-96af-5d8ba2eb9908 · outbound
Hardware-Efficient Attention for Fast Decoding HellaSwag : Can a machine really finish your sentence? In Anna Korhonen, David R
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58724d79-09c6-4ab7-8ef7-da461af97e40 · outbound
Hardware-Efficient Attention for Fast Decoding Tensor product attention is all you need, 2025
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b36fd59c-01be-41ec-84e9-33fb8e59e3e3 · outbound
Hardware-Efficient Attention for Fast Decoding H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2144f0f0-0145-4270-afb6-4cc0b32e59be · outbound
Hardware-Efficient Attention for Fast Decoding SGLang: Efficient Execution of Structured Language Model Programs
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 830b33b4-b87e-4a4d-a26e-bfd33221fa38 · outbound
Hardware-Efficient Attention for Fast Decoding Sglang: Efficient execution of structured language model programs
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5f4dc038-a52c-45e7-8213-b076abe14c63 · inbound
TransMLA: Multi-Head Latent Attention Is All You Need Hardware-Efficient Attention for Fast Decoding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292cce17-8c9a-4178-887e-7497c8c511e0 · inbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hardware-Efficient Attention for Fast Decoding
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a655ef5-cf0b-481a-a0e3-b783519b25c6 · inbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Hardware-Efficient Attention for Fast Decoding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c97ed32f-bce7-49e3-9860-9a60cc640eea · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Hardware-Efficient Attention for Fast Decoding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bc4302b9-6ff6-4770-8263-b52508b85bad · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Hardware-Efficient Attention for Fast Decoding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.